Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,986 characters · 17 sections · 22 citation commands
A Random Attention Model
Keywords: revealed preference, limited attention models, random utility models, nonparametric identification, partial identification.
\thispagestyle{empty}
\doublespacing \setcounter{page}{1} \pagestyle{plain}
Revealed preference theory is not only a cornerstone of modern economics, but is also the source of important theoretical, methodological and policy implications for many social and behavioral sciences. This theory aims to identify the preferences of a decision maker (e.g., an individual or a firm) from her observed choices (e.g., buying a house or hiring a worker). In its classical formulation, revealed preference theory assumes that the decision maker selects the best available option after full consideration of all possible alternatives presented to her. This assumption leads to specific testable implications based on observed choice patterns but, unfortunately, empirical testing of classical revealed preference theory shows that it is not always compatible with observed choice behavior Hauser_Wernerfelt_1990_JCR,Goeree_2008_ECMA,vanNierop_et_al_2010_JMR,Honka_Hortacsu_Vitorino_2017_RAND. For example, \citet*{Reutskaja_et_al_2011_AER} provides interesting experimental evidence against the full attention assumption using eye tracking and choice data.
Motivated by these findings, and the fact that certain theoretically important and empirically relevant choice patterns can not be explained using classical revealed preference theory based on full attention, scholars have proposed other economic models of choice behavior. An alternative is the limited attention model \citep*{Masatlioglu-Nakajima-Ozbay_2012_AER, Lleras_et_al_2017_JET, Dean_Kibris_Masatlioglu_2017_JET}, where decision makers are assumed to select the best available option from a subset of all possible alternatives, known as the consideration set. This framework takes the formation of the consideration set, also known as attention rule or consideration map, as unobservable and hence as an intrinsic feature of the decision maker. Nonetheless, it is possible to develop a fruitful theory of revealed preference within this framework, employing only mild and intuitive nonparametric restrictions on how the decision maker decides to focus attention on specific subsets of all possible alternatives presented to her.
Until very recently, limited attention models have been deterministic, a feature that diminished their empirical applicability: testable implications via revealed preference have relied on the assumption that the decision maker pays attention to the same subset of options every time she is confronted with the same set of available alternatives. This requires that, for example, an online shopper uses always the same keyword and the same search engine (e.g. Google) on the same platform (e.g. tablet) to look for a product. This is obviously restrictive, and can lead to predictions that are inconsistent with observed choice behavior. Aware of this fact, a few scholars have improved deterministic limited attention models by allowing for stochastic attention \citep*{ Manzini-Mariotti_2014_ECMA,Aguiar_2015,Brady-Rehbeck_2016_ECMA,Horan_2018}, which permits the decision maker to pay attention to different subsets with some non-zero probability given the same set of alternatives to choose from. All available results in this literature proceed by first parameterizing the attention rule (i.e., committing to a particular parametric attention rule), and then studying the revealed preference implications of these parametric models.
In contrast to earlier approaches, we introduce a Random Attention Model (RAM) where we abstain from any specific parametric (stochastic) attention rule, and instead consider a large class of nonparametric random attention rules. Our model imposes one intuitive condition, termed Monotonic Attention, which is satisfied by many stochastic attention rules. Given that consideration sets are unobservable, this feature is crucial for applicability of our revealed preference results, as our findings and empirical implications are valid under many different, particular attention rules that could be operating in the background. In other words, our revealed preference results are derived from nonparametric restrictions on the attention rule and hence are more robust to misspecification biases.
RAM is best suited for eliciting information about the preference ordering of a single decision-making unit when her choices are observed repeatedly.\footnote{The finding that individual choices frequently exhibit randomness was first reported in \citet*{Tversky_1969_PR} and has now been illustrated by \citet*{Agranov_Ortoleva_2017} and numerous other studies. Similar to our work, \citet*{Manzini-Mariotti_2014_ECMA}, \citet*{Fudenberg_Iijima_Strzalecki_2015_ECMA}, and \citet*{Brady-Rehbeck_2016_ECMA}, among others, have developed models which allow the analyst to reveal information about the agent's preferences from her observed random choices.} For example, scanner data keeps track of the same single consumer's purchases across repeated visits, where the grocery store adjusts product varieties and arrangements regularly. Another example is web advertising on digital platforms, such as search engines or shopping sites, where not only abundant records from each individual decision maker are available, but also it is common to see manipulations/experiments altering the options offered to them. A third example is given in \citet*{Kawaguchi-Uetake-Watanabe_2016_wp}, where large data on each consumer's choices from vending machines (with varying product availability) is analyzed. In addition, our model can be used empirically with aggregate data on a group of distinct decision makers, provided each of them may differ on what they pay attention to but all share the same preference.
Our key identifying assumption, Monotonic Attention, restricts the possibly stochastic attention formation process in a very intuitive way: each consideration set competes for the decision maker's attention, and hence the probability of paying attention to a particular subset is assumed not to decrease when the total number of possible consideration sets decreases. We show that this single nonparametric assumption is general enough to nest most (if not all) previously proposed deterministic and random limited attention models. Furthermore, under our proposed monotonic attention assumption, we are able to develop a theory of revealed preference, obtain specific testable implications, and (partially) identify the underlying preferences of the decision maker by investigating her observed choice probabilities. Our revealed preference results are applicable to a wide range of attention rules, including the parametric ones currently available in the literature which, as we show, satisfy the monotonic attention assumption.
Based on these theoretical findings, we also develop econometric results for identification, estimation, and inference of the decision maker's preferences, as well as specification testing of RAM. We show that RAM implies that the set of partially identified preference orderings containing the decision maker's true preferences is equivalent to a set of inequality restrictions on the choice probabilities (one for each preference ordering in the identified set). This result allows us to employ the identifiable/estimable choice probabilities to (i) develop a model specification test (i.e., test whether there exists a non-empty set of preference orderings compatible with RAM), (ii) conduct hypothesis testing on specific preference orderings (i.e., test whether the inequality constraints on the choice probabilities are satisfied), and (iii) develop confidence sets containing the decision maker's true preferences with pre-specified coverage (i.e., via test inversion). Our econometric methods rely on ideas and results from the literature on partially identified models and moment inequality testing: see \citet*{Canay_Shaikh_2017_ChBook}, \citet*{Ho_Rosen_2017_ChBook} and Molinari_2019_Handbook for recent reviews and further references.
RAM is fully nonparametric and agnostic because it relies on the monotonic attention assumption only. As a consequence, it may lead to relatively weak testable implications in some applications, that is, “little” revelation or a “large” identified set of preferences. However, RAM also provides a basis for incorporating additional (parametric and) nonparametric restrictions that can substantially improve identification power. In this paper, we illustrate how RAM can be combined with additional, mild nonparametric restrictions to tighten identification in non-trivial ways: in Section (ref), we incorporate an additional restriction on attention rule for binary choice problems, and show that this alone leads to important revelation improvements within RAM. We also illustrate this result numerically in our simulation study.
Finally, we implement our estimation and inference methods in the general-purpose software package ramchoice for R---see \url{https://cran.r-project.org/package=ramchoice} for details. Our novel identification results allow us to develop inference methods that avoid optimization over the possibly high-dimensional space of attention rules, leading to methods that are very fast and easy to implement when applied to realistic empirical problems. See the Supplemental Appendix for numerical evidence.
Our work contributes to both economic theory and econometrics. We describe several examples covered by our model in Section (ref) after we introduce our proposed RAM. We also discuss in detail the connections and distinctions between this paper and the economic theory literature in Section SA.1 of the Supplemental Appendix. In particular, we show how RAM nests and/or connects to the recent work by \citet*{Manzini-Mariotti_2014_ECMA}, \citet*{Brady-Rehbeck_2016_ECMA}, \citet*{Gul_Natenzon_Pesendorfer_2014_ECMA}, \citet*{Echenique_Saito_Tserenjigmid_2018}, \citet*{Echenique_Saito_2019}, \citet*{Fudenberg_Iijima_Strzalecki_2015_ECMA}, and \citet*{Aguiar_Boccardi_Dean_2016_JET}, among others.
This paper is also related to a rich econometric literature on nonparametric identification, estimation and inference both in the specific context of random utility models, and more generally. See \citet*{Matzkin_2013_ARE} for a review and further references on nonparametric identification, \citet*{Hausman-Newey_2017_ARE} for a recent review and further references on nonparametric welfare analysis, and \citet*{Blundell-Kristensen-Matzkin_2014_JOE}, Kawaguchi_2017_JoE, \citet*{Kitamura-Stoye_2018_ECMA}, and \citet*{Deb-Kitamura-Quah-Stoye_2018_wp} for a sample of recent contributions and further references. As mentioned above, a key feature of RAM is that our proposed monotonic attention condition on attention rule nests previous models as special cases, and also covers many new models of choice behavior. In particular, RAM can accommodate more choice behaviors/patterns than what can be rationalized by random utility models. This is important because numerous studies in psychology, finance and marketing have shown that decision makers exhibit limited attention when making choices: they only compare (and choose from) a subset of all available options. Whenever decision makers do not pay full attention to all options, implications from revealed preference theory under random utility models no longer hold in general, implying that empirical testing of substantive hypotheses as well as policy recommendations based on random utility models will be invalid. On the other hand, our results may remain valid.
In contemporaneous work, a few scholars have also developed identification and inference results under (random) limited attention, trying to connect behavioral theory and econometric methods, as we do in this paper. Three recent examples of this new research area include \citet*{Abaluck_Adams_2017}, \citet*{Dardanoni_Manzini_Mariotti_Tyson_2018}, and \citet*{Barseghyan-Coughlin-Molinari-Teitelbaum_2018_wp}. These papers are complementary to ours insofar different assumptions on the random attention rule and preference(s) are imposed, which lead to different levels of (partial) identification of preference(s) and (random) attention rule(s). For a further discussion on the relationship with these papers, see Section SA.1 of the Supplemental Appendix.
The rest of the paper proceeds as follows. Section (ref) presents the basic setup, where our key monotonicity assumption on the decision maker's stochastic attention rule is presented in Section (ref). Section (ref) discusses in detail our random attention model, including the main revealed preference results. Section (ref) presents our main econometrics methods, including nonparametric (partial) identification, estimation, and inference results. In Section (ref), we consider additional restrictions on the attention rule for binary choice problems, which can help improve our identification and inference results considerably. We also consider random attention filters in Section (ref), which are one of the motivating examples of monotonic attention rules. In this case, however, there is no additional identification. Section (ref) summarizes the findings from a simulation study. Finally, Section (ref) concludes with a discussion of directions for future research. A companion online Supplemental Appendix includes more examples, extensions and other methodological results, omitted proofs, and additional simulation evidence.
We designate a finite set $X$ to act as the universal set of all mutually exclusive alternatives. This set is thus viewed as the grand alternative space, and is kept fixed throughout. A typical element of $X$ is denoted by $a$ and its cardinality is $|X|=K$. We let $\mathcal{X}$ denote the set of all non-empty subsets of $X$. Each member of $\mathcal{X}$ defines a choice problem.
Thus, $\pi(a |S)$ represents the probability that the decision maker chooses alternative $a$ from the choice problem $S$. Our formulation allows both stochastic and deterministic choice rules. If $\pi(a |S)$ is either $0$ or $1$, then choices are deterministic. For simplicity in the exposition, we assume that all choice problems are potentially observable throughout the main paper, but this assumption is relaxed in Section SA.3 of the Supplemental Appendix to account for cases where only data on a subcollection of choice problems is available.
The key ingredient in our model is probabilistic consideration sets. Given a choice problem $S$, each non-empty subset of $S$ could be a consideration set with certain probability. We impose that each frequency is between $0$ and $1$ and that the total frequency adds up to $1$. Formally,
Thus, $\mu(T|S)$ represents the probability of paying attention to the consideration set $T \subset S $ when the choice problem is $S$. This formulation captures both deterministic and stochastic attention rules. For example, $\mu(S|S)=1$ represents an agent with full attention. Given our approach, we can always extract the probability of paying attention to a specific alternative: For a given $a \in S$, $\sum\limits_{a \in T \subset S} \mu(T|S) $ is the probability of paying attention to $a$ in the choice problem $S$. The probabilities on consideration sets allow us to derive the attention probabilities on alternatives uniquely.
We consider a choice model where a decision maker picks the maximal alternative with respect to her preference among the alternatives she pays attention to. Our ultimate goal is to elicit her preferences from observed choice behavior without requiring any information on consideration sets. Of course, this is impossible without any restrictions on her (possibly random) attention rule. For example, a decision maker's choice can always be rationalized by assuming she only pays attention to singleton sets. Because the consumer never considers two alternatives together, one cannot infer her preferences at all.
We propose a property (i.e., an identifying restriction) on how stochastic consideration sets change as choice problems change, as opposed to explicitly modeling how the choice problem determines the consideration set. We argue below that this nonparametric property is indeed satisfied by many problems of interest and mimics heuristics people use in real life (see examples below and in Section SA.2 of the Supplemental Appendix). This approach makes it possible to apply our method to elicit preference without relying on a particular formation mechanism of consideration sets.
Monotonic $\mu$ captures the idea that each consideration set competes for consumers' attention: the probability of a particular consideration set does not shrink when the number of possible consideration sets decreases. Removing an alternative that does not belong to the consideration set $T$ results in less competition for $T$, hence the probability of $T$ being the consideration set in the new choice problem is weakly higher. Our assumption is similar to the regularity condition proposed by Suppes-Luce_1965_Handbook. The key difference is that their regularity condition is defined on choice probabilities, while our assumption is defined on attention probabilities.
To demonstrate the richness of the framework and motivate the analysis to follow, we discuss six leading examples of families of monotonic attention rules, that is, attention rules satisfying Assumption (ref). We offer several more examples in Section SA.2 of the Supplemental Appendix. The first example is deterministic (i.e., $\mu(T|S) $ is either $0$ or $1$), but the others are all stochastic.
These six examples give a sample of different limited attention models of interest in economics, psychology, marketing, and many other disciplines. While these examples are quite distinct from each other, all of them are monotonic attention rules.\footnote{To provide an example where Assumption (ref) might be violated, consider a generalization of Independent Consideration of Manzini-Mariotti_2014_ECMA. In this generalization, the degree of brand awareness for a product is not only a function of the product but also a function of the context, that is, $\gamma_S(a)$. Then, the frequency of each set being the consideration set is calculated as in Independent Consideration rule. Due to this contextual dependence, further restrictions on $\gamma_{S}(a)$ and $\gamma_{S-b}(a)$ are needed to ensure Assumption (ref).} As a consequence, our revealed preference characterization will be applicable to a wide range of choice rules without committing to a particular attention mechanism, which is not observable in practice and hence untestable. Furthermore, as illustrated by the examples above (and those in Section SA.2 of the Supplemental Appendix), our upcoming characterization, identification, estimation, and inference results nest important previous contributions in the literature.
We are ready to introduce our random attention model based on Assumption (ref). We assume the decision maker has a strict preference ordering $\succ$ on $X$. To be precise, we assume the preference ordering is an asymmetric, transitive and complete binary relation. A binary relation $\succ$ on a set $X$ is (i) asymmetric, if for all $x, y \in X$, $x \succ y$ implies $y \not\succ x$; (ii) transitive, if for all $x, y,z \in X$, $x \succ y$ and $y \succ z$ imply $x \succ z$; and (iii) complete, if for all $x \neq y \in X $, either $x \succ y$ or $y \succ x$ is true. Consequently, the decision maker always picks the maximal alternative with respect to her preference among the alternatives she pays attention to. Formally,
While our framework is designed to model stochastic choices, it captures deterministic choices as well. In classical choice theory, a decision maker chooses the best alternative according to her preferences with probability $1$, hence choice is deterministic. In our framework, this case is captured by a monotone attention rule with $\mu(S|S)=1$. Figure (ref) gives a graphical representation of RAM.
We now derive the implications of our random attention model. They can be used to test the model in terms of observed choice rules/probabilities. In this section, we treat the choice rule as known/observed in order to facilitate the discussion of preference elicitation. In practice, the researcher may only observe a set of choice problems and choices thereof. We discuss econometric implementation in Section (ref): even if the choice rule is not directly observed, it is identified (consistently estimable) from choice data.
In the literature, there is a principle called regularity Suppes-Luce_1965_Handbook, according to which adding a new alternative should only decrease the probability of choosing one of the existing alternatives. However, empirical findings suggest otherwise. \citet*{Rieskamp_Busemeyer_Mellers_2006_JEL} provide a detailed review of empirical evidence on violations of regularity and alternative theories explaining these violations. Importantly, our model allows regularity violations.
The next example illustrates that adding an alternative to the feasible set can increase the likelihood that an existing alternative is selected. This cannot be accommodated in the Luce (multinomial logit) model, nor in any random utility model. In RAM, the addition of an alternative changes the choice set and therefore the decision maker's attention, which could increase the probability of an existing alternative being chosen.
This example shows that RAM can explain choice patterns that cannot be explained by the classical random utility model. Given that the model allows regularity violations, one might think that the model has very limited empirical implications, i.e. that it is too general to have empirical content. However, it is easy to find a choice rule $\pi$ that lies outside RAM with only three alternatives. Here we provide an example where our model makes very sharp predictions.
One might wonder that the model makes a strong prediction due to the cyclical binary choices, i.e., $\pi(a|\{a,b\})=\pi(b|\{b,c\})=\pi(c|\{a,c\})=1$. We can generate a similar prediction where the individual is perfectly rational in the binary choices, i.e., $\pi(a|\{a,b\})=\pi(a|\{a,c\})=\pi(b|\{b,c\})=1$. In this case, our model predicts that the individual cannot chose both $b$ and $c$ with strictly positive probability when the choice problem is $\{a,b,c\}$. Therefore, we obtain similar predictions. Given that RAM has non-trivial empirical content, it is natural to investigate to what extent Assumption (ref) can be used to elicit (unobserved) strict preference orderings given (observed) choices of decision makers.
In general, a choice rule can have multiple RAM representations with different preference orderings and different attention rules. When multiple representations are possible, we say that $a$ is revealed to be preferred to $b$ if and only if $a$ is preferred to $b$ in all possible RAM representations. This is a very conservative approach as it ensures that we never make false claims about the preference of the decision maker.
We now show how revealed preference theory can still be developed successfully in our RAM framework. If all representations share the same preferences $\succ$ (or if there is a unique representation), then the revealed preference will be equal to $\succ$. In general, if one wants to know whether $a$ is revealed to be preferred to $b$, it would appear necessary to identify all possible $(\succ_{j},\mu_{j})$ representations. However, this is not practical, especially when there are many alternatives. Instead, we shall now provide a handy method to obtain the revealed preference completely.
Our theoretical strategy parallels that of \citet*{Masatlioglu-Nakajima-Ozbay_2012_AER} (MNO) in their study of a deterministic model of inattention. MNO identifies $a$ as revealed preferred to $b$ whenever $a$ is chosen in the presence of $b$, and removing $b$ causes a choice reversal. This particular observation, in conjunction with the structure of attention filters, ensures that the decision maker considers $b$ while choosing $a$. Here, we show that $a$ is revealed preferred to $b$ if removing $b$ causes a regularity violation, that is, $\pi(a|S) > \pi(a|S-b)$. To see this, assume $(\succ,\mu)$ represents $\pi$ and $\pi(a|S) > \pi(a|S-b)$. By definition, we have
Hence, we have the following inequality: $$ {{\pi(a|S ) - \pi(a|S -b)}} \leq \sum_{\substack{ {{b \in T \subset S}}, \\ a\text{ is }{{\succ}} \text{-best in }T}}{\mu(T|S)} $$ Since $\pi ( a|S ) - \pi(a|S -b) > 0$, there must exist at least one $T$ such that (i) $b \in T$, (ii) $a\text{ is }{{\succ}} \text{-best in }T$, and (iii) $\mu(T|S) \neq 0$. Therefore, there exists at least one occasion that the decision maker pays attention to $b$ while choosing $a$ (Revealed Preference). The next lemma summarizes this interesting relationship between regularity violations and revealed preferences. It simply illustrates that the existence of a regularity violation informs us about the underlying preference.
Lemma (ref) allows us to define the following binary relation. For any distinct $a$ and $b$, define:
By Lemma (ref), if $a\mathsf{P}b$ then $a$ is revealed to be preferred to $b$. In other words, this condition is sufficient to reveal preference. In addition, since the underlying preference is transitive, we also conclude that she prefers $a$ to $c$ if $a\mathsf{P}b$ and $b\mathsf{P}c$ for some $b$, even when $a\mathsf{P}c$ is not directly revealed from her choices. Therefore, the transitive closure of $\mathsf{P}$, denoted by $\mathsf{P}_\mathtt{R}$, must also be part of her revealed preference. One may wonder whether some revealed preference is overlooked by $\mathsf{P}_\mathtt{R}$. The following theorem, which is our first main result, shows that $\mathsf{P}_\mathtt{R}$ includes all preference information given the observed choice probabilities, under only Assumption (ref).
Theorem (ref) establishes the empirical content of revealed preferences under monotonic attention only. Our resulting revealed preferences could be incomplete: it may only provide coarse welfare judgments in some cases. At one extreme, there is no preference revelation when there is no regularity violation. This is because the decision maker's behavior can be attributed fully to her preference or to her inattention (i.e., never considering anything other than her actual choice). This highlights the fact that our revealed preference definition is conservative, which guarantees no false claims in terms of revealed preference especially when there are alternative explanations for the same choice behavior. The following example illustrates that we might make misleading inferences if we wrongly believe the decision maker uses a particular attention rule.
Example (ref) is an example where a specific cnsideration set formation model leads to wrong conclusions on the revealed preferences. This example highlights the importance of knowledge about the underlying choice procedure when we conduct welfare analysis. In other words, welfare analysis is more delicate a task than it looks. Notice that, in the above example, monotonic attention is satisfied as engines do not change their presentations of first page results when an alternative outside of the first page becomes unavailable. Hence, Theorem (ref) is applicable. Since $\pi(a|\{a,b,c\})> \pi(a|\{a,c\})$, our model correctly identifies her true preference between $a$ and $b$. However, our model is silent about the relative ranking of $c$. Therefore, while our revealed preference is conservative, it does not make misleading claims.
We now illustrate that Theorem (ref) could be very useful to understand the attraction effect phenomena. The attraction effect introduced by \citet*{Huber_1982} was the first evidence against the regularity condition. It refers to an inferior product's ability to increase the attractiveness of another alternative when this inferior product is added to a choice set. In a typical attraction effect experiment, we observe $\pi(a|\{a,b,c\}) > \pi(a|\{a,b\})$. Assume that we have no information about the alternatives other than the frequency of choices. Then, by simply using observed choice, Theorem (ref) informs us that the third product $c$ is indeed an inferior alternative compared to $a$ ($a \succ c$). This is exactly how these alternatives are chosen in these experiments. While alternatives $a$ and $b$ are not comparable, alternative $c$, which is also not comparable to $b$, is dominated by $a$. Theorem (ref) informs us about the nature of products by just observing choice frequencies.
Our revealed preference result includes the one in \citet*{Masatlioglu-Nakajima-Ozbay_2012_AER} for attention filters (i.e., non-random monotonic attention rules). In their model, $a$ is revealed to be preferred to $b$ if there is a choice problem such that $a$ is chosen and $b$ is available, but it is no longer chosen when $b$ is removed from the choice problem. This means we have $1=\pi(a|S) > \pi(a|S- b)=0$. Given Theorem (ref), this reveals that $a$ is better than $b$. On the other hand, generalizing this result to non-deterministic attention rules allows for a broader class of empirical and theoretical settings to be analyzed, hence our revealed preference result (Theorem (ref)) is strictly richer than those obtained in previous work. For example, in a deterministic world with three alternatives, there is no data revealing the entire preference. On the other hand, we illustrate that it is possible to reveal the entire preference in RAM with only three alternatives. This discussion makes clear the connection between deterministic and probabilistic choice in terms of revealed preference.
Example (ref) also illustrates that one can achieve unique identification of preferences by utilizing Assumption (ref) even when observed choices cannot be explained by well known random attention models such as the logit attention model of Brady-Rehbeck_2016_ECMA and the independent attention model of Manzini-Mariotti_2014_ECMA. To see this point, assume that $\max\{1-\lambda_a,\lambda_c\}>0$. One can show that neither Brady-Rehbeck_2016_ECMA nor Manzini-Mariotti_2014_ECMA can explain observed choices in this example. First, notice that since both models satisfy Assumption (ref) and the preference is uniquely revealed as $a\succ b\succ c$ under Assumption (ref), if the observed choice data can be explained by either model, then their revealed preference must also be $a\succ b\succ c$. That is $c$ must be the worst alternative. On the other hand, $c$ is chosen with zero probability in $\{a,b,c\}$. These models then imply that $c$ must also be chosen with zero probability in $\{a,c\}$ and $\{b,c\}$. This contradicts our assumption that $\max\{1-\lambda_a,\lambda_c\}>0$.
Theorem (ref) characterizes the revealed preference in our model. However, it is not applicable unless the observed choice behavior has a random attention representation, which motivates the following question: how can we test whether a choice rule is consistent with RAM? It turns out that RAM can be simply characterized by only one behavioral postulate of choice: acyclicity. Our characterization is based on an idea similar to Houthakker_1950_Economica. Choices reveal information about preferences. If these revelations are consistent in the sense that there is no cyclical preference revelation, the choice behavior has a RAM representation.
Recall that Example (ref) is outside of our model. Theorem (ref) implies that $\mathsf{P}_\mathtt{R}$ must have a cycle. Indeed, we have $a \mathsf{P} b$ due to the regularity violation $\pi(a|\{a,b,c\}) =\lambda_a >0 = \pi(a|\{a,c\})$. Similarly, we have $b \mathsf{P} c$ by $\pi(b|\{a,b,c\}) =\lambda_b >0 = \pi(b|\{a,b\})$) and $c \mathsf{P} a$ by ($\pi(c|\{a,b,c\}) =\lambda_c >0 = \pi(c|\{b,c\})$). Since $\mathsf{P}$ has a cycle, Example (ref) must be outside of our model. Therefore, Theorem (ref) provides a very simple test of RAM.
Our characterization result also helps us to understand the relation between our model and random utility models. It is well-known in the literature that any choice rule that has a random utility model representation satisfies regularity. On the other hand, for any choice rule that satisfies regularity, $\mathsf{P}$ will trivially have no cycle. Hence, any choice rule that has a random utility model representation also has a RAM representation. However, in terms of modeling purposes, RAM assumes random attention with a deterministic preference whereas random utility model assumes random preference and deterministic (full) attention.
Before closing this section, we sketch the proof of Theorem (ref), and provide a corollary which will be used in the next section for developing econometric methods. The “only if” part of Theorem (ref) follows directly from Lemma (ref). For the “if” part, we need to construct a preference and a monotonic attention rule representing the choice rule. Given that $\mathsf{P}$ has no cycle, there exists a preference relation $\succ$ including $\mathsf{P}_\mathtt{R}$. Indeed, we illustrate that any such completion of $\mathsf{P}_\mathtt{R}$ represents $\pi$ by an appropriately chosen $\mu$. The construction of $\mu$ depends on a particular completion of $\mathsf{P}_\mathtt{R}$, and is not unique in general. We then illustrate that the constructed $\mu$ satisfies Assumption (ref). At the last step, we show that $(\succ, \mu)$ represents $\pi$. In Corollary (ref), we provide one specific construction of the attention rule. We first make a definition.
Theorem (ref) shows that if the choice probability $\pi$ is a RAM then preference revelation is possible. Theorem (ref) gives a falsification result, based on which a specification test can be designed. The challenge for econometric implementation, however, is that our main assumption, monotonic attention, is imposed on the attention rule, and that the attention rule is not identified from a typical choice data and has a much higher dimension than the identified (consistently estimable) choice rule. To circumvent this difficulty, we rely on Corollary (ref), which states that if $\pi$ has a random attention representation $(\succ,\mu)$, then there exists a unique monotonic triangular attention rule $\tilde{\mu}$ such that $(\succ,\tilde{\mu})$ is also a representation of $\pi$. This latter result turns out to be useful for our proposed identification, estimation, and inference methods, as it allows us to construct, for each given preference ordering, a mapping from the identified choice rule to a triangular attention rule, for which we can test whether Assumption (ref) holds. This test turns out to be a test on moment inequalities.
We first define the set of partially identified preferences, which mirrors Definition (ref), with the only difference that now we fix the choice rule to be identified/estimated from data. More precisely, let $\pi$ be the underlying choice rule/data generating process. Then a preference $\succ$ is compatible with $\pi$, denoted by $\succ\ \in\Theta_{\pi}$,\footnote{$\Theta_{\pi}$ is not the same as $\mathsf{P}_{\mathtt{R}}$ (defined in Section (ref)): $\mathsf{P}_{\mathtt{R}}$ contains all revealed preferences, while $\Theta_{\pi}$ is the set of preferences compatible with the choice probability (i.e., all possible completions of $\mathsf{P}_{\mathtt{R}}$). For example, when there is no preference revelation, $\Theta_{\pi}$ contains all preference orderings, and $\mathsf{P}_{\mathtt{R}}$ will be empty. For the other extreme that the choice probability is not compatible with our RAM, $\Theta_{\pi}$ will be empty and $\mathsf{P}_{\mathtt{R}}$ will involve cycles.} if there exists some monotonic attention rule $\mu$ such that $(\pi,\succ,\mu)$ is a RAM.
When $\pi$ is known, it is possible to employ Theorem (ref) directly to construct $\Theta_{\pi}$. For example, consider the specific preference ordering $a\succ b$, which can be checked by the following procedure. First, check whether $\pi(b|S)\leq \pi(b|S-a)$ is violated for some $S$. If so, then we know the preference ordering is not compatible with RAM and hence does not belong to $\Theta_{\pi}$ (Lemma (ref)). On the other hand, if the preference ordering is not rejected in the first step, we need to check along “longer chains” (Theorem (ref)). That is, whether $\pi(b|S)\leq \pi(b|S-c)$ and $\pi(c|T)\leq \pi(c|T-a)$ are simultaneously violated for some $S$, $T$ and $c$. If so, the preference ordering is rejected (i.e., incompatible with RAM), while if not then a chain of length three needs to be considered. This process goes on for longer chains until either at some step we are able to reject the preference ordering, or all possibilities are exhausted. In practice, additional comparisons are needed since it is rarely the case that only a specific pair of alternatives is of interest. This algorithm, albeit feasible, can be hard to implement in practice, even when the choice probabilities are known. The fact that $\pi$ has to be estimated makes the problem even more complicated, since it becomes a sequential multiple hypothesis testing problem.
Another possibility is to employ the J-test approach, which stems from the idea that, given the choice rule, compatibility of a preference is equivalent to the existence of an attention rule satisfying monotonicity. To implement the J-test, one fixes the choice rule (identified/estimated from the data) and the preference ordering (the null hypothesis to be tested), and search the space of all monotonic attention rules and check if Definition (ref) applies. The J-test procedure can be quite computationally demanding, due to the fact that the space of attention rules has high dimension. We further discuss the J-test approach in Section SA.4.3 of the Supplemental Appendix and how it is related to our proposed procedure.
One of the main purposes of this section is to provide an equivalent form of identification, which (i) is simple to implement, and (ii) remains statistically valid even when applied using estimated choice rules. For ease of exposition, we rewrite the choice rule $\pi$ as a long vector $\boldsymbol{\pi}$, whose elements are simply the probability of each alternative $a\in X$ being chosen from a choice problem $S\in\mathcal{X}$. For example, one can label the choice problems as $S_1$, $S_2$, $\cdots$, the alternatives as $a_1$, $a_2$, $\cdots$, $a_K$, and then the vector $\boldsymbol{\pi}$ simply consists of $\pi(a_1|S_1)$, $\pi(a_2|S_1)$, $\cdots$, $\pi(a_K|S_1)$, $\pi(a_1|S_2)$, $\pi(a_2|S_2)$, etc. See Example (ref) for a concrete illustration.
This theorem states that in order to decide whether a preference $\succ$ is compatible with the (identifiable) choice rule $\pi$, it suffices to check a collection of inequality constraints. In particular, it is no longer necessary to consider the sequential and multiple testing problems mentioned earlier, or numerically searching in the high dimensional space of attention rules. Moreover, as we discuss below, given the large econometric literature on moment inequality testing, many techniques can be adapted when Theorem (ref) is applied to estimated choice rules. An algorithmic construction of the constraint matrix $\mathbf{R}_{\succ}$ is given in Algorithm (ref).
As can be seen, the only input needed is the preference $\succ$, which we are interested in testing against. Each row of $\mathbf{R}_{\succ}$ consists of one “$+1$”, one “$-1$”, and $0$ otherwise. The constraint matrix $\mathbf{R}_{\succ}$ is non-random and does not depend on the estimated choice probabilities, but rather determined by the collection of (fixed, known to the researcher) restrictions on the estimable choice probabilities. Next we compute the number of constraints (i.e. rows) in $\mathbf{R}_\succ$ for the complete data case (i.e., when all choice problems are observed): \[ \text{\#row}(\mathbf{R}_\succ) = \sum_{S\in\mathcal{X}} \sum_{a,b\in S}\mathds{1}(b\prec a) = \sum_{S\in\mathcal{X},\ |S|\geq 2} \binom{|S|}{2} = \sum_{k=2}^{K} \binom{K}{k}\binom{k}{2}, \] where $K = |X|$ is the number of alternatives in the grand set $X$. Not surprisingly, the number of constraints increases very fast with the size of the grand set. However, once the matrix $\mathbf{R}_{\succ}$ has been constructed for one preference $\succ$, the constraint matrices for other preference orderings can be obtained by column permutations of $\mathbf{R}_{\succ}$. This is particular useful and saves computation if there are multiple hypotheses to be tested, as the above algorithm only needs to be implemented once.
Finally, we illustrate that, in simple examples, the constraint matrix $\mathbf{R}_{\succ}$ can be constructed intuitively.
Given the identification result in Theorem (ref), we can replace the identifiable choice rule with its estimate to conduct estimation and inference of the (partially identifiable) preferences. We can also conduct specification testing by evaluating whether the identified set $\Theta_{\pi}$ is empty. To proceed, we assume the following data structure.
We only assume the data is generated from some choice rule $\pi$. We allow for the possibility that it is not a RAM, since our identification result permits falsifying the RAM representation: $\pi$ has a RAM representation if and only if $\Theta_{\pi}$ is not empty according to Theorem (ref). In addition, we only assume that the choice problem $Y_i$ and the corresponding selection $y_i\in Y_i$ are observed for each unit, while the underlying (possibly random) consideration set for the decision maker remains unobserved (i.e., the set $T$ in Definition (ref) and Figure (ref)). For simplicity, we discuss the case of “complete data” where all choice problems are potentially observable, but in Section SA.3 and SA.4.4 of the Supplemental Appendix we extend our work to the case of incomplete data.
The estimated choice rule is denoted by $\hat{\pi}$, \[\hat{\pi}(a|S) = \frac{ \sum_{1\leq i\leq N} \mathds{1}(y_i=a,\ Y_i=S) }{ \sum_{1\leq i\leq N} \mathds{1}(Y_i=S) }, \qquad a\in S, \quad S\in \mathcal{X}. \] For convenience, we represent $\hat{\pi}(\cdot|S)$ by the vector $\hat{\boldsymbol{\pi}}_S$, and its population counterpart by $\boldsymbol{\pi}_{S}$. The choice rules are stacked into a long vector, denoted by $\hat{\boldsymbol{\pi}}$ with the population counterpart $\boldsymbol{\pi}$.
We consider Studentized test statistics, and hence we introduce some additional notation. Let $\boldsymbol{\sigma}_{\pi,\succ}$ be the standard deviation of $\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}}$, and $\hat{\boldsymbol{\sigma}}_{\succ}$ be its plug-in estimate. That is, \[ \boldsymbol{\sigma}_{\pi,\succ} = \sqrt{\mathrm{diag}\Big( \mathbf{R}_{\succ} \boldsymbol{\Omega}_{\pi} \mathbf{R}_{\succ}^\prime \Big)} \qquad\text{and}\qquad \hat{\boldsymbol{\sigma}}_{\succ} = \sqrt{\mathrm{diag}\Big( \mathbf{R}_{\succ} \hat{\boldsymbol{\Omega}}\mathbf{R}_{\succ}^\prime \Big)}, \] where $\mathrm{diag}(\cdot)$ denotes the operator that extracts the diagonal elements of a square matrix, or constructs a diagonal matrix when applied to a vector. Here $\boldsymbol{\Omega}_{\pi}$ is block diagonal, with blocks given by $\frac{1}{\mathbb{P}[Y_i=S]}\boldsymbol{\Omega}_{\pi,S}$, and $\boldsymbol{\Omega}_{\pi,S}=\mathrm{diag}(\boldsymbol{\pi}_{S}) - \boldsymbol{\pi}_{S}\boldsymbol{\pi}_{S}^\prime$. The estimator $\hat{\boldsymbol{\Omega}}$ is simply constructed by plugging in the estimated choice rule.
Consider the null hypothesis $\mathsf{H}_0:\,\succ\ \in\Theta_{\pi}$. This null hypothesis is useful if the researcher believes a certain preference represents the underlying data generating process. It also serves as the basis for constructing confidence sets or for ranking preferences according to their (im)plausibility in repeated sampling (for example, via employing associated p-values). Given a specific preference, the test statistic is constructed as the maximum of the Studentized, restricted sample choice probabilities: \[ \mathscr{T}(\succ) = \sqrt{N}\cdot \max\Big\{(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}}) \oslash \hat{\boldsymbol{\sigma}}_\succ,\ 0\Big\}, \] where $\oslash$ denotes elementwise division (i.e, Hadamard division) for conformable matrices. The test statistic is the largest element of the vector $\sqrt{N}(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}}) \oslash \hat{\boldsymbol{\sigma}}_\succ$ if it is positive, or zero otherwise. The reasoning behind such construction is straightforward: if the preference is compatible with the underlying choice rule, then in the population we have $\mathbf{R}_{\succ}\boldsymbol{\pi}\leq \mathbf{0}$, meaning that the test statistic, $\mathscr{T}(\succ)$, should not be “too large.”
Other test statistics have been proposed for testing moment inequalities, and usually the specific choice depends on the context. When many moment inequalities can be potentially violated simultaneously, it is usually preferred to use a statistic based on truncated Euclidean norm. In our problem, however, we expect only a few moment inequalities to be violated, and therefore we prefer to employ $\mathscr{T}(\succ)$. Having said this, the large sample approximation results given in Theorem (ref) can be adapted to handle other test statistics commonly encountered in the literature on moment inequalities.
The null hypothesis is rejected whenever the test statistic is “too large,” or more precisely, when it exceeds a critical value, which is chosen to guarantee uniform size control in large samples. We describe how this critical value leading to uniformly valid testing procedures is constructed based on simulating from multivariate normal distributions. Our construction employs the Generalized Moment Selection (GMS) approach of Andrews-Soares_2010_ECMA; see also Canay_2010_JoE and Bugni_2016_ET for closely related methods. The literature on moment inequalities testing includes several alternative approaches, some of which we discuss briefly in Section SA.4.5 of the Supplemental Appendix.
To illustrate the intuition behind the construction, first rewrite the test statistic $\mathscr{T}(\succ)$ as the following: \[ \mathscr{T}(\succ) = \max\Big\{\big(\mathbf{R}_{\succ}\sqrt{N}(\hat{\boldsymbol{\pi}}-\boldsymbol{\pi}) + \sqrt{N}\mathbf{R}_{\succ}\boldsymbol{\pi}\big) \oslash \hat{\boldsymbol{\sigma}}_\succ,\ 0\Big\}. \] By the central limit theorem, the first component $\sqrt{N}(\hat{\boldsymbol{\pi}}-\boldsymbol{\pi})$ is approximately distributed as $\mathcal{N}(\mathbf{0},\ \boldsymbol{\Omega}_{\pi})$. The second component, $\mathbf{R}_{\succ}\boldsymbol{\pi}$, although unknown, is bounded above by zero under the null hypothesis. Motivated by these observations, we approximate the distribution of $\mathscr{T}(\succ)$ by simulation as follows: \[\mathscr{T}^\star(\succ) = \sqrt{N}\cdot \max\Big\{ (\mathbf{R}_{\succ}\mathbf{z}^\star) \oslash \hat{\boldsymbol{\sigma}}_\succ + \psi_N(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}},\hat{\boldsymbol{\sigma}}_\succ),\ 0\Big\}. \] Here $\mathbf{z}^\star$ is a random vector simulated from the distribution $\mathcal{N}(\mathbf{0}, \hat{\boldsymbol{\Omega}}/N)$, and $\sqrt{N}\psi_N(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}},\hat{\boldsymbol{\sigma}}_\succ)$ is used to replace the unknown moment conditions $(\sqrt{N}\mathbf{R}_{\succ}\boldsymbol{\pi}) \oslash \hat{\boldsymbol{\sigma}}_\succ$. Several choices of $\psi_N$ have been proposed. One extreme choice is $\psi_N(\cdot)=0$, so that the upper bound $0$ is used to replace the unknown $\mathbf{R}_\succ\boldsymbol{\pi}$. Such a choice also delivers uniformly valid inference in large samples, and is usually referred to as “critical value based on the least favorable model.” However, for practical purposes it is better to be less conservative. In our implementation we employ \[ \psi_N(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}},\hat{\boldsymbol{\sigma}}_\succ) = \frac{1}{\kappa_N}\Big(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}} \oslash \hat{\boldsymbol{\sigma}}_\succ\Big)_- , \] where $(\mathbf{a})_-=\mathbf{a}\odot\mathds{1}(\mathbf{a}\leq 0)$, with $\odot$ denoting the Hadamard product, the indicator function $\mathds{1}(\cdot)$ operating element-wise on the vector $\mathbf{a}$, and $\kappa_N$ diverges slowly. That is, the function $\psi_N(\cdot)$ retains the non-positive elements of $(\mathbf{R}_{\succ}\hat{\boldsymbol{\pi}} \oslash \hat{\boldsymbol{\sigma}}_\succ)/\kappa_N$, since under the null hypothesis all moment conditions are non-positive. We use $\kappa_N=\sqrt{\ln N}$, which turns out to work well in the simulations described in Section (ref). For other choices of $\psi_N(\cdot)$, see Andrews-Soares_2010_ECMA.
In practice, $M$ simulations are conducted to obtain the simulated statistics $\{\mathscr{T}_m^\star(\succ) :\ 1\leq m\leq M \}$. Then, given some $\alpha\in(0,1)$, the critical value is constructed as \[ c_{\alpha}(\succ) = \inf\Big\{ t:\ \frac{1}{M}\sum_{m=1}^M \mathds{1}\big(\mathscr{T}^\star_m(\succ) \leq t \big)\geq 1-\alpha \Big\}, \] and the null hypothesis $\mathsf{H}_0:\ \succ\,\in\Theta_{\pi}$ is rejected if and only if $\mathscr{T}(\succ)>c_{\alpha}(\succ)$. Alternatively, one can compute the p-value as \[ \text{pVal}(\succ) = \frac{1}{M}\sum_{m=1}^M \mathds{1}\Big(\mathscr{T}^\star_m(\succ) > \mathscr{T}(\succ)\Big). \]
To justify the proposed critical values, it is important to address uniformity issues. A testing procedure is (asymptotically) uniform among a class of data generating processes, if the asymptotic size does not exceed the nominal level across this class. Testing procedures that are valid only pointwise but not uniformly may yield bad approximations to the finite sample distribution, because in finite samples the moment inequalities could be close to binding. The following theorem shows that conducting inference using the critical values above is uniformly valid.
The proof is postponed to Appendix (ref). The only requirement is that each moment condition is nondegenerate so that the normalized statistics are well-defined in large samples, but no restrictions on correlations among moment conditions are imposed.
We discuss some extensions based on Theorem (ref), including how to construct uniformly valid confidence sets via test inversion, and how to conduct uniformly valid specification testing, both based on testing individual preferences.
Given the uniformly valid hypothesis testing procedure already developed in Theorem (ref), we can obtain a uniformly valid confidence set for the (partially) identified preferences by test inversion: \[ \mathscr{C}(\alpha) = \Big\{\succ \ :\ \mathscr{T}(\succ) \leq c_\alpha(\succ) \Big\}. \] The resulting confidence set $\mathscr{C}(\alpha)$ exhibits an asymptotic uniform coverage rate of at least $1-\alpha$: \[ \liminf_{N\to \infty}\inf_{\pi\in\Pi}\min_{\succ\,\in\Theta_{\pi}}\mathbb{P}\big[ \succ\,\in \mathscr{C}(\alpha) \big] \geq 1-\alpha. \] This inference method offers a uniformly valid confidence set for each member of the partially identified set with pre-specified coverage probability, which is a popular approach in the partial identification literature Imbens-Manski_2004_ECMA.
Given a collection of preferences, an empirically relevant question is whether any of them is compatible with the data generating process---a basic model specification question. That is, the question is whether the null hypothesis $\mathsf{H}_0:\mathcal{P}\cap\Theta_{\pi}\neq\emptyset$ should be rejected. If the null hypothesis is rejected, then certain features shared by the collection of preferences is incompatible with the underlying decision theory (up to Type I error). See \citet*{Bugni-Canay-Shi_2015_JoE}, \citet*{Kaido-Molinari-Stoye_2019_ECMA} and references therein for further discussion of this idea and related methods.
For a concrete example, consider the question that whether $a\succ b$ is compatible with the data generating process. As long as there are more than 2 alternatives in the grand set, a question like this can be accommodated by setting $\mathcal{P} = \{ \succ:\ a\succ b \}$. Rejection of this null hypothesis provides evidence in favor of $b$ being preferred to $a$, (up to Type I error). Of course with more preferences included in the collection, it becomes more difficult to reject the null hypothesis.
The test is based on whether the confidence set intersects with $\mathcal{P}$: \[ \text{$\mathsf{H}_0$ is rejected} \quad \text{if and only if} \quad \mathscr{C}(\alpha)\cap \mathcal{P}=\emptyset. \] We note that, since $\mathscr{C}(\alpha)$ covers elements in the identified set asymptotically and uniformly with probability $1-\alpha$, the above testing procedure will have uniform size control. Indeed, if $\mathcal{P}\cap \Theta_{\pi}\neq \emptyset$, there exists some $\succ\,\in \mathcal{P}\cap \Theta_{\pi}$, which will be included in $\mathscr{C}(\alpha)$ with at least $1-\alpha$ probability asymptotically.
One important application of this idea is to set $\mathcal{P}$ as the collection of all possible preferences, which leads to a specification testing. Then, the null hypothesis becomes $\mathsf{H}_0:\Theta_{\pi}\neq\emptyset$, and is rejected based on the following rule: \[ \text{$\mathsf{H}_0$ is rejected} \quad \text{if and only if} \quad \mathscr{C}(\alpha)=\emptyset. \] Rejection in this case implies that at least one of the underlying assumptions is violated, and the data generating process cannot be represented by a RAM (up to Type I error).
Our identification and inference results so far are obtained using RAM only, that is, all empirical content of our revealed preference theory comes from the weak nonparametric Assumption (ref). As mentioned before, our model provides a minimum benchmark for preference revelation, which sometimes may not deliver enough empirical content. However, it is easy to incorporate additional (nonparametric) assumptions in specific settings. In this section, we first illustrate one such possibility, where additional restrictions on the attentional rule are imposed for binary choice problems. This will improve our identification and inference results considerably. We then consider random attention filters, which are one of the motivating examples of monotonic attention rules, and show that in this case there is no identification improvement relative to the baseline RAM.
To motivate our approach, a policy maker may want to conclude that $a$ is revealed to be preferred to $b$ if the decision maker chooses $a$ over $b$ “frequently enough” in binary choice problems. “Frequently enough” is measured by a constant $\phi \geq 1/2$.\footnote{Even when the policy maker is least cautious, we need $\pi(a|\{a,b\}) > \pi(b|\{a,b\})$ to conclude $a$ is strictly better than $b$. This implies $\pi(a|\{a,b\})>1/2$. Hence $\phi$ must be greater than $1/2$.} For example, when $\phi=2/3$, it means that choosing $a$ twice more often than choosing $b$ implies $a$ is better than $b$. $\phi$ represents how cautious the policy maker is. Denote by \[a \mathsf{P}^{\phi} b\quad \text{if and only if}\quad \pi(a|\{a,b\})>\phi.\] To justify $\mathsf{P}^{\phi}$ as preference revelation, the policy maker inherently assumes that the decision maker pays attention to the entire set “frequently enough.” This is captured by the following assumption on the attention rule.
The quantity $\frac{1-\phi}{\phi}$ is a measure of full attention at binaries. When $\frac{1-\phi}{\phi}=0$ (or $\phi=1$), there is no constraint on $\mu(\{a,b\}|\{a,b\})$. In this case, it is possible that the decision maker only considers singleton consideration sets. When $\frac{1-\phi}{\phi}$ gets larger (or $\phi$ gets smaller), the probability of being fully attentive is strictly positive, which creates room for preference revelation. An alternative way to understand Assumption (ref) is as follows. Take $\phi = \max\{ \pi(a|\{a,b\}),\pi(b|\{a,b\}) \}$, then $\frac{1-\phi}{\phi} \max \{\mu(\{a\}|\{a,b\})\ ,\ \mu(\{b\}|\{a,b\})\}$ is a strict lower bound on the amount of attention that the decision maker has to pay to both options, for revelation to occur.
We now illustrate that, under Assumption (ref), if $\pi(a|\{a,b\}) > \phi$ then $a$ is revealed to be preferred to $b$. Let $(\succ,\mu)$ be a RAM representation of $\pi$ where $\mu$ satisfies Assumption (ref). First, Assumption (ref) necessitates that $\mu(\{a\}|\{a,b\}) $ cannot be higher than $ \phi$. (To see this, assume $\mu(\{a\}|\{a,b\}) > \phi$. By Assumption (ref), we must have $\mu(\{a,b\}|\{a,b\}) > 1- \phi$, which is a contradiction.) Then, $\pi(a|\{a,b\}) > \phi $ indicates that $a$ is chosen over $b$ whenever the decision maker pays attention to $\{a,b\}$ (revealed preference). Therefore, $a \succ b$.
To accommodate the revealed preference defined in the original model (i.e., to combine Assumption (ref) and (ref)), we now define the following binary relation:
$\mathsf{P}^{\phi}\cup \mathsf{P}$ includes our original binary relation $\mathsf{P}$, defined under the monotonic attention restriction (Assumption (ref)), as well as $\mathsf{P}^{\phi}$, characterized by the new attentive at binary assumption. Therefore, we can infer more.
The next theorem shows that acyclicity of $\mathsf{P}^{\phi}\cup \mathsf{P}$, or its transitive closure $(\mathsf{P}^{\phi}\cup \mathsf{P})_{\mathtt{R}}$, provides a simple characterization of the model we consider in this subsection.
For $ \phi<1$, the model characterized by Theorem (ref) has a higher predictive power (i.e., empirical content) compared to the model characterized by Theorem (ref). Hence the model will fail to retain some of its explanatory power. For example, Example (ref) with $\lambda_a, \lambda_b, \lambda_c < 1-\phi$ is outside of the model given here.
Under the assumption $\phi=1/2$ and $\pi(a|\{a,b\}) \neq 1/2$ for all $a, b$, Theorem (ref) yields that our framework reveals a unique preference while it allows regularity violation.
Now we discuss the econometric implementation. Recall from Section (ref) that, to test if a specific preference ordering is compatible with the observed (identifiable) choice rule and the monotonicity assumption, we first construct a triangular attention rule and then test whether the triangular attention rule satisfies Assumption (ref). This is formally justified in the proof of Theorem (ref).
This line of reasoning can be naturally extended to accommodate Assumption (ref) in our econometric implementation. Again, the researcher constructs a triangular attention rule based on a specific preference ordering and the identifiable choice rule. She then tests whether the triangular attention rule satisfies Assumption (ref) and (ref). This is formally justified in the proof of Theorem (ref). For testing, only minor changes have to be made when constructing the matrix $\mathbf{R}_\succ$. The precise construction is given in Algorithm (ref).
We now revisit Example (ref) to illustrate what additional (identifying) restrictions are imposed by Assumption (ref).
Assumption (ref) improves considerably the empirical content of our benchmark RAM (Assumption (ref)). However, this assumption is just one of many possible assumptions that could be used in addition to our general RAM. The main takeaway is that our proposed RAM offers a baseline for specific, empirically relevant models of choice under random limited attention. In Section (ref) we compare using simulations the empirical content of our benchmark RAM, which employs only Assumption (ref), and the model that incorporates Assumption (ref) as well.
We now consider random attention filters, which are one of the motivating examples of monotonic attention rules. Recall from Section (ref) that an attention filter is a deterministic attention rule that satisfies Assumption (ref), and a random attention filter is a convex combination of attention filters, and hence a random attention filter will also satisfy Assumption (ref). For example, the same individual might be utilizing different platforms during her Internet search. Each platform yields a different attention filter, and the usage frequency of each platform is equal to the weight of that attention filter. Random attention filters also give a different interpretation of our model.
The set of all random attention filters is a strict subset of monotonic attention rules. This is not surprising given that the class of monotonic attention rules is very large. What is (arguably) surprising is the following fact that we are able to show: if $(\pi,\succ,\mu)$ is a RAM with $\mu$ being a monotonic attention rule, there exists a random attention filter $\mu'$ such that $(\pi,\succ,\mu')$ is still a RAM (see Remark (ref)). Before presenting this result, however, we observe that $\mu$ and $\mu'$ need not be the same, which means that there are monotonic attention rules that cannot be written as a convex combination of attention filters.
We now show that if we restrict our attention to a certain type of monotonic attention rules, then we can show that within that class every attention rule is a random attention filter (i.e., convex combination of deterministic attention filters). Let $\mathcal{MT}({\succ})$ denote the set of all attention rules that are both monotonic (Assumption (ref)) and triangular with respect to $\succ$ (Definition (ref) in the Appendix), and let $\mathcal{AF}({\succ})$ denote all attention filters that are triangular with respect to $\succ$. We are now ready to state the main result of this section.
The proof of Theorem (ref) is long and hence left to Appendix, but here we provide a sketch of it. First, $\mathcal{MT}({\succ})$ is a compact and convex set, and thus the above theorem can alternatively be stated as follows: The set of extreme points of $\mathcal{MT}({\succ})$ is $\mathcal{AF}(\succ)$. (An attention rule $\mu\in \mathcal{MT}({\succ})$ is an extreme point of $\mathcal{MT}({\succ})$ if it cannot be written as a nondegenerate convex combination of any $\mu',\mu''\in \mathcal{MT}({\succ})$.) Then, Minkowski's Theorem guarantees that every element of $\mathcal{MT}({\succ})$ lies in the convex hull of $\mathcal{AF}(\succ)$.
Obviously, every element of $\mathcal{AF}(\succ)$ is an extreme point of $\mathcal{MT}({\succ})$. We then show that non-deterministic triangular attention rules cannot be extreme points, i.e. given any $\mu \in \mathcal{MT}({\succ}) - \mathcal{AF}(\succ)$ we can construct $\mu',\mu''\in \mathcal{MT}({\succ})$ such that $\mu=\frac{1}{2}\mu'+\frac{1}{2}\mu''$. The key step is to show that both $\mu'$ and $\mu''$ that we construct are monotonic. After this step, we have shown that no $\mu\in \mathcal{MT}({\succ}) - \mathcal{AF}(\succ)$ can be an extreme point, thus concluding the proof.
This section gives a summary of a simulation study conducted to assess the finite sample properties of our proposed econometric methods. We consider a class of logit attention rules indexed by $\varsigma$:
where $|T|$ is the cardinality of $T$. Thus the decision maker pays more attention to larger sets if $\varsigma>0$, and pays more attention to smaller sets if $\varsigma<0$. When $\varsigma$ is very small (negative and large in absolute magnitude), the decision maker almost always pays attention to singleton sets, hence nothing will be learned about the underlying preference from the choice data.
Other details on the data generating process used in the simulation study are as follows. First, the grand set $X$ consists of five alternatives, $a_1$, $a_2$, $a_3$, $a_4$ and $a_5$. Without loss of generality, assume the underlying preference is $a_1\succ a_2\succ a_3\succ a_4\succ a_5$. Second, the data consists of choice problems of size two, three, four and five. That is, there are in total $26$ choice problems. Third, given a specific realization of $Y_i$, a consideration set is generated from the logit attention model with $\varsigma=2$, after which the choice $y_i$ is determined by the aforementioned preference. We also report simulation evidence for $\varsigma\in\{0,1\}$ in the Supplemental Appendix. Finally, the observed data is a random sample $\{(y_i,Y_i):\ 1\leq i\leq N\}$, where the effective sample size can be $50$, $100$, $200$, $300$ and $400$. (Effective sample size refers to the number of observations for each choice problem. Because there are $26$ choice problems, the overall sample size is $N\in\{1300, 2600, 5200, 7800, 10400\}$.)
For inference, we employ the procedure introduced in Section (ref) and test whether a specific preference ordering is compatible with the basic RAM (Assumption (ref)). We also incorporate the attentive at binaries assumption introduced in Section (ref). Recall from Assumption (ref) that $(1-\phi)/\phi$ is a measure of full attention at binaries, and specifying a larger value (i.e., a smaller value of $\phi$) implies that the researcher is more willing to draw information from binary comparisons. Note that with $\phi=1$, imposing Assumption (ref) does not bring any additional identification power. Before proceeding, we list five hypotheses (preference orderings), and whether they are compatible with our RAM and specific values of $\phi$.
As can be seen, $\mathsf{H}_{0,1}$ always belongs to the identified set of preferences, as it is the preference ordering used in the underlying data generating process. $\mathsf{H}_{0,2}$, however, may or may not belong to the identified set depending on the value of $\phi$: with $\phi$ close to 0.5, the researcher is confident enough using information from binary comparisons, and she will be able to reject this hypothesis; for $\phi$ close to 1, Assumption (ref) no longer brings too much additional identification power beyond the monotonic attention assumption, and monotonic attention alone is not strong enough to reject this hypothesis. Indeed, with $\phi=1$ (i.e., Assumption (ref) alone), the set of identified preference is $\{ \succ: a_2\succ a_3\succ a_4\succ a_5\}$, which contains $\mathsf{H}_{0,2}$. The other three hypotheses, $\mathsf{H}_{0,3}$, $\mathsf{H}_{0,4}$ and $\mathsf{H}_{0,5}$, do not belong to the identified set even with $\phi=1$.
Overall, our simulation has 5 (different $N$) $\times$ 5 (different preference orderings) $\times$ 11 (different $\phi$) $=275$ designs. For each design, 5,000 simulation repetitions are used, and the five null hypotheses are tested using our proposed method at the 5% nominal level. Simulation results are summarized in Figure (ref).
We first focus on $\mathsf{H}_{0,1}$ (panel a). As this preference ordering is compatible with our RAM, one should expect the rejection probability to be less than the nominal level. Indeed, the rejection probability is far below 0.05: this illustrates a generic feature of any (reasonable) procedure for testing moment inequalities---to maintain uniform asymptotic size control, empirical rejection probability is below the nominal level when the inequalities are far from binding. Next consider $\mathsf{H}_{0,2}$ (panel b). For $\phi$ lager than 0.85, the rejection probability is below the nominal size, which is compatible with our theory, because this preference belongs to the identified set when only Assumption (ref) is imposed. With smaller $\phi$, the researcher relies more heavily on information from binary comparisons/choice problems, and she is able to reject this hypothesis much more frequently. This demonstrates how additional restrictions on the attention rule can be easily accommodated by our basic RAM, which in turn can bring additional identification power. The other three hypotheses (panel c, d and e) are not compatible with our RAM, and we do see that the rejection probability is much larger than the nominal size even for $\phi=1$, showing that even our basic RAM has non-trivial empirical content in this case.
We introduced a limited attention model allowing for a general class of monotonic (and possibly stochastic) attention rules, which we called a Random Attention Model (RAM). We showed that this model nests several important recent contributions in both economic theory and econometrics, in addition to other classical results from decision theory. Using our RAM, we obtained a testable theory of revealed preferences and developed partial identification results for the decision maker's unobserved strict preference ordering. Our results included a precise constructive characterization of the identified set for preferences, as well as uniformly valid inference methods based on that characterization. Furthermore, we showed how additional nonparametric restriction can be easily incorporated into RAM to obtain tigher empirical implications, and more powerful accompying econometric procedures. We found good finite sample performance of our econometric methods in a simulation experiment. Last but not least, we provide the general-purpose R software package ramchoice, which allows other researchers to easily employ our econometric methods in empirical applications.