EconBase
← Back to paper

Discrete Choice under Risk with Limited Consideration

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,445 characters · 27 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Discrete Choice under Risk with Limited Consideration

titlepage\begin{abstract} This paper is concerned with learning decision makers' preferences using data on observed choices from a finite set of risky alternatives. We propose a discrete choice model with unobserved heterogeneity in consideration sets and in standard risk aversion. We obtain sufficient conditions for the model's semi-nonparametric point identification, including in cases where consideration depends on preferences and on some of the exogenous variables. Our method yields an estimator that is easy to compute and is applicable in markets with large choice sets. We illustrate its properties using a dataset on property insurance purchases.\\ \noindentKeywords: discrete choice, limited consideration, semi-nonparametric identification \\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\setcounter{page}{1} \setcounter{section}{0}

Introduction

This paper is concerned with learning decision makers' (DMs) preferences using data on observed choices from a finite set of risky alternatives with monetary outcomes. The prevailing empirical approach to study this problem merges expected utility theory (EUT) models with econometric methods for discrete choice analysis. Standard EUT assumes that the DM evaluates all available alternatives and chooses the one yielding the highest expected utility. The DM's risk aversion is determined by the concavity of her Bernoulli utility function. The set of all alternatives -- the choice set -- is assumed to be observable by the researcher.

We depart from this standard approach by proposing a discrete choice model with unobserved heterogeneity in preferences and unobserved heterogeneity in consideration sets. Specifically, preferences satisfy the classic Single Crossing Property (SCP) of mirrlees1971exploration and spence1974market, central to important studies of decision making under risk.\footnote{E.g., apesteguia2017single,chiappori2019aggregate. While our focus is on decision making under risk, the SCP property is satisfied in many contexts, ranging from single agent models with goods that can be unambiguously ordered based on quality, to multiple agents models (e.g., athey2001single).} That is, the preference order of any two alternatives switches only at one value of the preference parameter.\footnote{The EUT framework satisfies the SCP, which requires that if a DM with a certain degree of risk aversion prefers a safer lottery to a riskier one, then all DMs with higher risk aversion also prefer the safer lottery.} Given her unobserved preference parameter, each DM evaluates only the alternatives in her unobserved consideration set, which is a subset of the choice set.

Our first contribution is to provide a general framework for point identification of these models. Our analysis relies on two types of observed data variation. In the first case, we assume that the data include a single (common) excluded regressor affecting the utility of each alternative. In the second case, we assume that each alternative has its own excluded regressor. In both cases, the excluded regressor(s) is independent of unobserved preference heterogeneity. When the excluded regressor(s) also has large support it becomes a “special regressor” Lewbel00, lewbel2012overview. For reasons we explain, the case of the single common excluded regressor is the most demanding from an identification standpoint. Nonetheless, under classic conditions for identification of full-consideration discrete choice models Lewbel00,matzkin07 and the SCP, we obtain semi-nonparametric identification of the preference distribution given basically any consideration set formation mechanism (henceforth, consideration mechanism).\footnote{The identification results are semi-nonparametric because we specify the utility function up to a DM-specific preference parameter. We establish nonparametric identification of the distribution of the latter.} We also prove identification of the consideration mechanism for the widely used Alternative-specific Random Consideration (ARC) model of manski1977structure and man:mar14. The identification argument is constructive and applicable beyond the ARC model. We establish identification results for preferences that do not require large support of the excluded regressor(s). We also show that identification of both preferences and the consideration mechanism is attainable when consideration depends on preferences. In particular, we introduce (i) binary consideration types, and (ii) proportionally shifting consideration, both of which can capture the notion that the DM's attention probabilistically shifts from riskier to safer alternatives as her risk aversion increases. In these cases, identification requires that the distribution of the preference parameter admits a continuous density function.

We can significantly expand our results with alternative-specific excluded regressors. First, we can allow for essentially unrestricted dependence of consideration on preferences without assuming that the excluded regressors have large support. Second, we show that consideration can depend both on preferences and on some excluded regressors. We show this for two cases. In the first case, there is one alternative (the default) that is always considered. The probability of considering other alternatives can depend on the default-specific excluded regressor. This is a generalization of the models in heiss2016inattention,ho2017impact, Abaluck2019, where the consideration mechanism only allows for the possibility that either the default or the entire choice set is considered. We, however, allow for each subset of the choice set containing the default to have its own probability of being drawn and this probability can vary with the DM's preferences. In the second case, we allow the consideration of each alternative to depend on its own excluded regressor, but not on the regressors of other alternatives Goeree2008,Abaluck2019,kawaguchi2019designing. In addition, consideration may depend on preferences -- a feature unique to our paper.

Our second contribution is to provide a simple method to compute our likelihood-based estimator. Its computational complexity grows polynomially in the number of parameters governing the consideration mechanism. Because the SCP generates a natural ordering of alternatives akin to vertical product differentiation, our method does not require enumerating all possible subsets of the choice set. If it did, the computational complexity would grow exponentially with the size of the choice set. Moreover, we compute the utility of each alternative only once for a given value of the preference parameter, gaining enormous computational advantage similar to that of importance-sampling methods.

Our third contribution is to elucidate the applicability and the advantages of our framework over the standard application of full consideration random utility models (RUMs) with additively separable unobserved heterogeneity (e.g., Mixed Logit). First, our model can generate zero shares for non-dominated alternatives. Second, the model has no difficulty explaining relatively large shares of dominated alternatives. Third, in markets with many choice domains, our model can match not only the marginal but also the joint distribution of choices across domains. Forth, our framework is immune to an important criticism by ape:bal18 against using standard RUMs to study decision making under risk. As these authors note, combining standard EUT with additive noise results in non-monotonicity of choice probabilities in the risk preferences, a clearly undesirable feature.

Random preference models like the ones we consider are random utility models as envisioned by Mcfadden1974 Manski2007book. We show that our random preference models can be written as RUMs with unobserved heterogeneity in risk aversion and with an additive error that has a discrete distribution with support $\{-\infty,0\}$. Then, it is natural to draw parallels with the Mixed (random coefficient) Logit model McFaddenTrain2000. In our setting, the Mixed Logit boils down to assuming that, given the DM's risk aversion, her evaluation of an alternative equals its expected utility summed with an unobserved heterogeneity term capturing the DM's idiosyncratic taste for unobserved characteristics of that alternative. However, in some markets it is hard to envision such characteristics.\footnote{Many insurance contracts are identical in all aspects except for the coverage level and price, e.g., employer provided health insurance, auto, or home insurance offered by a single company. In other contexts, unobservable characteristics may affect choice mostly via consideration -- as we model -- rather than via “additive noise”. E.g., a DM may only consider those supplemental prescription drug plans that cover specific medications.} We show that limited consideration models and the Mixed Logit generate several contrasting implications. First, the Mixed Logit generally implies that each alternative has a positive probability of being chosen, while a limited consideration model can generate zero shares by setting the consideration probability of a given alternative to zero. Second, the Mixed Logit satisfies a Generalized Dominance Property that we derive: if for any degree of risk aversion alternative $j$ has lower expected utility than either alternative $k$ or $l$, then the probability of choosing $j$ must be no larger than the probability of choosing $k$ or $l$. Limited consideration models do not necessarily abide Generalized Dominance. Third, in limited consideration models choice probabilities depend on the ordinal expected utility rankings of the alternatives, while in the Mixed Logit it depends on the cardinal ranking. This difference implies that choice probabilities may be monotone in risk preferences in the limited consideration models we propose, while in the Mixed Logit they are not ape:bal18.

We illustrate our method in a study of households' deductible choices across three lines of insurance: auto collision, auto comprehensive, and home (all perils). We aim to estimate the distribution of risk preferences and the consideration parameters and to assess the resulting fit of the models. We find that the $\modA$ model does a remarkable job at matching the distribution of observed choices, and because of its aforementioned properties, outperforms the Mixed Logit. Under the ARC model, we find that although households are on average strongly risk averse, they consider lower coverages more often than higher coverages. We also find support for proportionally shifting consideration. In particular, risk-neutral DMs consider each of the safer alternatives $15\%$ ($11\%$) less often than do extremely risk averse DMs (DMs with median risk aversion).

The rest of the paper is organized as follows. We describe the model of DMs' preferences in Section (ref), and study identification in Section (ref). In Section (ref) we describe the computational advantages of our approach. Section (ref) compares our model to the Mixed Logit. Section (ref) presents our empirical application. Section (ref) contextualizes our contribution relative to the extant literature and offers concluding remarks.

Preferences

Decision Making under Risk in a Market Setting: An Example

Consider as an example the following insurance market, which mimics the setting of our empirical application. There is an underlying risk of a loss that occurs with probability $\mu$ that may vary across DMs. A finite number of alternatives are available to insure against this loss. Conditional on risk type, i.e., given $\mu$, each alternative $j \in \Dc \equiv \{1,\dots,D\}$ is fully characterized by the pair $(d_j,p_j)$. The first element is the insurance deductible, which is the DM's out of pocket expense in the case a loss occurs. Deductibles are decreasing with index $j$, and all deductibles are less than the lowest realization of the loss. The second element is the price (insurance premium), which also varies across DMs. For each DM there is a baseline price $\bar{p}$ that determines prices for all alternatives faced by the DM according to the multiplication rule $p_{j}=g_j\cdot \bar{p}+\delta$. Lower deductibles provide more coverage and cost more, so $g_j$ is increasing with $j$. Both $g_j$ and $\delta$ are invariant across DMs. The lotteries that DMs face are $L_{j}(x)\equiv \left( -p_{j},1-\mu ;-p_{j}-d_j,\mu \right)$, where $x\equiv\bar{p}$. DMs are expected utility maximizers. Given initial wealth $w$, the expected utility of deductible lottery $L_{j}(x)$ is

equation*[equation* omitted — 135 chars of source]

where $u_{\pparam}(\cdot)$ is a Bernoulli utility function defined over final wealth states. We assume that $u_{\pparam}(\cdot)$ belongs to a family of utility functions that are fully characterized by a scalar $\pparam$ (e.g. Constant Absolute Risk Aversion (CARA), Constant Relative Risk Aversion (CRRA), or Negligible Third Derivative (NTD)), which varies across DMs.\footnote{Under CRRA, it is implied that DMs' initial wealth is known to the researcher. NTD utility is defined in Cohen2007 and in Barseghyan2013.}

Given the risk type, the relationship between risk aversion and prices is standard. At sufficiently high $\bar{p}$, less coverage is always preferred to more coverage for all $\pparam$ on the support: $U_{\pparam}(L_{1}(\covx))>U_{\pparam}(L_{2}(\covx)) >\dots> U_{\pparam}(L_{D}(\covx))$. At sufficiently low $\bar{p}$, we have the opposite ordering for all $\pparam$ on the support: $U_{\pparam}(L_{D}(\covx))>U_{\pparam}(L_{D-1}(\covx)) >\dots> U_{\pparam}(L_{1}(\covx))$. At moderate prices, for each pair of deductible lotteries $j<k$ there is a cutoff value $c_{j,k}(x)$ in the interior of $\pparam$'s support, found by solving $U_{\pparam}(L_{j}(\covx))=U_{\pparam}(L_{k}(\covx))$ for $\pparam$. On the left of this cutoff the higher deductible is preferred and on the right the lower deductible is preferred. In other words, $c_{j,k}(x)$ is the unique coefficient of risk aversion that makes the DM indifferent between $L_j(\covx)$ and $L_k(\covx)$, known to the researcher at any given $\covx$. Those with lower $\pparam$ choose the riskier alternative $L_j(\covx)$, while those with higher $\pparam$ choose the safer alternative $L_k(\covx)$. Provided $U_{\pparam}(\cdot)$ is smooth in $\pparam$, $c_{j,k}(x)$ is smooth in $\covx$. In fact, under CARA, CRRA, or NTD, $c_{j,k}(x)$ is a continuously differentiable monotone function. The prices are such that, under CARA, CRRA, or NTD, whenever $U_{\pparam}(L_{1}(\covx))>U_{\pparam}(L_{j}(\covx))$ it is also the case that $U_{\pparam}(L_{1}(\covx))>U_{\pparam}(L_{j+1}(\covx))$.\footnote{We analytically verify this claim for our application in Appendix (ref), but it can also be checked numerically for any given dataset.} As we show below, this can be stated as $c_{1,j}(x)<c_{1,j+1}(x)$. That is, if the DM's risk aversion is so low that she prefers the riskiest lottery to a safer one, then she also prefers it to an even safer one. Finally, there are no three-way ties. That is, for a given $\covx$ there are no alternatives $\{j,k,l\}$ such that $U_{\pparam}(L_{j}(\covx))=U_{\pparam}(L_{k}(\covx))=U_{\pparam}(L_{l}(\covx))$.\footnote{It is straightforward to very this condition, and we do so in our application.}

Preferences with Single Crossing Property

There is a continuum of DMs. Each of them faces a choice among a finite number of alternatives, i.e., a choice set, which is denoted ${\cal{D}}=\{1,\dots,D\}$. The number of alternatives is invariant across DMs. Alternatives vary by their utility-relevant characteristics and are distinguished by (at least) one characteristic, $\dor_j \in \R,~ j \in {\cal{D}}$, which is DM invariant. This characteristic reflects the quality of alternative $j$ (e.g., insurance deductible). When it is unambiguous, we may write $\dor_j$ instead of “alternative $j$”. Other characteristics may vary across DMs or across alternatives. Our analysis rests on the excluded regressor(s) $\covx$. To keep the notation as lean as possible, we state our assumptions and results implicitly conditioning on all remaining characteristics. Hence, alternative $j$ is fully characterized by $(\dor_j,\covx_j)$. We consider two cases. In one case, all $\covx_j$'s are perfectly correlated with a single (common) excluded regressor, $\covx$ (e.g., $\bar{p}$ in our insurance example). In the other case, each $\covx_j$ has its own variation conditional on all other $\covx_k,~ k\neq j$ (e.g., each alternative on the market exhibits locally independent price variation). \phantomsection

assumptionSP{T0} The random variable (or vector) $\covx$ has a strictly positive density on a set $\mathcal{S} \subset \R$ $\left(\mathcal{S} \subset \R^D, ~ \dim \mathcal{S}=D\right)$.

Each DM's valuation of the alternatives is defined by a utility function $U_{\pparam}(d_j,x)$, which depends on a DM-specific index $\pparam$ distributed according to $F(\cdot)$ over a bounded support.\footnote{We assume that while $\pparam$ has bounded support, the utility function is well defined for any real valued $\pparam$.} \phantomsection

assumptionSP{T1} The density of $F(\cdot)$, denoted $f(\cdot)$, is continuous and strictly positive on $[0,\bar \pparam]$ and zero everywhere else.

The DMs' draws of $\pparam$ are not observed by the researcher. We require that DMs' preferences satisfy the Single Crossing Property (SCP). \phantomsection

assumptionSP{T2} [Single Crossing Property] For any two alternatives, $d_j$ and $d_k$, there exists a continuously differentiable function $\cmap_{L,R}: \mathcal{S} \to \R_{[-\infty,\infty]}$ such that \begin{align*} &\util_\pparam(d_L,\covx) > \util_\pparam(\dor_R,\covx) \quad \forall \pparam \in (-\infty,\cmap_{L,R}(\covx)) \\ &\util_\pparam(d_L,\covx) = \util_\pparam(\dor_R,\covx) \quad \pparam = \cmap_{L,R}(\covx) \\ &\util_\pparam(d_L,\covx) < \util_\pparam(\dor_R,\covx) \quad \forall \pparam \in (\cmap_{L,R}(\covx),\infty). \end{align*} where $(L,R)=(j,k)$ or $(L,R)=(k,j)$. We refer to $\cmap_{L,R}(\cdot)$ as the cutoff between $d_L$ and $d_R$.

The SCP implies that the DM's ranking of alternatives is monotone in $\pparam$. In the context of risk preferences, if a DM with a certain level of risk aversion prefers a safer asset to a riskier one, then all DMs with higher risk aversion also prefer the safer asset. Since the cutoffs may be infinite, the SCP does not exclude dominated alternatives.

definition[Dominated Alternatives] Given $\covx$, alternative $d_j$ is dominated if there exists an alternative $d_k$ such that $\forall \pparam \in \R$, $\util_\pparam(d_k,\covx) > \util_\pparam(\dor_j,\covx)$.

We now establish some useful facts that follow from Assumption (ref). First, the index $L$ in $c_{L,R}(\cdot)$ indicates the alternative that is preferred on the left of the cutoff. It is without loss of generality to assume $L=\min(j,k)$ and $R=\max(j,k)$ because of the following fact:

fact[Natural Ordering of Alternatives] Suppose Assumption (ref) holds. Then alternatives can be enumerated such that as $\pparam\rightarrow -\infty$, $\util_\pparam(d_1,\covx) > \util_\pparam(\dor_2,\covx) >\cdots>\util_\pparam(d_D,\covx)$ for all $\covx$ at which no alternative is dominated.

We assume that alternatives are enumerated according to the Natural Ordering of Alternatives.\footnote{Under this enumeration, $d_{j}$ will be ordered in either ascending or descending order. In our example from the previous section, since $d_j$ refers to the deductible and $\pparam$ is the risk aversion coefficient, the natural ordering implies $d_1>d_2>\dots>d_D.$} As the next fact shows, for high values of $\pparam$ the preference over the Natural Ordering of Alternatives is reversed.

fact[Rank Switch] Suppose Assumption (ref) holds. Consider any $\covx$ such that no alternative is dominated. As $\pparam\rightarrow \infty$, $\util_\pparam(d_1,\covx) < \util_\pparam(\dor_2,\covx) <\cdots<\util_\pparam(d_D,\covx).$

The SCP also has implications for the relative position of the cutoffs. For readability, we state them for alternatives $\{d_1, d_2,d_3\}$, but they hold for any $\{d_j,d_k,d_l\}$, $j<k<l$.

fact[Simple Relative Order of Cutoffs] Suppose Assumption (ref) holds. Given $\covx$, if $\cmap_{1,2}(x)<\cmap_{1,3}(x)$, then $\cmap_{1,3}(x)<\cmap_{2,3}(x)$ or both $d_1$ and $d_2$ dominate $d_3$ ($\cmap_{1,3}(x)=\cmap_{2,3}(x)=\infty$).

The next fact concerns the relative order of cutoffs for non-dominated alternatives. Before stating it, it is convenient to define Never-the-First-Best Alternatives.

definition[Never-the-First-Best] Given $\covx$, alternative $d_j$ is Never-the-First-Best in $\cal{D}$ if for every $\pparam$ there exists another alternative $d_k(\pparam)$ in $\cal{D}$ such that $\util_\pparam(d_k(\pparam),\covx)>\util_\pparam(d_j,\covx)$.
fact[Cutoff Relative Order] Suppose that Assumption (ref) holds. If, given $\covx$, alternatives $d_1$, $d_2$, and $d_3$ are not dominated, then one and only one of the following cases holds: \begin{enumerate} • $\cmap_{1,2}(x)<\cmap_{1,3}(x)<\cmap_{2,3}(x)$ and $d_2$ is the first best in $\{d_1,d_2,d_3\}$, $\forall \pparam \in (\cmap_{1,2}(x),\cmap_{1,3}(x))$; • $\cmap_{1,2}(x)>\cmap_{1,3}(x)>\cmap_{2,3}(x)$ and $d_2$ is Never-the-First-Best in $\{d_1,d_2,d_3\}$; • $\cmap_{1,2}(x)=\cmap_{1,3}(x)=\cmap_{2,3}(x)$ and $d_2$ is strictly worse than either $d_1$ or $d_3$ for all $\pparam$ except for $\pparam=\cmap_{1,2}(x)$ where there is a three-way tie among these alternatives. \end{enumerate}

Fact (ref) is a convenient way to distill and exploit the SCP. In particular, for any $\covx$, the complete preference order of the alternatives is known for all DMs as well as the identity (of the preference parameter) of the DM indifferent between any two alternatives $d_j$ and $d_k$.

Identification

The classic identification argument for discrete choice under full consideration rests on the following four canonical assumptions. \phantomsection

assumptionSP{I0} The random variable (or vector) $x$ is independent of preferences.

\phantomsection

assumptionSP{I1} $\exists \covxs \subset \mathcal{S}$ s.t. $\cmap_{1,2}(x)$ covers the support of $\pparam$: $[0,\bar{\pparam}]\subset \{\cmap_{1,2}(x), x\in \covxs\}$.

\phantomsection

assumptionSP{I2} Consideration is independent of preferences.

\phantomsection

assumptionSP{I3} Consideration is independent of $\covx$.

The last two conditions are vacuous in the standard full consideration model, while the first two are typically stated as data requirements.

We first discuss how to obtain identification and the role of Assumptions (ref)-(ref) in the simplest case of two alternatives (Section (ref)). We then consider the general model with $D$ alternatives. Table (ref) organizes our results by assumptions imposed, the consideration mechanism assumed, data availability, and the theorems' conclusions. Theorems (ref)-(ref) in Section (ref) demonstrate that the preference distribution and some features of the consideration mechanism are identified with a single excluded regressor. Next, we show that alternative-specific variation allows for identification of both the preference distribution and the consideration mechanism when consideration depends on preferences and one of the excluded regressors (Theorem (ref) and Corollary (ref) in Section (ref)). We discuss testing for limited consideration in Section (ref). We then turn to the ARC model in Section (ref). We show that the full model is identified with a single excluded regressor (Theorem (ref)). Moreover, identification attains for a particular case where consideration depends on preferences (Theorem (ref)). Finally, Theorem (ref) shows that with alternative-specific variation, identification attains when consideration of each alternative depends both on preferences and its own regressor, without requiring full support.

table[table omitted — 2,258 chars of source]

The Role of the Canonical Assumptions

Let the choice set be binary and suppose that the DM considers both alternatives. In addition, let $\covx$ be a scalar so that there is a single excluded regressor. Under Assumptions (ref)-(ref) and (ref)-(ref), any realization of $x$ is associated with a single conditional moment in the data:

equation*[equation* omitted — 85 chars of source]

because the DM chooses $d_1$ if and only if her preference parameter is less than $\cmap_{1,2}(x)$. The distribution $F(\cdot)$ is non-parametrically identified, since for any $\pparam$ on the support there is an $x$ such that $\pparam=\cmap_{1,2}(x)$.

We emphasize two points. First, given a family of utility functions, for any $x$ the value of the cutoff can be solved for. Hence, the function $\cmap_{1,2}(x)$ (and its derivatives) can be treated as data. Second, Assumption (ref) requires that the cutoff reaches both ends of the support: there exist $x^0$ and $x^1$ such that $F(\cmap_{1,2}(x^0))=0$ and $F(\cmap_{1,2}(x^1))=1$.

Turning to limited consideration, suppose that $d_1$ is considered with probability $0<\aparam_{1}\leq 1$, and whenever it is considered so is $d_2$.\footnote{With two alternatives this implies that $d_2$ is always considered.} Then, $d_1$ is chosen when it is considered and it is preferred to $d_2$, yielding:

align[align omitted — 175 chars of source]

At first glance, it appears that the distribution of preferences is identified up to a constant. Yet, at the boundary of the support $\Pr(d=d_1|x^1) = \aparam_1 F(\cmap_{1,2}(x^1)) = \aparam_1$, so that $\aparam_{1}$ is identified. Once $\aparam_1$ is known, the distribution $F(\cdot)$ is identified by varying $\cmap_{1,2}(\covx)$ over the support of $\pparam$, similar to the full consideration case. We now explore what happens to identification if Assumptions (ref)--(ref) are not satisfied.

Assumption (ref) fails: the variation in $x$ is not independent of preferences. Then $F(\cdot)$ is not non-parametrically identified under either full or limited consideration.

Assumption (ref) fails: the variation in $\covx$ is such that $\cmap_{1,2}(x)$ only covers an interval $[\pparam^{l},\pparam^{u}] \subset [0,\bar{\pparam}]$. Then the data provide no information about preferences outside of the interval $[\pparam^{l},\pparam^{u}]$. Inside the interval, the conditional distribution $F(\pparam|\pparam \in [\pparam^{l},\pparam^{u}])=\frac{F(\pparam)-F(\pparam^{l})}{F(\pparam^{u})-F(\pparam^{l})}$ is identified under both limited and full consideration. The consideration probability (and hence the scale of $F(\cdot)$) is partially identified and satisfies the bounds $\Pr(d=d_1|\covx^{u}) \leq \aparam_1 \leq 1$, where $\covx^{u}$ is such that $\cmap_{1,2}(\covx^{u})=\pparam^{u}$. Point identification can be attained if an additional assumption is maintained to pin down the scale of $F(\cdot)$. For example, one can simply assume full consideration and set $\aparam_{1}=1$.

Assumption (ref) fails: $\aparam_{1}$ depends on preferences and this dependence is arbitrary. Then identification breaks down completely as there is one data moment to identify two unknown objects. However, since we assume -- as it is common in the econometrics literature -- that the density function of $\pparam$ is continuous and strictly positive, identification is possible for some types of dependence between consideration and preferences. Suppose there are two consideration types:

align*[align* omitted — 198 chars of source]

where $\pparam^{*}$ is an unobserved breakpoint. We show that $\underline{\aparam}_1$, $\overline{\aparam}_1$, and $\pparam^{*}$ are identified. First, the product $\aparam_1(\pparam)f(\pparam)$ is identified under Assumptions (ref), (ref), and (ref), since

equation[equation omitted — 186 chars of source]

at $\pparam=\cmap_{1,2}(x)$. The product $\aparam_1(\pparam)f(\pparam)$ is discontinuous only at the point $\pparam^{*}$. Thus, the breakpoint is identified by continuously varying $c_{1,2}(\covx)$ across $[0,\bar{\pparam}].$ Next, the ratio $\frac{\underline{\aparam}_1}{\overline{\aparam}_1}$ is identified by the ratio of the right and left derivatives of $\Pr(d=d_1|x)$ at the breakpoint $\covx^{*}$ ($\pparam^{*}=\cmap_{1,2}(\covx^{*})$). The quantity $F(\pparam^{*})$ is identified by the ratio:

align*[align* omitted — 205 chars of source]

Hence, $\underline{\aparam}_1$ and $\overline{\aparam}_1$ are identified. Identification of $F(\cdot)$ on the entire support follows from Assumption (ref). The same argument above applies if the probability of considering an alternative discretely jumps in $\covx$ (i.e., Assumption (ref) fails). Concretely, suppose there is a breakpoint in $\aparam_{1}(\covx)$ at $\covx^{*}$ and let $\pparam^{*}=\cmap_{1,2}(x^{*})$. The breakpoint $\covx^{*}$ is identified by the point of discontinuity in Equation (ref), and the rest follows.

To summarize the case of the binary choice set, the only seemingly real difference in identification is that without large support the scale of the preference distribution $F(\cdot)$ is partially identified under limited consideration, while it is assumed to be known under full consideration. The key to identification is a one-to-one mapping from a data moment, $\Pr\left(d=d_1|\covx\right)$, and the preference distribution $F(\cdot)$ at a single point on the support, $\cmap_{1,2}(x)$. As we will show next, even with just a single excluded regressor, Assumptions (ref)-(ref) allow for such a mapping to be constructed for a generic consideration mechanism and a choice set of arbitrary size.

Single Common Excluded Regressor

We start by introducing general notation for consideration probabilities.

definitionLet $\mathcal{Q}_{\pparam}^{\covx}(\cal{K})$ be the probability that, given $\covx$, the DM with preference parameter $\pparam$ draws consideration set $\cal{K} \subset \Dc$ conditional on $\covx$. Let $\mathcal{O}_{\pparam}^{\covx}(\mathcal{A};\mathcal{B})$ be the probability that, given $\covx$, every alternative in set $\mathcal{A}$ is in the consideration set and every alternative in set $\mathcal{B}$ is not for the DM with preference parameter $\pparam$: \[ \mathcal{O}_{\pparam}^{\covx}(\mathcal{A};\mathcal{B})\equiv\sum_{\Kc: ~\mathcal{A} \subset \Kc, ~ \mathcal{B} \cap \Kc =\emptyset}\mathcal{Q}_{\pparam}^{\covx}(\Kc). \]

The subscript is suppressed when consideration does not depend on preferences, and the superscript is suppressed when it does not depend on the excluded regressor(s).

To ease exposition, we build our discussion around a choice set with three alternatives, $ \Dc = \{d_1,d_2,d_3\}$, such that $\cmap_{1,2}(x)<\cmap_{1,3}(x)<\cmap_{2,3}(x)$ for all $\covx$. That is, by Fact (ref), if $U_{\pparam}(d_1,\covx)>U_{\pparam}(d_2,\covx)$ then $U_{\pparam}(d_1,\covx)>U_{\pparam}(d_3,\covx)$ for all $\covx$. Suppose consideration is independent of preferences and of the excluded regressor. Then the choice frequencies of $d_1$ and $d_3$ are

align*[align* omitted — 333 chars of source]

Consider the expression for $\Pr(d=d_1|x)$. Its RHS has three terms. The first term captures the case when $d_1$ is considered along with $d_2$, which happens with probability $\mathcal{O}(\{d_1,d_2\};\emptyset).$ Given the relative position of the cutoffs, whether $d_3$ is considered or not is irrelevant. The DM will choose $d_1$ over $d_2$ if and only if her preference parameter is below $\cmap_{1,2}(x)$. The second term captures the case when $d_1$ is considered along with $d_3$, but $d_2$ is not considered, which happens with probability $\mathcal{O}(\{d_1,d_3\};d_2)$. Then the relevant cutoff for choosing $d_1$ is $\cmap_{1,3}(x)$. Third, when $d_1$ is the only alternative considered, it is chosen regardless of the DM's risk aversion. This event occurs with probability $\mathcal{O}(d_1;\{d_1,d_2\})$.

Since there are two cutoffs, $\cmap_{1,2}(x)$ and $\cmap_{1,3}(x)$, that enter the moment $\Pr(d=d_1|x)$, there is not, without additional assumptions, a one-to-one mapping between the moment and the preference distribution at one point on the support, as it was the case in Section (ref). That is, as $\covx$ changes, the observed choice frequency of $d_1$ may change because of two types of marginal DMs: those indifferent between $d_1$ and $d_2$, and those indifferent between $d_1$ and $d_3$. This is apparent in the following derivative:

align[align omitted — 193 chars of source]

The corresponding equation for $\Pr(d=d_3|x)$ does not immediately help, as it brings about $f(\cdot)$ evaluated at yet another cutoff, $\cmap_{2,3}(x)$:

align[align omitted — 176 chars of source]

Identification with Large Support

When Assumption (ref) holds, we can construct a one-to-one mapping sequentially. The algorithm for doing so consists of four steps. First, we rewrite Equation (ref) as

equation[equation omitted — 183 chars of source]

where $\phi \equiv\frac{ \mathcal{O}(\{d_1,d_3\};d_2)}{\mathcal{O}(\{d_1,d_2\};\emptyset)}$ and $\hat f(\pparam) \equiv \mathcal{O}(\{d_1,d_2\};\emptyset) f(\pparam)$. Second, for $\pparam$'s near the far end of the support, we can find $\covx$ and $\covx'$ such that $\pparam= \cmap_{1,2}(x)<\bar{\pparam}<\cmap_{1,3}(x)$ and $\pparam=\cmap_{1,3}(x')<\bar{\pparam}<\cmap_{2,3}(x')$. For any such pair, $f(\cmap_{1,3}(x))=f(\cmap_{2,3}(x'))=0$, and, hence, by Equations (ref) and (ref):

align*[align* omitted — 224 chars of source]

The first equation identifies $\hat{f}(\pparam)$, while the ratio of the two equations identifies $\phi$. Third, whenever $\hat{f}(\cmap_{1,3}(\covx))$ is known, $\hat{f}(\cmap_{1,2}(\covx))$ is uniquely pinned down by Equation (ref). Because $\cmap_{1,2}(\covx)<\cmap_{1,3}(\covx)$, $\forall \covx$, we can learn $\hat{f}(\cdot)$ sequentially:

enumerate• Take an $\covx^1$ such that $\hat{f}(\cmap_{1,3}(\covx^1))$ is already known, learn $\hat{f}(\cmap_{1,2}(x^1))$; • Take $\covx^{2}$ such that $\cmap_{1,3}(\covx^{2})=\cmap_{1,2}(\covx^{1})$, learn $\hat{f}(\cmap_{1,2}(\covx^{2}))$; • Let $\covx^{1}=\covx^{2}$. Repeat Step 2 until the entire support has been covered, i.e., $\cmap_{1,2}(\covx^{2})\leq 0$.

For this approach to work, $ \cmap_{1,3}(\covx)$ cannot “catch up” to $\cmap_{1,2}(\covx)$ (i.e., as assumed, $\cmap_{1,2}(x) < \cmap_{1,3}(\covx)$ whenever $\cmap_{1,2}(x)$ is on the support). This requires that DMs with preference coefficients on the support are never indifferent between $d_1$ and two other alternatives -- i.e. there are no three way ties involving $d_1$. Fourth, integration of $\hat{f}(\pparam)$ over the entire support recovers the scale and the true density. Indeed, \[ \int_{0}^{\bar{\pparam}} \hat{f}(\pparam)d\pparam=\mathcal{O}(\{d_1,d_2\};\emptyset)\int_{0}^{\bar{\pparam}} f(\pparam)d\pparam =\mathcal{O}(\{d_1,d_2\};\emptyset) \] pins down $\mathcal{O}(\{d_1,d_2\};\emptyset)$, and hence $f(\cdot)$ is identified. A generalization of this strategy yields our first formal result.

theoremSuppose Assumptions (ref), (ref), (ref), (ref)-(ref) hold, and \begin{enumerate} • The consideration mechanism is s.t. with positive probability $d_1$ and $d_2$ are considered together; • Assumption (ref) holds for $\mathcal{X} \subset \mathcal{S}$ s.t. $\forall \covx \in \covxs$ \[ U_{\pparam}(d_1,x)>U_{\pparam}(d_j,\covx) \Rightarrow U_{\pparam}(d_1,\covx)>U_{\pparam}(d_{j+1},\covx), \quad \forall j>1. \] \end{enumerate} Then $f(\cdot)$ is identified and so are $\mathcal{O}(d_1;\emptyset)$ and $\mathcal{O}(\{d_1,d_2\};\emptyset)$. For $j>2$, if $\Pr(d=d_j|\covx)>0$ for some $\covx$, then $\mathcal{O}(\{d_1,d_j\};\{d_2,\dots, d_{j-1}\})$ is identified.

The first assumption of the theorem ensures that a generalized version of Equation (ref) is informative. The second assumption implies that the cutoffs for alternative $d_1$ are ordered: $\cmap_{1,j}(x)<\cmap_{1,j+1}(x)$. While Theorem (ref) requires large support for the excluded regressor, it does not generally require it to exhibit variation that forces alternative $d_1$ to go from being the first best to the least preferred. Rather, the theorem requires that at one extreme of the support alternative $d_1$ dominates all others. However, at the other extreme we only require that $d_2$ is preferred to $d_1$ for all DMs. Identification is attained for any consideration mechanism that allows $d_1$ and $d_2$ to be considered together with positive probability. Moreover, if the probability of being considered together is zero for $d_1$ and $d_2$, but positive for $d_1$ and $d_3$, the theorem still holds as long as the assumptions of the theorem hold for $d_3$ instead of $d_2$. Theorem (ref) identifies some features of the consideration mechanism. These features may be sufficient for identifying the entire mechanism. In particular, as shown in Section (ref), Theorem (ref) yields identification of the ARC model, including the consideration mechanism.

Dependence between consideration and preferences. We next generalize the example in Section (ref) by allowing for high/low consideration types.

assumptionSP{I2.BCT}[Binary Consideration Types] For some unknown $\pparam^* \in (0, \bar{\pparam})$: \[ \mathcal{Q_{\pparam}(K)} = \begin{cases} \mathcal{\underline{Q}(K)} & \text{if } \pparam < \pparam^* \\ \mathcal{\overline{Q}(K)} & \text{if } \pparam > \pparam^* \\ \end{cases} \] where, $\forall \pparam$ and $\forall \mathcal{K} \subset \Dc$, $\sum_{\mathcal{K} \subset \Dc} \mathcal{Q_{\pparam}(K)}=1$ and $\mathcal{Q_{\pparam}(K)}\geq 0$.
theoremSuppose Assumptions (ref), (ref), (ref), (ref)-(ref), and Condition 2 of Theorem (ref) hold. Suppose Condition 1 of Theorem (ref) holds for all $\pparam$. Then $f(\cdot)$ is identified and so is $\mathcal{O}_{\pparam}(\{d_1,d_2\};\emptyset)$. Suppose $\frac{d\Pr(d=d_1|\covx)}{dx}$ is discontinuous. Then $\pparam^*$ is identified. If, in addition, $\cmap_{1,j}(\covx) < \pparam^*$ for some $\covx \in \mathcal{X}$ and $j>2$, then $\mathcal{O}_{\pparam}(\{d_1,d_j\};\{d_2,\dots, d_{j-1}\})$ is also identified.

A discontinuity in $\frac{d\Pr(d=d_1|\covx)}{dx}$ may occur when a cutoff $\cmap_{1,j}(\covx)$ crosses $\pparam^*$. In some cases it may not happen despite binary consideration. For example, the probability of considering $d_1$ and $d_2$ may jump but in a way that $\mathcal{O}_{\pparam}(\{d_1,d_2\};\emptyset)$ remains constant. In such a case, $f(\cdot)$ is identified but not necessarily the breakpoint $\pparam^*$.

The theorem holds if Assumption (ref) is replaced with \[ \mathcal{Q}^{\covx}(\mathcal{K}) =

cases\mathcal{Q(K)} & if \covx < \covx^* \\ \mathcal{\overline{Q}(K)} & if \covx > \covx^*

\] for some unknown $\covx^* \in \mathcal{S}$. In sum, preferences can be identified even when there are threshold effects affecting consideration. Assumption (ref) is one instance where Assumption (ref) does not hold but identification attains. Another instance, which we establish for the ARC model in Section (ref), is proportionately shifting consideration.

Identification without large support

Returning to our example with three alternatives, it is immediate to see that if whenever $d_1$ is considered so is $d_2$, i.e. $\mathcal{O}(\{d_1,d_3\};d_2) = 0$, the one-to-one mapping is restored. Indeed, the second term on the RHS of Equation (ref) disappears and we are back to Equation (ref).

propositionSuppose Assumptions (ref), (ref), (ref), (ref)-(ref) hold, and \begin{enumerate} • The consideration mechanism is such that $d_1$ is considered with positive probability and whenever it is considered so is $d_2$; • There exists $\covxs \subset \mathcal{S}$ such that $\cmap_{1,2}(\covx)$, $\covx \in \covxs$, covers $[\pparam^{l},\pparam^{u}]\subset [0,\bar \pparam]$ and $\forall \covx \in \covxs$ \[ U_{\pparam}(d_1,\covx)>U_{\pparam}(d_2,\covx) \Rightarrow U_{\pparam}(d_1,\covx)>U_{\pparam}(d_j,\covx), \quad \forall j>2. \] \end{enumerate} Then $F(\pparam|\pparam \in [\pparam^{l},\pparam^{u}])$ is identified.

The proposition above uses (the derivative of) $\Pr(d=d_1|x)$ to create the one-to-one mapping from data to the preference density function. Depending on the consideration mechanism, the same can be achieved using the derivative of $\Pr(d\in \{d_1,d_2,\dots,d_j\}|x)$.

definition[Loosely Ordered Consideration] The consideration mechanism is loosely ordered around $j$, $j<D$, if whenever alternatives $d_k$ and $d_l$, $k\leq j<l$, are both considered, so are $d_{j}$ and $d_{j+1}$. In addition, $d_{j}$ and $d_{j+1}$ have a positive probability of being considered together.
theoremSuppose Assumptions (ref), (ref), (ref), (ref)-(ref) hold, and \begin{enumerate} • The consideration mechanism is loosely ordered around $j$. • There exists $\covxs \subset \mathcal{S}$ such that $\cmap_{j,j+1}(\covx)$, $\covx \in \covxs$, covers $[\pparam^{l},\pparam^{u}]\subset [0,\bar \pparam]$ and $\forall \covx \in \covxs$ \begin{align*} U_{\pparam}(d_j,\covx)>U_{\pparam}(d_{j+1},\covx) &\Rightarrow U_{\pparam}(d_j,\covx)>U_{\pparam}(d_k,\covx), \quad \forall k>j+1,\\ U_{\pparam}(d_{j+1},\covx)>U_{\pparam}(d_j,\covx) &\Rightarrow U_{\pparam}(d_{j+1},\covx)>U_{\pparam}(d_k,\covx), \quad \forall k<j. \end{align*} \end{enumerate} Then $F(\pparam|\pparam \in [\pparam^{l},\pparam^{u}])$ is identified.

Condition 1 in Theorem (ref) -- a loosely ordered consideration mechanism -- splits the choice set into “low quality” and “high quality” sets. Any subset of the low quality set can form the consideration set and so can any subset of the high quality set. However, if a consideration set contains both high and low quality alternatives, then it must also contain the “bridging” alternatives $\{d_{j},d_{j+1}\}$. The following mechanisms can generate loosely ordered consideration:

enumerate• Bottom-Up consideration: Alternative $d_k$ is considered only if $d_{k-1}$ is considered; • Top-Down consideration: Alternative $d_k$ is considered only if $d_{k+1}$ is considered; • Center-to-edges consideration: Alternative $d_{j}$, $1<j<D$, is always considered. Alternative $d_k$, $k>j$, is considered only if $d_{k-1}$ is considered. Alternative $d_k$, $k<j$, is considered only if $d_{k+1}$ is considered; • Trimmed-from-the-edges consideration: Only consideration sets of the form $\mathcal{K} = \{d_k,d_{k+1},\dots,d_{k+l}\}$ can occur with positive probability.

The identification result in Theorem (ref) extends to mixtures of these mechanisms. They cover a wide array of models including versions of threshold models kimya2018choice, (partial) elimination-by-aspects Tversky72, extremeness aversion simonson_tversky_1992, and edge aversion teigen1983studies,christenfeld1995choices,rubinstein1997naive,attali2003guess, as well as models that embed budget or liquidity constraints.

Condition 2 in Theorem (ref) requires that whenever a DM prefers $d_j$ to $d_{j+1}$, she also prefers $d_j$ to all high quality alternatives; and whenever a DM prefers $d_{j+1}$ to $d_{j}$, she also prefers $d_{j+1}$ to all low quality alternatives. This condition can be tested in any given dataset and is automatically satisfied if no alternative is never-the-first-best.

The fundamental difference between Theorems (ref) and (ref) is that the former imposes the large support requirement, while the latter does not. On the other hand, Theorem (ref) imposes less restrictions on the consideration mechanism than Theorem (ref).

Alternative-specific Excluded Regressors

With alternative-specific excluded regressors we can allow for consideration to depend on preferences. To illustrate, we continue to assume that the choice set is $\{d_1,d_2,d_3\}$. However, now each alternative has its own regressor $x_j$ that only affects the utility of alternative $j$: $\covx=(\covx_{1},\covx_{2},\covx_{3})$. In addition, these regressors vary independently of one another and each consideration set contains at least two alternatives.

Identification is built on the following insight. Consider the change in the choice frequency of alternative $d_1$ in response to an incremental change in $\covx_2$ (e.g., a price increase for alternative $d_2$). The DMs who may switch to $d_1$ are those indifferent between $d_1$ and $d_2$ and consider them both. If these DMs prefer $d_1$ and $d_2$ to $d_3$, whether $d_3$ is considered is irrelevant; otherwise, for the response to occur, $d_3$ should not be considered. These two cases translate to the following statements: (i) $\cmap_{1,2}(x)<\cmap_{1,3}(x)<\cmap_{2,3}(x)$; and (ii) $\cmap_{2,3}(x)<\cmap_{1,3}(x)<\cmap_{1,2}(x)$ and $d_3$ is not considered. No other ordering of cutoffs can occur by Fact (ref). With alternative specific variation we can construct two vectors of regressors, $\covx^i$ and $\covx^{ii}$, such that $\pparam = c_{1,2}(\covx^i) <\cmap_{1,3}(\covx^i)<\cmap_{2,3}(\covx^i)$ and $c_{2,3}(\covx^{ii}) <\cmap_{1,3}(\covx^{ii})<\cmap_{1,2}(\covx^{ii})=\pparam$.\footnote{ First, we can construct an $\covx$ such that $U_{\pparam}(d_1,\covx) = U_{\pparam}(d_2,\covx) = U_{\pparam}(d_3,\covx)$. To do so, we fix the price of the first alternative, $\covx_1$, and find a price for the second alternative that makes the DM with preference parameter $\pparam$ indifferent between $d_1$ and $d_2$. Since $\covx_3$ does not affect the utility of $d_1$ nor $d_2$, we can find an $\covx_3$ so that the DM is indifferent between $d_3$ and $d_1$, and hence she is indifferent between all three alternatives. The two cases are then constructed by taking a small perturbation of $\covx_3$. Taken in a direction that reduces $U_{\pparam}(d_3,\covx)$ generates Case (i); and in the opposite direction Case (ii).} The derivative of the choice frequency of $d_1$ with respect to $\covx_2$ for these cases are, respectively:

align*[align* omitted — 450 chars of source]

It follows that $\mathcal{Q}_{\pparam}(\{d_1,d_2\})f(\pparam)$ and $\mathcal{Q}_{\pparam}(\{d_1,d_2,d_3\})f(\pparam)$ are identified. In a similar fashion, $Q_{\pparam}(\{d_1,d_3\})f(\pparam)$ and $Q_{\pparam}(\{d_2,d_3\})f(\pparam)$ are identified. Hence, the consideration probability of each non-singleton set is identified up to the same scale. The scale, however, is identified because consideration probabilities must sum to one: $f(\pparam)= \sum_{\mathcal{K}} \mathcal{Q}_{\pparam}(\mathcal{K})f(\pparam)$. Hence, $\mathcal{Q}_{\pparam}(\mathcal{K})$ is also identified for each $\mathcal{K}$. The following theorem generalizes this idea.

definition[Alternative-Specific Variation] We say that there is alternative-specific variation if $U_{\pparam}(d_j,\covx)$ depends only on $\covx_j$: $\frac{\partial U_{\pparam}(d_j,\covx)}{\partial x_k}\neq 0 \Leftrightarrow k=j$.
theoremSuppose Assumptions (ref), (ref), (ref)-(ref) hold, there is alternative-specific variation, and the choice set contains at least three alternatives. Suppose \begin{enumerate} • Each consideration set contains at least two alternatives and $\mathcal{Q}_{\pparam}(\cdot)$ is measurable; • For a given value of $\pparam$, there exists an $\covx$ with an open neighborhood around it in $\mathcal{S}$ s.t. \begin{align*} & U_{\pparam}(d_1,\covx) = U_{\pparam}(d_2,\covx) =\cdots= U_{\pparam}(d_D,\covx). \end{align*} \end{enumerate} Then $f(\pparam)$ is identified and so are $\mathcal{Q}_{\pparam}(\cal{K})$, $\forall \mathcal{K}\ \subset \Cs$.

The assumptions of the theorem above rule out singleton (and empty) consideration sets: identification is impossible with singleton consideration sets and arbitrary dependence on preferences, because any empirical choice frequency can be explained by such consideration sets. An alternative approach is to have one alternative -- the “default” -- that is always considered as the following corollary demonstrates. The identification argument exploits the response of $\Pr(d=d_j|\covx)$ to changes in $x_k$, but not the response of $\Pr(d=d_k|\covx)$ to changes in $\covx_j$. Hence, $D-1$ excluded regressors are sufficient for identification, allowing for arbitrary dependence of consideration on one (the default's) excluded regressor.

corollarySuppose Assumptions (ref), (ref)-(ref) hold, there is alternative-specific variation, the choice set contains at least three alternatives, all consideration sets contain $d_1$, and \begin{enumerate} • Consideration is independent of $\covx_{-1}\equiv(\covx_2,\dots,\covx_D)$: $\mathcal{Q}_{\pparam}^{\covx}(\mathcal{K})=\mathcal{Q}_{\pparam}^{\covx_1}(\mathcal{K})$, and $\mathcal{Q}_{\pparam}^{\covx_1}(\cdot)$ are measurable functions, continuous in $\covx_{1}$; • The consideration of $\mathcal{K}=\{d_1\}$ is independent of $\pparam$: $\mathcal{Q}_{\pparam}^{\covx_1}(d_1) =\mathcal{Q}^{\covx_1}(d_1)<1$, $\forall \pparam$; • For a given value of $\covx_1$ and each value of $\pparam \in [0,\bar{\pparam}]$, there exists an $\covx_{-1}$ and an open neighborhood around $x=(x_1,x_{-1})$ in $\mathcal{S}$ s.t. \begin{align*} & U_{\pparam}(d_1,\covx) = U_{\pparam}(d_2,\covx) =\cdots= U_{\pparam}(d_D,\covx). \end{align*} \end{enumerate} Then $f(\pparam)$ is identified and so are $\mathcal{Q}_{\pparam}^{\covx_{1}}(\cal{K})$, $\forall \mathcal{K} \subset \Cs$, for all $\pparam$ on the support.

Corollary (ref) generalizes the model of heiss2016inattention,ho2017impact in two dimensions. First, here each subset of the choice set containing the default has its own probability of being drawn. Second, this probability can vary with the DM's preferences as well as with the excluded regressor of the default alternative.

Testing for limited consideration

Since full consideration is a special case of limited consideration, it follows from the identification results above that under the SCP one can test for full consideration. The theorem below states one way of doing so without: (1) relying on large support; (2) specifying a consideration mechanism; or (3) invoking the independence assumptions (ref) and (ref).

propositionSuppose Assumptions (ref), (ref)-(ref) hold. Suppose there exist $\covx,\covx' \in \mathcal{S}$, and sets $\Lc, \Lc' \subset \Dc$ s.t. for some $\pparam^* \in [0,\bar{\pparam}]$ \begin{enumerate} • $\argmax_{j \in \Dc} \util_\pparam(d_j,x) \in \Lc$, $\forall \pparam \in [0,\pparam^*)$, and $\argmax_{j \in \Dc} \util_\pparam(d_j,x) \in \Dc \setminus \Lc$, $\forall \pparam \in (\pparam^*,\bar \pparam]$$\argmax_{j \in \Dc} \util_\pparam(d_j,x') \in \Lc'$, $\forall \pparam \in [0,\pparam^*)$, and $\argmax_{j \in \Dc} \util_\pparam(d_j,x') \in \Dc \setminus \Lc'$, $\forall \pparam \in (\pparam^*,\bar \pparam]$ \end{enumerate} If $\Pr(d\in \Lc|x)\neq \Pr(d\in \Lc'|x')$, then there is limited consideration.

Condition 1 of the theorem requires that, given $\covx$, the first-best alternative belongs to $\Lc$ for all DMs with $\pparam <\pparam^*$ and to $\Dc\setminus\Lc$ for all DMs with $\pparam > \pparam^*$. Condition 2 is the identical requirement, but given $\covx'$ and stated for $\Lc'$. Under these conditions and full consideration, the probability of choosing an alternative in $\Lc$ or, respectively, $\Lc'$ should be $F(\pparam^*)$ in both cases. Thus, if $\Pr(d\in \Lc|x)\neq \Pr(d\in \Lc'|x')$, then there is a limited consideration mechanism pushing DMs' choices away from $\Lc$ and $\Lc'$ at different rates.

The $\modA$ Model

We now introduce a specific consideration mechanism, while maintaining the preference structure, including the SCP, from Section (ref). We refer to this model as the Alternative-specific Random Consideration ($\modA$) model manski1977structure,man:mar14. Each alternative $\dor_j$ appears in the consideration set with probability $\aparam_j$ independently of other alternatives. For now, we assume that these probabilities do not depend on DMs' preferences or the excluded regressor. Once the consideration set is drawn, the DM chooses the best alternative according to her preferences. To avoid empty consideration sets, following manski1977structure, we assume that at least one alternative whose identity is unknown to the researcher is always considered.\footnote{In the previous version of this paper BaMoTh19 this completion rule is called Preferred Option(s). There we also provide identification results for other completion rules, including Coin Toss (if the empty consideration set is drawn, the DM randomly uniformly picks one alternative from the choice set, i.e. each alternative has probability $1/D$ of being chosen), Default Option (there is a preset alternative that is chosen if the empty set is drawn), and Outside Option (the DM exits the market if the empty set is drawn).} \phantomsection

assumptionSP{ARC}[The Basic ARC Model] The probability that the consideration set takes realization $\Ks$ is \begin{align*} \mathcal{Q}(\Ks) & \equiv \prod_{k \in \Kc }\aparam_k \prod_{k \in \Dc \setminus \Kc} (1-\aparam_k), \quad \forall \Ks \subset \Cs, \end{align*} where $\aparam_j>0, \, \forall j$, and $\exists d^*$ s.t. $\aparam_{d^*}=1$.

By assuming $\aparam_{j}>0$, we omit never-considered alternatives from the choice problem. Since a never-considered alternative is never compared to any other alternative, whether it is in the choice set or not does not affect the DM's problem. Hence, never-considered alternatives have no impact on what we can learn about preferences.

Under the assumptions of Theorem (ref), identification attains. Notably, each consideration parameter $\aparam_{j}$ is identified (as long as $d_j$ is chosen with positive probability at some $\covx$).

theoremSuppose Assumptions (ref), (ref)-(ref), (ref)-(ref), (ref) hold, and Assumption (ref) holds for $\covxs \subset \mathcal{S}$ s.t. $\forall \covx\in \covxs$ \[ U_{\pparam}(d_1,x)>U_{\pparam}(d_j,\covx) \Rightarrow U_{\pparam}(d_1,\covx)>U_{\pparam}(d_{j+1},\covx), \quad \forall j>1. \] Then $f(\cdot)$ is identified and so are $\aparam_{1}$ and $\aparam_{2}$. In addition, if $\Pr(d = d_j|\covx) \neq 0$ for some $x$, then $\aparam_j$ is identified.

Preference-Dependent Consideration

Returning to our example with three alternatives, recall that we have an additional moment $\Pr(d=d_3|\covx)$. The information it provides allows us to identify some forms of dependence between consideration and preferences, i.e., to relax Assumption (ref). To see how, suppose $d_2$ is always considered. Then, with preference dependence, the choice frequencies become:

align*[align* omitted — 198 chars of source]

The ratio of the derivatives of these two moments yields $\frac{\aparam_1(\pparam)}{\aparam_3(\pparam)}$. More assumptions are required to obtain point identification of the $\aparam_{j}(\pparam)$'s. In Section (ref) we provided identification results for Binary Consideration Types. Here, leveraging the additional structure provided by the ARC model, we can allow for more flexible dependence between consideration and preferences. We do so through a proportionally shifting consideration mechanism, formally defined below. This mechanism may arise when there is a cost to evaluate each alternative. In such a case, the DMs may consider alternatives that they ex-ante deem more aligned with their preferences (e.g., the DM's consideration shifts away from riskier to safer alternatives as her risk aversion increases).

assumptionSP{ARC.P}[ARC with Proportional Consideration] The consideration mechanism follows the $\modA$ model with $\{D\geq 4 ~ \& ~ 1\leq d^*\leq D\}$ or $\{D=3 ~ \& ~ d^*=2\}$, and\[ \aparam_j(\pparam) = \begin{cases} \aparam_j(1 - \alpha(\pparam)) & \text{if } j < d^* \\ 1 & \text{if } j = d^* \\ \aparam_j(1 + \alpha(\pparam)) & \text{if } j > d^* \end{cases} \] s.t. $\alpha(\cdot)$ is differentiable a.e., $\alpha'(\cdot) \neq 0$ a.e., $\alpha(\bar{\pparam}) = 0$, $0<\aparam_j(\pparam)< 1, ~ \forall j \neq d^*$, $\forall\pparam \in [0,\bar \pparam]$.

In the case with three alternatives, $\frac{\aparam_1(\pparam)}{\aparam_3(\pparam)}=\frac{\aparam_1(1 - \alpha(\pparam))}{\aparam_3(1 + \alpha(\pparam))} $. From this, $\frac{\aparam_1}{\aparam_3}$ is identified when $\covx$ and $\covx'$ are chosen such that $\cmap_{1,2}(\covx) = \cmap_{2,3}(\covx') =\bar{\pparam}$. Once $\frac{\aparam_1}{\aparam_3}$ is identified, $\frac{1-\alpha(\pparam)}{1+\alpha(\pparam)}$ is known for all $\pparam$; hence, $\alpha(\pparam)$ can be solved for. Identification of $f(\pparam)$ follows from substituting $\alpha(\pparam)$ into the expression for $\frac{\Pr(d=d_1|x)}{dx}$. The theorem below generalizes this argument.

definition(No Three Way Ties) For a given $\covx$, there are no-three way ties if $\nexists \pparam \in [0,\bar{\pparam}]$ and $\{j,k,l\}$ s.t. $U(d_j,x)=U(d_k,x)=U(d_l,x)$.

\phantomsection

theoremSuppose Assumptions (ref), (ref), (ref)-(ref), (ref) hold, and Assumption (ref) holds for $\covxs$ s.t. $\forall \covx\in \covxs$ there are no three-way ties and \begin{align*} U_{\pparam}(d_1,x)>U_{\pparam}(d_j,x) &\Rightarrow U_{\pparam}(d_1,x)>U_{\pparam}(d_{j+1},x), \quad \forall j>1,\\ U_{\pparam}(d_D,x)>U_{\pparam}(d_j,x) &\Rightarrow U_{\pparam}(d_D,x)>U_{\pparam}(d_{j-1},x), \quad \forall j<D, \end{align*} and $\exists x \in \covxs$ s.t. $\cmap_{j,k}(x)\leq0, ~ \forall j,k, ~j<k$. Then $f(\cdot)$ and $\{\aparam_{j}(\cdot)\}_{j=1}^D$ are identified.

The conditions of the theorem are stronger than in Theorem (ref), as they impose relative order of the cutoffs not only for alternative $d_1$ but also for $d_D$. In many cases, the relative order of $\cmap_{1,j}(x)$'s alone is sufficient, for example when $d_1$ is always considered.

Identification with Alternative-specific Excluded Regressors

By leveraging features of the ARC model, the identification results in Section (ref) can be extended to the case where the consideration of $d_j$ is a function both of $\covx_j$ and preferences. This differs from Corollary (ref), which restricted the consideration of alternative $d_j$ to depend only on the default alternative's excluded regressor. We continue to assume that the choice set is $\{d_1,d_2,d_3\}$ and that $d_2$ is always considered. Let the consideration of $d_j$ be a measurable function of $\covx_j$ and $\nu$, continuous in its first argument: $\aparam_{j} =\aparam_j(\covx_j,\pparam)$. Similar to the example in Section (ref), we construct two vectors, $\covx^i$ and $\covx^{ii}$, such that: (i) $\pparam = \cmap_{1,2}(\covx^i)<\cmap_{1,3}(\covx^i)<\cmap_{2,3}(\covx^i)$; and (ii) $\cmap_{2,3}(\covx^{ii})<\cmap_{1,3}(\covx^{ii})<\cmap_{1,2}(\covx^{ii}) = \pparam$. The derivative of the choice frequency of $d_1$ with respect to $\covx_2$ for these cases are, respectively:

align[align omitted — 425 chars of source]

The ratio of the expressions in Equation (ref) identifies $\aparam_3(\covx_3,\pparam)$. Using a similar logic, we can identify $\aparam_1(\covx_1,\pparam)$. Plugging these consideration probabilities into Equation (ref) identifies $f(\pparam)$. In sum, alternative-specific variation yields identification without large support and without the independence Assumptions (ref) and (ref). It is also possible to allow consideration of $d_1$ (and $d_3$) to depend on $x_1$, $\pparam$, as well as $x_3$. The key exclusion restriction in this case is that the consideration of $d_2$ is independent of all components of $\covx$. Our last identification result generalizes this example.

assumptionSP{ARC.AS} The consideration mechanism follows the $\modA$ model. The consideration probability of each alternative $d_j$ is a measurable function of $\covx_j$ and preferences: $\aparam_{j}= \aparam_{j}(\covx_j,\pparam)$, continuous in the first argument. Default alternative $d^*$ is s.t. $\aparam_{d^*}(\covx_{d^*},\pparam)=1$ for all $\covx_{d^*} \in\mathcal{S}$ and for all $\pparam \in [0,\bar \pparam]$.
theoremSuppose Assumptions (ref), (ref)-(ref), (ref) hold. Suppose there is alternative-specific variation and the choice sets contain at least three alternatives. Suppose for a given value of $\pparam$ there exists an $\covx=(\covx_1,\covx_2,\dots,\covx_D)$, and an open neighborhood around it in $\mathcal{S}$, s.t. \[ U_{\pparam}(d_1,\covx) = U_{\pparam}(d_2,\covx) =\cdots= U_{\pparam}(d_D,\covx). \] Then $f(\pparam)$ and $\{\aparam_{j}(\covx_j,\pparam)\}_{j=1}^{D}$ are identified.

Existing identification results that rely on alternate-specific variation Goeree2008,Abaluck2019,kawaguchi2019designing allow for consideration dependence on its own regressor, but not preferences. Theorem (ref) states identification for a general version of the ARC model where the alternative-specific consideration probability can depend on both its own regressor and DMs' preferences.

Likelihood and Tractability

We now turn to the computational aspects of limited consideration models under the SCP and, in particular, of their likelihood function. Consider a generic consideration mechanism.

A computationally appealing way to write the likelihood function is to determine the probability that a DM with preference parameter $\pparam$ chooses alternative $\dor_j$ conditional on $\covx$. Alternative $\dor_{ j}$ is chosen if and only if $d_j$ is in the consideration set and every alternative that is preferred to $d_j$ is not. Denote the set of alternatives that are preferred to $\dor_{j}$ by

align*[align* omitted — 111 chars of source]

Then,

align[align omitted — 156 chars of source]

The object on the RHS does not require evaluating the utility of each alternative within each possible consideration set. In fact, $U_{\pparam}(d_j,x)$ needs to be computed only once for each $\pparam$, $d_j$, and $\covx$ to create $\Bc_{\pparam}(\dor_j,x)$, which does not vary with the consideration set. Hence the computational complexity lies in the mapping from $\mathcal{O}_{\pparam}^{\covx}(\cdot)$ to the parameters governing the consideration mechanism. This, however, may not even require enumerating all possible consideration sets. To demonstrate this with a concrete example, we proceed with the basic ARC model. In this case, the RHS of Equation (ref) is:

align[align omitted — 119 chars of source]

Given $\{\aparam_j\}_{j=1}^D$, the integrand $\prod_{k\in \Bc_{\pparam}(\dor_j,x)}(1-\aparam_k)$ is piecewise constant in $\pparam$ with at most $D-1$ breakpoints, corresponding to indifference points between alternatives $j$ and $k$, i.e., $\cmap_{j,k}(x)$, that are computed only once for each observed $\covx$. There are at least two methods to compute this integral. First, for every $d_j$ and $\covx$, we can directly compute the breakpoints and hence write $I(\dor_j|\covx)$ as a weighted sum:

align*[align* omitted — 159 chars of source]

where $\pparam_{h}$'s are the sequentially ordered breakpoints augmented by the integration endpoints: $\pparam_{0}=0$ and $\pparam_{D}=\bar{\pparam}$. This expression is trivial to evaluate given $F(\cdot)$ and breakpoints $\{\pparam_{h}\}_{h=0}^{\Cn}$. More importantly, since the breakpoints are invariant with respect to the consideration probabilities, they are computed only once for each $\covx$. This simplifies the likelihood maximization routine by orders of magnitude, as each evaluation of the objective function involves a summation over products with at most $\Cn$ terms. A second approach is to compute $I(\dor_j|\covx)$ using Riemann approximation:

align*[align* omitted — 165 chars of source]

where $M$ is the number of intervals in the approximating sum, $\frac{\bar{\pparam}}{M}$ is the intervals' length, $\pparam_{m}$'s are the intervals' midpoints, and $f(\cdot)$ is the density of $F(\cdot)$. Again, one does not need to evaluate the utility from different alternatives in the likelihood maximization. Instead, one a priori computes the utility rankings for each $\pparam_m$, $m=1,\dots,M$. These rankings determine $\Bc_{\pparam_m}(d_j,x)$. The likelihood maximization is now a standard search routine over $\{\aparam_j\}_{j=1}^{D}$ and $f(\cdot)$. Our theory restricts $f(\cdot)$ to the class of continuous and strictly positive functions. In practice, the search is over a class of non-parametric estimators for $f(\cdot)$.\footnote{One could use a mixture of Beta distributions ghosal2001, as we do in Section (ref).} If the density is parameterized, i.e., $f(\pparam_m)\equiv f(\pparam_m;\theta^f)$, then the maximization is over $\{\aparam_j\}_{j=1}^{D}$ and $\theta^f$. Finally, the interval midpoints are the same across all DMs as they do not depend on $\covx$, further reducing computational burden.\footnote{Depending on the class of $f(\cdot)$, it may be more accurate to compute $I(\dor_j|\covx)$ by substituting $\frac{\bar{\pparam}}{M}f({\pparam_m})$ with $F(\overline{\pparam}_m)-F(\underline{\pparam}_m)$, where $\overline{\pparam}_m$ and $\underline{\pparam}_m$ are the endpoints of the corresponding interval.}

Allowing consideration to depend on preferences (or on $\covx$) introduces only minimal adjustments to the likelihood function. For example, let each consideration function be parameterized by $\theta_j$: $\aparam_{j}(\pparam)\equiv\aparam_{j}(\pparam;\theta_{j})$. Then, at each $\pparam$, we can substitute $\aparam_{j}$ with the corresponding $\aparam_{j}(\pparam;\theta_{j})$, and the likelihood maximization is now over $\{\theta_j\}_{j=1}^{D}$ and $\theta^f$. Given the desired level of parameterization -- i.e., the dimensionality of the parameter vectors $\theta_{j}$ and $\theta^{f}$ -- the computational complexity of the problem grows polynomially in $D$.

As a final remark, if alternative $d_j$ is never chosen, then one can conduct estimation as if $d_j$ were not in the choice set. Indeed, per Equation (ref), $\aparam_j$ contributes positively to the likelihood if and only if alternative $d_j$ is chosen. When it is never chosen, it may only enter via the term $(1-\aparam_j)$; hence, the likelihood will be maximized by setting $\aparam_j=0$. Therefore, setting $\aparam_{j}=0$ for all zero-share alternatives, regardless of why they were not chosen, has no impact on estimation. This too may speed up estimation.

Limited Consideration and RUM: A Comparison

We focus on a standard application of the RUM with full consideration in the context of our example in Section (ref). The final evaluation of the utility that the DM derives from alternative $j$ now includes a separately additive error term:

align[align omitted — 84 chars of source]

where, as before, $\pparam$ captures unobserved heterogeneity in preferences, and $\varepsilon_j$ is assumed independent of the random coefficients (in this application, $\pparam$).

Typical implementations of this model further specify that $\varepsilon_j$ is i.i.d. across alternatives (and DMs) with a Type 1 Extreme Value distribution, following the seminal work of Mcfadden1974. This yields a Mixed Logit that is distinct from the commonly used one in McFaddenTrain2000. In their model, random coefficient(s) enter the utility function linearly, while in the context of expected utility they enter nonlinearly. We now discuss two properties of the Mixed Logit that hinder its applicability in our context.

Monotonicity

Coupling utility functions in the hyperbolic absolute risk aversion (HARA) family, for example CARA or CRRA, with a Type 1 Extreme Value distributed additive error yields:

propositionape:bal18,Wilcox08 In Model (ref) with HARA preferences and $\varepsilon_j$ i.i.d. Type 1 Extreme Value, as the DM's risk aversion increases, the probability that she chooses a riskier alternative declines at first but eventually starts to increase.

To see why, consider two non-dominated alternatives $\dor_j$ and $\dor_k$ such that $\dor_j$ is riskier than $\dor_k$. A risk neutral DM prefers $\dor_j$ to $\dor_k$, and hence will choose the former with higher probability. As risk aversion increases, the DM eventually becomes indifferent between $\dor_j$ and $\dor_k$ and chooses either of these alternatives with equal probability. As risk aversion increases further, she prefers $\dor_k$ to $\dor_j$ and chooses the latter with lower probability. However, as risk aversion gets even larger, the expected utility under HARA of any lottery with finite stakes converges to zero. Consequently, the choice probabilities of all alternatives, regardless of their riskiness, converge to a common value.\footnote{Recall that in the Mixed Logit the magnitude of the utility differences is tied to differences in (log) choice probabilities, $ U_{\pparam}(L_k(x))-U_{\pparam}(L_j(x))=\log(\Pr(d=d_k|x,\pparam))-\log(\Pr(d=d_j|x,\pparam))$, so that as $\nu\to\infty$ the choice probabilities are predicted to be all equal.} Hence, at some point the probability of choosing $\dor_j$ is increasing in risk aversion.

To the contrary, our model with a limited consideration mechanism that is independent of preferences yields choice probabilities that are monotone in the preference parameter.

property[Generalized Preference Monotonicity] A model satisfies generalized preference monotonicity if for any $\pparam_1 < \pparam_2$ and $J \in \{1,2,\dots,\Cn\}$: \begin{align*} \Pr\left(\bigcup_{j=1}^J \dor_j \bigg|\covx,\pparam_1\right) \geq \Pr\left(\bigcup_{j=1}^J \dor_j \bigg|\covx,\pparam_2\right). \end{align*}

In the context of risk preferences, Property (ref) states that the probability of choosing one of the $J$ riskiest alternatives declines as $\pparam$ increases. Since Property (ref) is satisfied for any choice set under the SCP and full consideration, it is also satisfied under limited consideration:

propositionA model that satisfies the SCP (i.e., Assumption (ref)) and Assumption (ref) satisfies Generalized Preference Monotonicity.

Generalized Dominance

Next, we establish the relation between utility differences across two alternatives and their respective choice probabilities. Because our random expected utility model features unobserved preference heterogeneity, we work with an analog of the rank order property in Manski75 that is conditional on $\pparam$:

definition(Conditional Rank Order of Choice Probabilities) The model yields conditional rank order of the choice probabilities if for given $\nu$ and alternatives $j,k\in \cal{D}$, \begin{align*} U_{\pparam}(L_j(x))>U_{\pparam}(L_k(x)) &\Rightarrow \Pr(d=d_j|x,\pparam)>\Pr(d=d_k|x,\pparam). \end{align*}

We show that the conditional rank order property implies the following upper bound on the probability that suboptimal alternatives are chosen.

property(Generalized Dominance) A model satisfies Generalized Dominance if for any $\covx$, $d_j$, and set $\mathcal{K} \subset {\cal{D}}\setminus \{d_j\}$ s.t. alternative $d_j$ is never-the-first-best in $\mathcal{K} \cup \{d_j\}$ \[\Pr(d=d_j|x)<\sum_{k \in \mathcal{K}}\Pr(d=d_{k}|x). \]

Generalized Dominance holds in the Mixed Logit model and, more broadly, in models that satisfy the conditional rank order property. However, it may not hold in some limited consideration models. For example, Generalized Dominance is violated if $d_j$ is never-the-first best among $\{d_j,d_k,d_l\}$, is almost always considered, and alternatives $d_k$ and $d_l$ are rarely considered.

Limited Consideration as Ordinal RUM

In the Mixed Logit, the cardinality of the differences in the (random) expected utility of alternatives plays a crucial role in the determination of choice probabilities, as it interacts with the realization of the additive error. In contrast, in models that satisfy the SCP, the DMs' choices are determined by the ordinal expected utility ranking of the alternatives. Hence, limited consideration models can be recast as Ordinal Random Utility models (ORUM), where the key departure from standard RUMs is the distribution of the additive error term.

proposition(Limited Consideration as ORUM) A Limited Consideration Model is equivalent to an additive error random utility model with unobserved preference heterogeneity where all alternatives are considered, the DM's utility of each alternative $d_j \in \mathcal{D}$ is given by\[ V_{\pparam}(d_j,x)=U_{\pparam}(d_j,x)+\varepsilon_{j}(\pparam,x), \] and $\vec{\varepsilon}(\pparam,x)=\{\varepsilon_{j}(\pparam,x)\}_{j=1}^D$ is distributed on $\{-\infty,0\}^D$ according to $\Pr\left(\vec{\varepsilon}(\pparam,x)\right)= \mathcal{Q}_{\pparam}^{\covx}(\cal{K})$ for $\cal{K}$ s.t. $d_j \in \mathcal{K}$ if $\varepsilon_j(\pparam,x) = 0$ and $d_k \in \mathcal{D} \setminus \mathcal{K}$ if $\varepsilon_j(\pparam,x) = -\infty$.

Casting our limited consideration model as an ORUM clearly demonstrates its flexibility. In particular, our results show how to obtain identification when the errors are correlated with the excluded regressors, the preference parameter, and across alternatives.

table[table omitted — 1,125 chars of source]

We conclude this section with Table (ref), listing the differences across the Mixed Logit and limited consideration models. The first two columns summarize the differences between the basic $\modA$ model and the Mixed Logit. The third column and fourth column remind the reader our two models with consideration depending on preferences. Finally, the last column highlights the fact that with alternative-specific variation we may also have dependence of the error term on the excluded regressor(s) as well as on the preference parameter.

Application

We offer an empirical analysis of households' decisions under risk. This analysis aims to illustrate how our method works and its ability to fit the data.

Data

We study households' deductible choices across three lines of property insurance: auto collision, auto comprehensive, and home all perils. The data come from a U.S. insurance company. Our analysis uses a sample of 7,736 households who purchased their auto and home policies for the first time between 2003 and 2007 and within six months of each other.\footnote{The dataset is an updated version of the one used in Barseghyan2013. It contains information for an additional year of data and puts stricter restrictions on the timing of purchases across different lines. These restrictions are meant to minimize potential biases stemming from non-active choices, such as policy renewals, and temporal changes in socioeconomic conditions.} Table (ref) provides descriptive statistics for households' observable characteristics, which we use later to estimate households' preferences.\footnote{These are the same variables that are used in Barseghyan2013 to control for households' characteristics. See discussion there for additional details.} We observe the exact menu of alternatives available at the time of the purchase for each household and each line of coverage. The deductible alternatives vary across lines of coverage but not across households. Table (ref) presents the frequency of chosen deductibles in our data.

table[table omitted — 563 chars of source]

Premiums are set coverage-by-coverage as in the example from Section (ref). Table (ref) reports the average premium by context and deductible, and Table (ref) summarizes the premium distributions for the $\$500$ deductible. Premiums vary dramatically. The 99th percentile of the $\$500$ deductible is more than ten times the corresponding 1st percentile in each line of coverage.

Claim probabilities stem from barseghyan2018different, who derived them using coverage-by-coverage Poisson-Gamma Bayesian credibility models applied to a large auxiliary panel. Predicted claim probabilities (summarized in Table (ref)) exhibit extreme variation: The 99th percentile claim probability in collision (comprehensive and home) is 4.3 (12 and 7.6) times higher than the corresponding 1st percentile. Finally, the correlation between claim probabilities and premiums for the $\$500$ deductible is 0.38 for collision, 0.15 for comprehensive, and 0.11 for home all perils. Hence, there is independent variation in both.

normalsize\begin{table} \scriptsize \begin{threeparttable} \caption{Claim Probabilities Across Contexts} \begin{tabular}{lc c c c c c c} \hline \hline \\[-1.5ex] Quantiles & 0.01 & 0.05 & 0.25 & 0.50 & 0.75 & 0.95 & 0.99 \\\\[-1.5ex] \hline \\[-1.5ex] Collision & 0.036 & 0.045 & 0.062 & 0.077 & 0.096 & 0.128 & 0.156 \\\\[-1.5ex] Comprehensive & 0.005 & 0.008 & 0.014 & 0.021 & 0.030 & 0.045 & 0.062 \\\\[-1.5ex] Home & 0.024 & 0.032 & 0.048 &0.064 & 0.084 &0.130 &0.183 \\\\[-1.5ex] \hline \end{tabular} \end{threeparttable} \end{table}

Estimation Results

The basic $\modA$ Model: Collision

We start by presenting estimation results in a simple setting where the only choice is the collision deductible and observable demographics do not affect preferences. To execute our estimation procedure we set $\bar{\pparam}=0.02$, which is conservative Barseghyan2016. We ex post verify that this does not affect our estimation by checking that the density of the estimated distribution is close to zero at the upper bound. We approximate $F(\cdot)$ non-parametrically through a mixture of Beta distributions. In practice, however, both AIC/BIC criteria indicate that a single component is sufficient for our analysis, resulting in a total of seven parameters to be estimated. We let the data speak to the identity of the always-considered alternative.\footnote{In fact, the estimation is run under the Coin Toss completion rule that nests the possibility that any alternative can be always considered. The data chooses $\aparam_{1000} = 1$.}

The estimated distribution and consideration parameters are reported in Table (ref). As the first panel in Figure (ref) shows, the model closely matches the aggregate moments observed in the data. The second panel in Figure (ref) illustrates side-by-side the frequency of predicted choices, consideration probabilities, and the distribution of households' first-best alternatives (i.e., the distribution of optimal choices under full consideration). Predicted choices are determined jointly by the preference induced ranking of deductibles and by the consideration probabilities: Limited consideration forces households' decision towards less desirable outcomes by stochastically eliminating better alternatives. The two highest deductibles ($\$1000$ and $\$500$) are considered at much higher frequency (1.00 and 0.92, respectively) than the other alternatives, suggesting that households have a tendency to regularly pay attention to the cheaper items in the choice set. Yet, the most frequent model-implied optimal choice under full consideration is the $\$250$ deductible, which is considered with low probability.

figure[figure omitted — 438 chars of source]

In this application, assuming full consideration leads to a significant downward bias in the estimation of the underlying risk preferences. To see why, consider increasing the consideration probabilities for the lower deductibles to the same levels as the $\$500$ deductible. Holding risk preferences fixed, the likelihood that the lower deductibles are chosen increases and therefore the higher deductibles are chosen with lower probability. Average risk aversion must decline to compensate for this shift. This is exactly the pattern we find when we estimate a near-full consideration model. In particular, we find that average risk aversion decreases by about 32% from 0.0037 to 0.0025 when all consideration parameters equal 0.9999.\footnote{We cannot assume that all consideration probabilities are equal to one, since the $\$200$ deductible is never the first best under full consideration and is chosen with positive probability.} To put these numbers into context, a DM with risk aversion equal to 0.0037 is willing to pay $\$431$ to avoid a $\$1000$ loss with probability 0.1, while a DM with risk aversion equal to 0.0025 is only willing to pay $\$300$ to avoid the loss.

The basic $\modA$ model's ability to match the data extends also to conditional moments. The first two panels of Figure (ref) show observed and predicted choices for the fraction of households facing low and high premiums, respectively, and the next two panels are for households facing low and high claim probabilities.\footnote{Low and high groups here are defined as households whose claim rate (or baseline price) are in the first and third terciles, respectively.} Finally, the last two panels display households who face both low claim probabilities and high prices and vice versa. It is transparent from Figure (ref) that the model matches closely the observed frequency of choices across different subgroups of households facing a variety of prices and claim probabilities, even though some of these frequencies are quite different from the aggregate ones.

figure[figure omitted — 191 chars of source]

The $\modA$ model's ability to violate Generalized Dominance is key in matching the data. In our dataset, because of the pricing schedule in collision, the $\$200$ is never-the-first best among $\{\$100,\$200,\$250\}$ for $99.84\%$ of all households and $100\%$ of households who have chosen the $\$200$ deductible. It costs the same to get an additional $\$50$ of coverage by lowering the deductible from $\$250$ to $\$200$ as it does to get an additional $\$100$ of coverage by lowering the deductible from $\$200$ to $\$100$. If a household's risk aversion is sufficiently small, then it prefers the $\$250$ deductible to the $\$200$ deductible. If, on the other hand, the household's level of risk aversion is such that it would prefer the $\$200$ deductible to the $\$250$ deductible, then it would also prefer getting twice the coverage for the same increase in the premium. That is, for any level of risk aversion, the $\$200$ deductible is dominated either by the $\$100$ deductible or by the $\$250$ deductible.\footnote{This pattern is at odds not only with EUT but also many non-EU models Barseghyan2016.} Yet, overall the $\$200$ deductible is chosen roughly as often as the $\$100$ and $\$250$ deductibles combined. More so, for certain sub-groups the $\$200$ deductible is chosen much more often than the $\$100$ and $\$250$ deductible combined. It follows that a model satisfying Generalized Dominance cannot rationalize these choices.

Next we relax the assumption that demographic variables, $\mathbf{Z}$, do not influence risk preferences. In particular, conditional on demographics, preferences are distributed Beta($\beta_{1}(\mathbf{Z}),\beta_{2}$), where $\log\frac{\beta_{1}(\mathbf{Z})}{\beta_{2}}= \mathbf{Z}\gamma$, yielding a conditional mean preference value $ E(\pparam|\mathbf{Z})=\frac{e^{\mathbf{Z}\gamma}}{1+e^{\mathbf{Z} \gamma}}\bar{\pparam} $. The details of this step and the results are reported in Appendix (ref). Both consideration and preference estimates remain close to those reported above.

Proportionally Shifting Consideration

We estimate the model with proportionally shifting consideration where the preference distribution and function $\alpha(\pparam)$ may depend on demographic variables (Table (ref)). Motivated by the findings of the previous section, we assume that the cheapest/riskiest alternative is always considered. The consideration probability of the remaining alternatives is equal to $\aparam_j(\covx,\pparam)= \aparam_j(1-\alpha(\pparam|\mathbf{Z}))$, where $\alpha(\pparam|\mathbf{Z}) = \xi_1(\mathbf{Z})\left(1-\frac{\pparam}{\bar{\pparam}}\right)^{\xi_2}$, $\xi_1(\mathbf{Z})=\frac{e^{\mathbf{Z}\rho}}{1+e^{\mathbf{Z}\rho}}$, and $\xi_2$ is positive. We continue to assume that preferences are distributed Beta($\beta_{1}(\mathbf{Z}),\beta_{2}$).

The estimated average value of $\xi_1(\mathbf{Z})$ is $0.15$, with the $95\%$ CI of $[0.02, 0.21]$. When $\xi_1(\mathbf{Z}) = 0.15$, a risk-neutral DM considers each of the safer alternatives $\{\$100, \$200, \$250, \$500\}$ $15\%$ less often than does an extremely risk averse DM. The estimated value of $\xi_2$ is $7.14$. This implies that as the risk aversion parameter increases from 0 to its estimated average median value of $0.0034$, consideration probability of the safer alternatives increases by $11\%$, and it is essentially flat after that rising by an additional $4\%$ as risk aversion reaches its upper bound.

The Mixed Logit Random Utility Model

As in the case of the $\modA$ model, we assume that $\pparam$ is Beta distributed on $[0,\bar{\pparam}]$, where $\bar{\pparam}=0.02$. The Mixed Logit satisfies the Generalized Dominance and smoothly spreads households' choices around their respective first bests. Consequently, it cannot match the observed distribution and, in particular, is unable to explain the relatively high observed share of the $\$200$ deductible. Table (ref) reports the estimation results and Figure (ref) compares the observed distribution of choices to the predicted choices. The predicted distribution is a much poorer fit relative to the $\modA$ model. In fact, the vuong1989likelihood test soundly rejects (at 1$\%$ level) the Mixed Logit in favor of the $\modA$ model.

The $\modA$ Model: All Coverages

We now proceed with estimation of the full model. We assume that households' consideration sets are formed over the entire deductible portfolio. There are 120 possible alternative triplets $(d^{coll},d^{comp},d^{home})$, each having its own probability of being considered. This model is flexible as it nests many rule of thumb assumptions such as only considering contracts with the same deductible level across the three contexts or only considering contracts with a larger collision deductible than comprehensive deductible. Figure (ref) and Table (ref) present estimation results. The first panel of the figure shows the predicted distribution of choices across triplets, ranked in descending order by observed frequencies. The second panel plots the differences between predicted and observed choice distributions. Clearly, the predicted distribution is close to the observed distribution.

figure[figure omitted — 596 chars of source]

The largest difference between the predicted and observed shares equals $0.38$ percentage points, which is for the ($\$250,\$250,\$500$) triplet that is chosen by $2.9\%$ of the households. The integrated absolute error across all triplets is $3.46\%$. In our data, 43 out of 120 triplets are never chosen (these are omitted from Figure (ref)). As discussed in Section (ref), the likelihood maximization implies that the consideration probabilities for these triplets must be zero, so their predicted shares are zero. Hence, the likelihood maximization routine is faster and more reliable as we do not need to search for $\aparam_j$ for these alternatives.

Another virtue of the $\modA$ model is that it effortlessly reconciles two sides of the debate on stability of risk preferences Barseghyan2011, Einav2012, Barseghyan2016. On the one hand, households' risk aversion relative to their peers is correlated across lines of coverage, implying that households preferences have a stable component. On the other hand, analyses based on revealed preference reject the standard models: under full consideration, for the vast majority of households one cannot find a level of (household-specific) risk aversion that justifies their choices simultaneously across all contexts. Limited consideration allows the model to match the observed joint distribution of choices, and hence their rank correlations.

The estimated risk preferences are similar to those estimated with collision only data, although the variance is slightly smaller. The triplet considered most frequently is the cheapest one: ($\$1000,\$1000,\$1000$). Its consideration probability is 0.76, while the next two most considered triplets are (\$500, \$500, \$1000) and (\$500, \$500, \$500). These are considered with probability 0.47 and 0.43, respectively. Overall, there is a strong positive correlation $(0.53)$ between the consideration probability and the sum of the deductibles in a given alternative.

We summarize once more the computational advantages of our procedure. First, estimation of our model remains feasible for a large choice set.\footnote{In our setting, it is feasible to estimate an additive error RUM assuming the DMs consider each deductible triplet as a separate alternative (Figure (ref) and Table (ref)). As the figure shows, the failure to match the data is evident. The Vuong test formally rejects it in favor of the ARC model.} Second, the model's parameters grow linearly with the size of the choice set -- one parameter per an additional alternative. Third, enlarging the choice set does not call for new independent sources of data variation. For example, in our model whether there are five deductible alternatives or one hundred twenty does not make any difference either from an identification or an estimation stand point: with sufficient variation in $\bar{p}$ and/or $\mu$, the model is identified and can be estimated. As a final remark, once the model is estimated, one can compute the average monetary cost of limited consideration. In our data it is $\$50$ (see Appendix (ref)).

Discussion

The literature concerned with the formulation, identification, and estimation of discrete choice models with limited consideration is vast. However, to our knowledge, there is no previous work applying such models to the study of decision making under risk, except for the contemporaneous work of Barseghyan2019. In particular, this paper is the first to exploit the SCP for identification purposes. As a result, several fundamental differences emerge between our work and existing papers. First, we achieve identification in the most challenging case where there is a single excluded regressor that affects the utility of all alternatives.\footnote{This setting is common in insurance markets, see, e.g.,Cohen2007, Einav2012, Sydnor2010, Barseghyan2011, Barseghyan2013, handel2013adverse, Bhargava2017.} Second, we allow for consideration to depend on preferences. Third, with alternative-specific excluded regressors, this dependence can be essentially unrestricted and can be combined with dependence of consideration on (some of) the excluded regressors. Fourth, we scrutinize the large support assumption, show why it may be necessary, and when and how it is possible to make progress when it is not satisfied. Fifth, our approach comes with an easy to implement and computationally fast estimation strategy. Finally, we make a contribution specific to the study of decision making under risk by proposing a model that is immune from ape:bal18 criticism and features two sources of unobserved heterogeneity -- risk aversion and limited consideration -- whose distributions are identified. More generally, the paper establishes that, as long as the DMs' preferences satisfy the SCP, allowing for limited consideration does not hinder the model's identifiability or applicability. Hence, we view our framework as a stepping stone for studies of consumer behavior in markets where limited consideration may be present coughlin2019.

Papers that allow for limited consideration or more broadly for choice set heterogeneity can be classified in four groups. The first relies on auxiliary information about the composition or distribution of DMs' choice sets, such as brand awareness Draganska2011, honka2017simultaneous or search activity honka2017simultaneous, DelosSantos2012, Kim2010, HonkaRAND2017.\footnote{For canonical cites see, e.g., Roberts1991 and Ben-Akiva1995.} We do not require such information.

The second group attains identification via two-way exclusion restrictions, i.e., by assuming that some variables impact consideration but not utility and vice versa. A well-known example of this approach is Goeree2008, who posits that advertising intensity affects the likelihood of considering a computer, but does not impact consumer preferences, while computer attributes such as CPU speed affect preferences but not consideration (see also vanNierop2010 and Gaynor2016). Hortacsu2017 create an exclusion restriction by exploiting the dynamic aspect of consumer choice.{\footnote{Time variation is used also in Crawford2019, who show that with panel data and preferences in the logit family, point identification of preferences is possible, without any exclusion restrictions, under the assumption that choice sets and preferences are independent conditional on observables and with restrictions on how choice sets evolve over time. These restrictions enable the construction of proper subsets of DMs' true choice sets (`sufficient sets') that can be utilized to estimate the preference model.} The consumer's decision to consider alternatives to her current service provider is a function of (her experiences with) the last period provider but not her next period provider (see also heiss2016inattention). In contrast, we achieve identification with as little as one common excluded regressor and a single cross section.

The third group relies on restricting the consideration mechanism to a specific class of models. Abaluck2019 consider two such models (and their hybrid): a variant of the ARC and a “default specific” model ho2017impact,heiss2016inattention in which each DM's consideration set comprises either a single default alternative or the entire feasible set. They assume that consideration and preferences are independent, and that each alternative has a characteristic with large support that is additively separable in utility and may only affect its own consideration but not the consideration of other alternatives.\footnote{The exception is the “default” alternative, whose characteristic may trigger the consideration of the entire choice set.} They exploit violations of symmetry in the Slutsky matrix (i.e., in cross-alternative demand responses to prices) to detect limited consideration. kawaguchi2019designing study beverage purchases from vending machines, allowing advertisement to be a driver of consideration, but also to affect utility. Their approach is close to that of Goeree2008, though they provide a formal argument for identification with large support and exclusion restrictions even when there is no choice set variation. A key assumption is that all beverages are considered with probability equal to one as the advertising intensity of each beverage becomes very large.

The methods we propose relate to the papers in the third group in two aspects. First, we too sometimes require large support as a “fail safe” assumption, but only in the most challenging case of a single common excluded regressor. Second, we too rely on exclusion restrictions. The reliance on these assumptions is inescapable given the econometrics literature on point identification of discrete choice models. Our approach elucidates the identifying power of a single excluded regressor in models that satisfy the SCP and, in particular, the relative ranking of alternatives encapsulated in Facts (ref) and (ref) lewbel2016identifying. We further exploit this structure to establish identification in models with substantially richer levels of unobserved heterogeneity, by allowing for dependence between consideration and preferences.

The fourth group of papers has a different goal than what we pursue here, as it provides partial rather than point identification results. Cattaneo2019 propose a random attention model with homogeneous preferences, and they require that the probability of each consideration set is monotone in the number of alternatives in the choice problem. Their analysis yields testable implications and partial identification for preference orderings. Barseghyan2019 study discrete choice models, where consideration may arbitrarily depend on preferences as well as on all observed characteristics. They show that such unrestricted forms of heterogeneity generally yield partial, but not point, identification of the preference distribution and obtain bounds on the distribution of consideration sets' size. Finally, Dardanoni2018 consider a stochastic choice model with homogeneous preferences and heterogeneous cognitive types. They show how one can learn the moments of the distribution of cognitive types from a single cross section of aggregate choice shares.

\linespread{1.22}

flushleft

\linespread{1.3} \numberwithin{equation}{section}