EconBase
← Back to paper

Risk Preference Types, Limited Consideration, and Welfare

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,784 characters · 17 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Risk Preference Types, Limited Consideration, and Welfare

\setstretch{1.2}

titlepage\begin{abstract} We provide sufficient conditions for semi-nonparametric point identification of a mixture model of decision making under risk, when agents make choices in multiple lines of insurance coverage (contexts) by purchasing a bundle. As a first departure from the related literature, the model allows for two preference types. In the first one, agents behave according to standard expected utility theory with CARA Bernoulli utility function, with an agent-specific coefficient of absolute risk aversion whose distribution is left completely unspecified. In the other, agents behave according to the dual theory of choice under risk Yaari1987 combined with a one-parameter family distortion function, where the parameter is agent-specific and is drawn from a distribution that is left completely unspecified. Within each preference type, the model allows for unobserved heterogeneity in consideration sets, where the latter form at the bundle level -- a second departure from the related literature. Our point identification result rests on observing sufficient variation in covariates across contexts, without requiring any independent variation across alternatives within a single context. We estimate the model on data on households' deductible choices in two lines of property insurance, and use the results to assess the welfare implications of a hypothetical market intervention where the two lines of insurance are combined into a single one. We study the role of limited consideration in mediating the welfare effects of such intervention. \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

Introduction

This paper is concerned with providing sufficient conditions for semi-nonparametric point identification of risk preferences from observation of agents' choices in property insurance markets, and with assessing the welfare impact of policy interventions in these markets. Property insurance includes a collection of lines of coverage (e.g., for automobiles: collision, comprehensive, liability, etc.), and we refer to each of them as a context. Within each context, a finite set of alternatives is offered for purchase. Researchers frequently observe agents choosing (at the same time) one alternative in each context, hence choosing a bundle. We assume that agents choose bundles based on preferences that are stable across contexts (i.e., a single agent-specific parameterization of the model governs that agent's choices in each context).\footnote{This assumption is sometimes viewed as an aspect of rationality Kahneman2003, and is credible in our empirical study of demand in very similar contexts (collision and comprehensive deductible insurance).} Our model allows for unobserved heterogeneity in preference types, with some agents behaving according to expected utility theory with CARA Bernoulli utility function (EU types), and others behaving according to the dual theory of choice under risk Yaari1987 combined with a one-parameter family distortion function (DT types). The coefficient of risk aversion of the EU types, and the parameter of the distortion function of the DT types, are random coefficients with unknown distribution functions that are left completely unspecified. The model also allows for unobserved heterogeneity in the bundles that agents consider before making a choice (their consideration set). In particular, whether an alternative offered in one context is considered can depend in unrestricted ways on whether another alternative offered in a distinct context is also considered.

Such rich unobserved heterogeneity makes identification analysis challenging. The multiple preference types, random coefficients within type, and agent-specific consideration sets, contribute three layers to a mixtures problem that we need to disentangle. Moreover, because we allow the consideration sets to form at the bundle level, even within a single preference type the choice problem does not inherit the standard single crossing property of mirrlees1971exploration and spence1974market, central to important studies of decision making under risk apesteguia2017single,chiappori2019aggregate, that BaMoTh21 show plays a key role in allowing for semi-nonparametric point identification of single preference type models. Our main methodological contribution amounts to showing how to resolve each of these challenges. In doing so, we also confront the fact that due to the structure of insurance markets and of data resulting from a single insurance company, while the covariates $\mathbf{x}$ characterizing products in each context do exhibit independent variation across contexts, they do not exhibit independent variation across alternatives within a context.\footnote{Within a single insurance company, typically in a given context if an agent faces a larger price than another agent for one alternative, the first agent faces a (proportionally) larger price for all other alternatives.}

One may wonder whether some aspects of unobserved heterogeneity that are present in our model could be dispensed with, thereby simplifying the identification problem. We argue that this is not the case, both in our empirical application and more broadly. A large literature in experimental economics documents that while some people exhibit behavior consistent with standard EU theory, others exhibit behavior that systematically deviates from it Starmer2000. And it reports substantial heterogeneity in risk preferences within type; see, e.g., Choi2007 and references therein. Moreover, people routinely make or stick to sub-optimal choices Handel2013,Bhargava2015,BMT16, or make choices across contexts that imply incompatible levels of risk aversion Barseghyan2011,Einav2012. The traditional additive error random utility model (Luce-McFadden model), or a “trembling hand" alternative Wilcox08 that is sometimes used to study insurance demand, often do not remedy the problem, as the model implied choice probabilities can be incompatible with their empirical counterpart. These incompatibilities are not specific to a particular utility model, but to an entire class of models that satisfy properties that are typically viewed as desirable.\footnote{See BaMoTh21 for a formal discussion and Section (ref) below for further details.} On the other hand, models of decision making under risk with limited consideration can rationalize agents' choices.

We illustrate the relevance of the rich unobserved heterogeneity that we allow for, by estimating risk preferences from data on household's choices in two contexts, auto collision and auto comprehensive. While currently U.S. property insurance companies offer these two lines of coverage as two separate products, we investigate the implications of offering a combined auto insurance product at a price that equals the sum of the prices for the two separate coverages. Such pricing arises if firms operate under perfect competition or if they use a constant markup rule. This counterfactual exercise is of substantive interest as combined lines of coverage already exist elsewhere (e.g., in Israel; see Cohen2007) and even in the U.S. auto insurance industry. For example, property damage and bodily injury coverage can be offered both as separate lines of coverage, as well as combined in the form of single limit liability coverage. The exercise has the virtue of illustrating the potentially different predictions of the EU model and of the DT model, as we explain in Section (ref), and the extent to which these predictions interact with whether consideration increases or decreases after the intervention. Moreover, the exercise informs the debate on the need to simplify insurance choice, and it clarifies how limited consideration interacts with the behavioral responses associated with this type of market intervention.

The rest of the paper is organized as follows. Section (ref) lays out the model, using our application as motivating example. Section (ref) presents our sufficient conditions for its semi-nonparametric point identification. Section (ref) describes our empirical model and the data. Section (ref) reports the results of our estimation exercise. Section (ref) reports the results of the welfare exercise. Section (ref) concludes by contextualizing our work in the broader literature.

Discrete Choice Under Risk in Multiple Contexts

Our starting point is the random utility model in McFadden1974, applied to study choices over risky alternatives with monetary outcomes. We further adapt the model to analyze the behavior of agents who make choices under risk in multiple distinct contexts.

Lotteries as objects of choice in property insurance

We use our empirical application as motivating example for the discrete choice framework that we analyze. We study deductible choices in two contexts: auto collision (context $\texttt{I}$) and auto comprehensive (context $\texttt{II}$). In each context $j=\texttt{I},\texttt{II}$, we assume that there are two states of the world: one that has probability $\mu_i^j$, where an accident happens and agent $i$ faces a loss; and the other that has probability $1-\mu_i^j$, where no accident happens. Auto collision coverage can be used to insure against loss in context $\texttt{I}$: it pays for damage in excess of the deductible to the insured vehicle caused by a collision with another vehicle or object, without regard to fault. Auto comprehensive coverage can be used to insure against loss in context $\texttt{II}$: it pays for damage in excess of the deductible to the insured vehicle from all other causes (e.g., theft, fire, flood, windstorm, or vandalism), without regard to fault. In each context, a finite set $\mathcal{D}^j$ of alternatives (insurance contracts) is offered.

Conditional on risk level, i.e., given $\mu_i^j$, each alternative $\ell \in \mathcal{D}^j$ is fully characterized by the pair $(\texttt{d}^{\ell j},\mathbf{x}_i^{\ell j})$. The first element is the insurance deductible, which is the agent's out of pocket expense if a loss occurs. All deductibles are assumed to be less than the lowest realization of the loss and $\texttt{d}^{1j}>\texttt{d}^{2j}>\dots>\texttt{d}^{M^j j}$, with $M^j$ the total number of deductibles in context $j$. In collision, $M^\texttt{I}=5$ and $\texttt{d}^\texttt{I}\in\{\$1000,\$500,\$250,\$200,\$100\}$; in comprehensive, $M^\texttt{II}=6$ and $\texttt{d}^\texttt{II}\in\{\$1000,\$500,\$250,\$200,\$100,\$50\}$, for a total of $30$ bundles of offered coverages in $\mathcal{D}=\mathcal{D}^\texttt{I}\times\mathcal{D}^\texttt{II}$.

The second element in $(\texttt{d}^{\ell j},\mathbf{x}_i^{\ell j})$ is the price (insurance premium), and varies across agents. It is important to understand the sources of such variation, because to obtain our point identification result we assume that premiums are exogenous to preferences (Assumption (ref) below) and exhibit substantial variation across households (Assumptions (ref)-(ref) below).

First, we note that an insurance company's rating plan is subject to state regulation and oversight. In particular, the regulations require that a company receive prior approval of its rating plan by the state insurance commissioner, and they prohibit the company and its agents from charging rates that depart from the plan.

Second, we describe the procedure applied by the company from which we obtained our data to rate a policy in each line of coverage.\footnote{See Section (ref) below for additional information on the data.} Under the plan, within each context $j$ the company determines a household's base price $\mathbf{x}_i^j$ according to a coverage-specific rating function, which takes into account agent $i$'s coverage-relevant characteristics and any applicable discounts. Using the base price, the company then generates the agent's pricing menu $\mathcal{M}^j\equiv\{(\texttt{d}^{\ell j},\mathbf{x}_i^{\ell j}):\ell\in\mathcal{D}^j\}$, which associates a premium $\mathbf{x}_i^{\ell j}$ with each deductible $\texttt{d}^{\ell j}$ in the coverage-specific set of alternatives in $\mathcal{D}^j$, according to an agent-invariant and coverage-specific multiplication rule, $\mathbf{x}_i^{\ell j}=(g^{\ell j}\cdot\mathbf{x}_i^j)+\delta^j$, where $\delta^j>0$ and $g^{\ell j}$ is increasing in $\ell$ and strictly greater than zero, so that $\mathbf{x}_i^{1j}<\mathbf{x}_i^{2j}<\dots<\mathbf{x}_i^{M^j j}$ ($\texttt{d}^{\ell j}$ is decreasing in $\ell$, so lower deductibles provide more coverage and cost more).\footnote{The multiplicative factors $\{g^{\ell j}:\ell\in\mathcal{D}^j\}$ are known as the deductible factors and $\delta^j$ is a small markup known as the expense fee.} As $\{g^{\ell j}:\ell\in\mathcal{D}^j;\delta^j\}$ are agent-invariant, there is no independent variation in covariates across alternatives within a context.

With this as background, for given $\mu_i^j$, alternatives can be represented as lotteries:

align[align omitted — 185 chars of source]

where $(\mathbf{x}_i^j,\mu_i^j)$ is observed by the researcher for each agent $i$ and context $j$. Throughout, we implicitly condition on $\mu_i^j$. We do not use variation in $\mu_i^j$ to establish our identification results, although doing so is potentially useful and the subject of ongoing research.

Preference types with unobserved heterogeneity within type

We allow the population of agents to be a mixture of preference types.\footnote{Multiple preference types are a focus of the literature that estimates risk preferences using experimental data (e.g., Bruhin2010,Conte2011; Harrison2010), although preferences are homogeneous within each type, at most conditioning on some observed demographic characteristics.} The literature has put forward many models of decision making under risk which can generate demand for insurance at actuarially unfair prices, including the workhorse expected utility theory model and a host of non-expected utility theory models. Each of these models has relative (de)merits in rationalizing observed choices, and may deliver different predictions for counterfactual policies Barseghyan2018. We hence think it important to provide identification results for a model where multiple preference types are allowed for, and where unobserved heterogeneity within type is also present.

For notational simplicity, we detail here the case with two preference types. The results extend to more than two types (even when one observes choices only in two contexts). Let each agent $i$ draw a preference type $t_i$ as follows:

align[align omitted — 158 chars of source]

with $\alpha\in(0,1)$ the unknown mixing probability.

Each realization of $t_i$ is associated with a family of utility functions with distinct functional forms, denoted $\mathcal{U}^1=\{U_{\nu},~\nu\in [0,\bar \nu]\}$ for $t_i=1$, and $\mathcal{U}^0=\{U_{\omega},~\omega\in[0,\bar \omega]\}$ for $t_i=0$. Functions in each family are known up to a scalar random coefficient that depends on type, denoted $\nu_i$ (with support $[0,\bar \nu]$) for agents with $t_i=1$, and $\omega_i$ (with support $[0,\bar \omega]$) for agents with $t_i=0$. For example, in our empirical application $\mathcal{U}^1$ is the collection of expected utility functions associated with preferences that exhibit constant absolute risk aversion (CARA) with agent-specific Arrow-Pratt coefficient $\nu_i$,\footnote{Other preferences that are characterized by a scalar parameter include ones exhibiting constant relative risk aversion (CRRA), or negligible third derivative Cohen2007,Barseghyan2013. Under CRRA, it is required that agents' initial wealth is known to the researcher. } and $\mathcal{U}^0$ is a family of non-expected utility functions that do not nest expected utility as a special case and are parametrized by $\omega_i$ (see Eqs. (ref)-(ref) and Assumptions (ref) & (ref) below). As the preference types are distinct, each agent either receives a draw of $\nu_i$ or a draw of $\omega_i$, hence by construction the two random coefficients are independent. We do not impose any parametric restrictions on their distributions. Rather, in Section (ref) we provide nonparametric point identification results for the two marginal distributions of preferences and for the share of each type.

assumption[Restrictions on distribution of random coefficients] The random coefficient $\nu_i$ (respectively, $\omega_i$) is distributed according to a cumulative distribution function $F$ (respectively, $G$) that satisfies the properties of CDFs, and admits a density function $f$ that is continuous and strictly positive on $\mathcal{V}\equiv[0,\bar \nu]$ (respectively, $g$ strictly positive on $\mathcal{W}\equiv[0,\bar \omega]$). Both $\nu_i$ and $\omega_i$ are independent of $\mathbf{x}_i$.\footnote{Recall that our analysis conditions on $\mu_i^j$, hence the distribution of preferences may depend on it.}

We make three fundamental assumptions about utility functions in both families. First, we assume that households' preferences are stable across contexts, which allows us to leverage variation in observed choices and covariates across contexts (recall that we have no covariate variation within each context).

assumption[Stability] The utility function $U_{\nu_i}$ of each agent $i$ with $t_i=1$ (respectively, $U_{\omega_i}$ for agents with $t_i=0$) is context-invariant.

Second, we need to take a stand on how agents make choices in multiple contexts. To this end, it is important to introduce notation for bundles of alternatives. Denote bundles as $\mathcal{I}_{\ell,q}$, where the first index refers to the alternative in context $\texttt{I}$ and the second one to that in context $\texttt{II}$. Let $CE_{\nu_i}(\mathcal{L}(\texttt{d}^{\ell j},\mathbf{x}^j))$ (respectively, $CE_{\omega_i}(\mathcal{L}(\texttt{d}^{\ell j},\mathbf{x}^j))$) denote the certainty equivalent of lottery $\mathcal{L}(\texttt{d}^{\ell j},\mathbf{x}^j)$ mas:whi:gre95 in context $j$ for an agent of type $t_i=1$ (respectively, $t_i=0$). We impose a standard, albeit sometimes implicit, assumption in the literature,\footnote{All papers that estimate risk preferences in the field as reviewed in Barseghyan2018 impose it. } according to which agents' choices are made without taking into account any background risk Read1999.

assumption[Narrow Bracketing] Agent $i$'s certainty equivalent for the lottery associated with bundle $\mathcal{I}_{\ell,q}$ is equal to $CE_{\zeta_i}(\mathcal{L}(\texttt{d}^{\ell \texttt{I}},\mathbf{x}^\texttt{I}))+CE_{\zeta_i}(\mathcal{L}(\texttt{d}^{q \texttt{II}},\mathbf{x}^\texttt{II}))$, with $\zeta_i=\nu_i$ if $t_i=1$ and $\zeta_i=\omega_i$ if $t_i=0$.

Third, we assume that each preference type satisfies the classic Single Crossing Property (SCP) of mirrlees1971exploration and spence1974market, central to important studies of decision making under risk apesteguia2017single,chiappori2019aggregate.\footnote{The SCP is satisfied in many contexts, ranging from single agent models with goods that can be unambiguously ordered based on quality, to multiple agents models athey2001single.} Formally,

assumption[Single Crossing Property] For a given context $j$ and any two lotteries $\mathcal{L}(\texttt{d}^{\ell j},\mathbf{x})$ and $\mathcal{L}(\texttt{d}^{k j},\mathbf{x})$, $\ell<k$, there exists a continuously differentiable and strictly monotone function $\mathcal{Z}_k^\ell: \operatorname{supp}(\mathbf{x}) \to \mathbb{R}_{[-\infty,\infty]}$, with $\operatorname{supp}(\mathbf{x})=\bigcup_{j=\texttt{I},\texttt{II}}\operatorname{supp}(\mathbf{x}^j)$, such that \begin{align*} U_\zeta(\mathcal{L}(d^{k j},\mathbf{x})) &< U_\zeta(\mathcal{L}(d^{\ell j},\mathbf{x})) \quad \forall \zeta \in (-\infty,\mathcal{Z}_k^\ell(\mathbf{x})), \\ U_\zeta(\mathcal{L}(d^{k j},\mathbf{x})) &= U_\zeta(\mathcal{L}(d^{\ell j},\mathbf{x})) \quad \zeta = \mathcal{Z}_k^\ell(\mathbf{x}), \\ U_\zeta(\mathcal{L}(d^{k j},\mathbf{x}))&> U_\zeta(\mathcal{L}(d^{\ell j},\mathbf{x})) \quad \forall \zeta \in (\mathcal{Z}_k^\ell(\mathbf{x}),\infty). \end{align*} where $\zeta=\nu_i$ for agents of type $t_i=1$ and $\zeta=\omega_i$ for type $t_i=0$. We refer to $\mathcal{Z}_k^\ell(\cdot)$ as the cutoff between $\mathcal{L}(\texttt{d}^{\ell j},\mathbf{x})$ and $\mathcal{L}(\texttt{d}^{k j},\mathbf{x})$, and denote it $\mathcal{V}_k^\ell(\cdot)$ for $t_i=1$ and $\mathcal{W}_k^\ell(\cdot)$ for $t_i=0$.\footnote{We assume that while $\nu$ and $\omega$ have bounded support, the utility functions in $\mathcal{U}^1$ and $\mathcal{U}^0$ are well defined for any real valued $\nu$ and $\omega$, respectively.}

Within a single context, the expected utility theory framework generally satisfies the SCP, which requires that if an agent with a certain degree of risk aversion (the random coefficient $\nu_i$) prefers a safer lottery to a riskier one, then all agents with higher risk aversion also prefer the safer lottery.\footnote{For a discussion of possible failures of SCP, see ape:bal18. } The same is true for the non-expected utility theory model that we use in our empirical analysis in Section (ref). The SCP implies that within a single context, the household's ranking of alternatives is monotone in $\nu_i$ for $t_i=1$ and in $\omega_i$ for $t_i=0$, yielding vertical differentiation of alternatives within each preference type.

Unobserved heterogeneity in consideration sets

Across contexts, the subset of alternatives actually available to each agent is unknown to the researcher, due, e.g., to unobserved budget constraints, liquidity constraints, etc. Moreover, agents face an overall large and potentially overwhelming universe of feasible alternatives, leading to choice overload, cognitive ability constraints, etc. Hence, we allow for unobserved heterogeneity in consideration sets, i.e., in the collection of alternatives that the agents evaluate when making their choices. We denote the overall universe of alternatives across contexts as $\mathcal{D}\equiv\mathcal{D}^\texttt{I}\times\mathcal{D}^\texttt{II}$, with $\mathcal{I}_{\ell,q}$ denoting each of the bundles in $\mathcal{D}$.

assumption[Consideration set formation mechanism] Conditional on $t_i$, agent $i$ draws a consideration set $C_i\subseteq\mathcal{D}$ independently from its random coefficient and from $\mathbf{x}_i$ s.t. \begin{align*} \mathcal{Q}_{1}(\mathcal{K})&\equiv\Pr(C_i=\mathcal{K}|t_i=1)=\Pr(C_i=\mathcal{K}|\mathbf{x}_i,\nu_i,t_i=1), \mathcal{K}\subseteq \mathcal{D},\\ \mathcal{Q}_{0}(\mathcal{K})&\equiv\Pr(C_i=\mathcal{K}|t_i=0)=\Pr(C_i=\mathcal{K}|\mathbf{x}_i,\omega_i, t_i=0), \mathcal{K}\subseteq \mathcal{D}. \end{align*}

The fundamental restrictions imposed in Assumption (ref) are that conditional on preference type, consideration is independent of the agent's random coefficient and of the observed covariate $\mathbf{x}$.\footnote{Recall that our analysis conditions on $\mu_i^j$, hence the distribution of consideration sets may depend on it.} However, the distribution of consideration sets may depend on preference type. Importantly, we allow consideration to be broad, as it is determined at the bundle level instead of within context. A significantly more restrictive approach would posit that consideration is narrow: agent $i$ draws a pair of consideration sets $C_i^j\in\mathcal{D}^j$, $j=\texttt{I},\texttt{II}$ independently across contexts, and forms $C_i=C_i^\texttt{I}\times C_i^\texttt{II}$. As we further discuss in Section (ref), allowing consideration sets to be drawn at the bundle level substantially complicates the identification analysis, but delivers a more realistic model.

Optimal choice within the consideration set

Once the consideration set is drawn, each agent chooses the best alternative in each context according to their preferences.

align[align omitted — 249 chars of source]

where $\zeta=\nu$ if $t=1$ and $\zeta=\omega$ if $t=0$. The bundle choice $\mathcal{I}^*$ depends on the agent's preference type, random coefficient, consideration set, associated premium-deductible tuples, and claim probabilities $\mu^j,~j=\texttt{I},\texttt{II}$.

The flexible model of consideration set formation that we allow for has important implications for the choice problem in Eq. (ref). If we were to assume narrow consideration, hence restrict agents to draw consideration sets independently across contexts, the choice problems would break into independent, context-specific decisions, with

align[align omitted — 141 chars of source]

Each context-specific choice problem satisfies the SCP in Assumption (ref). BaMoTh21 offer a comprehensive analysis of the implications of the SCP for semi-nonparametric identification of a model of discrete choice under risk that features a single preference type and unobserved heterogeneity in consideration sets. Even in the simplified framework where consideration is narrow, our analysis extends theirs as we allow for multiple preference types. More importantly, a narrow model of consideration implies that very similar alternatives in different contexts enter the consideration set independently.\footnote{For example, a \$500 deductible at price $\mathbf{x}^\texttt{I}$ in collision insurance and a \$500 deductible at price $\mathbf{x}^\texttt{II}$ in comprehensive insurance would enter the consideration set independently.} This assumption is unpalatable, particularly when analyzing demand for bundled products. We therefore allow for broad consideration. In doing so, we overcome a substantial hurdle relative to BaMoTh21. When consideration is broad and $C_i$ is formed at the bundle level, the SCP may not necessarily hold across tuples of alternatives, because alternatives may not be monotonically ranked against each other (with respect to $\nu_i$ or $\omega_i$). Hence, here we develop a new approach to obtain point identification of the distribution of preferences, shares of preferences types, and features of the distribution of consideration sets given type.

Identification Results

We begin by describing the conditions under which we can prove our point identification results.\footnote{The results extend easily to more than two contexts, at the cost of heavier notation.} We index bundles as $\mathcal{I}_{\ell,q}$ and $\mathcal{I}_{k,r}$, with $\ell,k\in\mathcal{D}^\texttt{I}$ alternatives in context $\texttt{I}$ and $q,r\in\mathcal{D}^\texttt{II}$ alternatives in context $\texttt{II}$. We recall that in each context, $\texttt{d}^{1j}>\texttt{d}^{2j}>\dots>\texttt{d}^{M^j j}$ and $\mathbf{x}_i^{1j}<\mathbf{x}_i^{2j}<\dots<\mathbf{x}_i^{M_j j}$, see Section (ref). We let $\mathcal{V}^{\ell,q}_{k,r}(\mathbf{x})$ and $\mathcal{W}^{\ell,q}_{k,r}(\mathbf{x})$ denote cutoff levels for $\nu_i$ and $\omega_i$, respectively, at which the agent is indifferent between bundles $\mathcal{I}_{\ell,q}$ and $\mathcal{I}_{k,r}$. Hence, under Assumption (ref), the cutoff $\mathcal{V}^{\ell,q}_{k,r}(\mathbf{x})$ is such that (and similarly for $\mathcal{W}^{\ell,q}_{k,r}(\mathbf{x})$):

multline[multline omitted — 473 chars of source]

Relative to the cutoffs introduced in Assumption (ref), which compared alternatives within a single context and we denoted $\mathcal{V}^\ell_k(\mathbf{x})$ (single superscript and subscript for a single context of choice), we have $\mathcal{V}^{m,q}_{m,r}(\mathbf{x})=\mathcal{V}^q_r(\mathbf{x})$ and $\mathcal{V}^{\ell,s}_{k,s}(\mathbf{x})=\mathcal{V}^\ell_k(\mathbf{x})$ for all $m\in\mathcal{D}^\texttt{I}$ and $s\in\mathcal{D}^\texttt{II}$ (and similarly for $\mathcal{W}^{\cdot,\cdot}_{\cdot,\cdot}(\mathbf{x})$). While cutoffs $\mathcal{V}^{\ell,q}_{k,r}(\mathbf{x})$ and $\mathcal{W}^{\ell,q}_{k,r}(\mathbf{x})$ for $\ell\neq k,q\neq r$ depend on both $\mathbf{x}^\texttt{I}$ and $\mathbf{x}^\texttt{II}$, cutoffs $\mathcal{V}^{\ell,s}_{k,s}(\mathbf{x})$ and $\mathcal{W}^{\ell,s}_{k,s}(\mathbf{x})$ depend only on $\mathbf{x}^\texttt{I}$, while cutoffs $\mathcal{V}^{m,q}_{m,r}(\mathbf{x})$ and $\mathcal{W}^{m,q}_{m,r}(\mathbf{x})$ depend only on $\mathbf{x}^\texttt{II}$. These properties will be used to establish our identification results.

We remark that the cutoffs $\mathcal{V}^{\ell,q}_{k,r}(\mathbf{x})$ and $\mathcal{W}^{\ell,q}_{k,r}(\mathbf{x})$ may not be unique if $\ell>k$ but $q<r$ (or vice versa). However, they are unique whenever $\mathcal{I}_{1,1}$ is compared with any other bundle (and similarly whenever $\mathcal{I}_{M^\texttt{I},M^\texttt{II}}$ is compared with any other bundle).

Throughout, we assume that the researcher has access to data that identify the joint distribution of chosen bundles and covariates. The consideration set, however, is not observed.

assumption[Observed data] A random sample $\{(\mathcal{I}^\ast_i,\mathbf{x}_i^\texttt{I},\mathbf{x}_i^\texttt{II}):i=1,\dots,n\}$ is observed, with $\mathcal{I}^\ast_i$, as defined in Eq. (ref).

Restrictions on variation in $\mathbf{x}^j$ across contexts

Identification of the model's functionals rests on the interplay between the model and the variation in the observed covariates. We only require the covariates $\mathbf{x}_i\equiv(\mathbf{x}_i^\texttt{I},\mathbf{x}_i^\texttt{II})$ to vary across agents and contexts, as formally stated below, but allow $\mathbf{x}_i^\texttt{I}$ (respectively, $\mathbf{x}_i^\texttt{II}$) to be constant across alternatives within $\mathcal{D}^\texttt{I}$ (respectively, $\mathcal{D}^\texttt{II}$). Hence, one needs sufficient variation across contexts to obtain point identification results.

assumption[Preferred within a triplet] In each context $j\in\{\texttt{I},\texttt{II}\}$, for any $\mathbf{x}$ and triplet $\{\texttt{d}^{1j},\texttt{d}^{kj},\texttt{d}^{(k+1)j}\}$, $\forall k\in \{2,...,M^j-1\}$, there are values of $\nu$ (and $\omega$) at which each alternative in this triplet is strictly preferred to the other two.

Assumption (ref) requires that given three coverage levels including the cheapest, each one is preferred by at least some agent. As shown in BaMoTh21, under Assumption (ref), this condition is satisfied for agents of type $t_i=1$ within context $\texttt{I}$ if and only if $-\infty<\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^\texttt{I})<\mathcal{V}^{1,1}_{3,1}(\mathbf{x}^\texttt{I})<\mathcal{V}^{1,1}_{4,1}(\mathbf{x}^\texttt{I})\dots<+\infty$ (and similarly for agents of type $t_i=0$, and for context $\texttt{II}$ with appropriate modifications in the compared bundles and evaluation at $\mathbf{x}^\texttt{II}$ instead of $\mathbf{x}^\texttt{I}$), with $\mathcal{V}^{\ell,q}_{k,r}$ defined through Eq. (ref). So, any agent of type $t_i=1$ who draws $\nu<\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^\texttt{I})$ unambiguously prefers alternative $\ell^{1 \texttt{I}}$ to any other in $\mathcal{D}^\texttt{I}$.

In what follows, an important role is played by the values of $\mathbf{x}=(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ at which the indifference cutoff for an agent of type $t_i$ between alternatives $\ell^{1 \texttt{I}}$ and $\ell^{2 \texttt{I}}$ (the two cheapest alternatives in context $\texttt{I}$) is equal to that agent's indifference cutoff between alternatives $\ell^{1 \texttt{II}}$ and $\ell^{2 \texttt{II}}$ (the two cheapest alternatives in context $\texttt{II}$). We first define these values of $\mathbf{x}$, and then make assumptions on the support of $\mathbf{x}$ to guarantee that it includes them.

definition[Covariate values delivering indifference] Given $t_i$, fix a value of $\nu\in[0,\bar\nu]$ if $t_i=1$ and of $\omega\in[0,\bar\omega]$ if $t_i=0$. Let the set of covariate values at which the agent has preference $\nu$ (respectively, $\omega$) and is indifferent between bundles $\mathcal{I}_{1,1}$, $\mathcal{I}_{1,2}$, and $\mathcal{I}_{2,1}$, be: \begin{align*} \mathbf{X}^1(\nu)&\equiv\{(\mathbf{x}^I,\mathbf{x}^II): \mathcal{V}^{1,1}_{2,1}(\mathbf{x}^I)=\mathcal{V}^{1,1}_{1,2}(\mathbf{x}^II)=\nu\},\\ \mathbf{X}^0(\omega)&\equiv\{(\mathbf{x}^I,\mathbf{x}^II): \mathcal{W}^{1,1}_{2,1}(\mathbf{x}^\texttt{I})=\mathcal{W}^{1,1}_{1,2}(\mathbf{x}^\texttt{II})=\omega\}. \end{align*}

The covariate values $\mathbf{X}^1(\nu)$ (respectively, $\mathbf{X}^0(\omega)$) are the values of $\mathbf{x}=(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ at which an agent with preferences $\nu$ (respectively, $\omega$) is indifferent between the two cheapest coverage levels in context $\texttt{I}$ and, at the same time, also in context $\texttt{II}$. In other words, the agent is indifferent between $\mathcal{I}_{1,1}$, $\mathcal{I}_{1,2}$, and $\mathcal{I}_{2,1}$ (and, hence, $\mathcal{I}_{2,2}$). Given the single crossing property in Assumption (ref), within each context it is immediate to see that both elements of $\mathbf{X}^1(\nu)$ (the covariate value in context $\texttt{I}$ and the covariate value in context $\texttt{II}$) are strictly monotone in $\nu$ (and, similarly, both elements of $\mathbf{X}^0(\omega)$ are monotone in $\omega$). For example, the higher is $\nu$, the higher is the base price in context $\texttt{I}$ at which the agent with random coefficient $\nu$ is indifferent between $\mathcal{I}_{1,1}$ and $\mathcal{I}_{2,1}$. Hence, we can represent $\mathbf{X}^1(\nu)$ (respectively, $\mathbf{X}^0(\omega)$) as a strictly monotone function on the support of $(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$.\footnote{See Figure (ref) and its discussion below.} We assume that these strictly monotone functions intersect on a set of measure zero.

assumption[Distinct contexts] The contexts are distinct, in the sense that: \begin{enumerate}[label=(\Roman*)] • $\mathbf{X}^1(\nu)\neq\mathbf{X}^0(\omega)~a.e.$ • The following four conditions are satisfied: \begin{align} \mathcal{V}^{1,1}_{\ell,q}(\mathbf{X}^0(\omega))&\neq\mathcal{V}^{1,1}_{k,r}(\mathbf{X}^0(\omega)) a.e.\\ \mathcal{W}^{1,1}_{\ell,q}(\mathbf{X}^1(\nu))&\neq\mathcal{W}^{1,1}_{k,r}(\mathbf{X}^1(\nu)) a.e. \\ \mathcal{V}^{1,1}_{\ell,q}(\mathbf{X}^1(\nu))&\neq\mathcal{V}^{1,1}_{k,r}(\mathbf{X}^1(\nu)) a.e. \forall \{\ell,q,k,r\} s.t. \{\ell,q,k,r\}\setminus \{1,2\}\neq\emptyset.\\ \mathcal{W}^{1,1}_{\ell,q}(\mathbf{X}^0(\omega))&\neq\mathcal{W}^{1,1}_{k,r}(\mathbf{X}^0(\omega)) a.e. \forall \{\ell,q,k,r\} s.t. \{\ell,q,k,r\}\setminus \{1,2\}\neq\emptyset. \end{align} \end{enumerate}

Assumption (ref)-(ref) implies Assumption (ref)-(ref), as Eqs. (ref)-(ref) for $\ell,q=2,1$ and $k,r=1,2$ imply $\mathbf{X}^1(\nu)\neq\mathbf{X}^0(\omega)~a.e.$ Both conditions require that at any value of $(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ at which indifference across $\mathcal{I}_{1,1}$, $\mathcal{I}_{1,2}$, and $\mathcal{I}_{2,1}$ occurs for an agent of type $t_i=1$, such indifference cannot occur for an agent of type $t_i=0$. Additionally, Assumption (ref)-(ref) requires that at any value of $(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ at which indifference across $\mathcal{I}_{1,1}$, $\mathcal{I}_{1,2}$, and $\mathcal{I}_{2,1}$ occurs, no other triplet of bundles including $\mathcal{I}_{1,1}$ can generate a three-way tie in utility ranking. Given the data and utility models across preference types, one can directly check whether Assumption (ref) is satisfied.

figure[figure omitted — 211 chars of source]

Finally, we require that the support of $\mathbf{x}$ is sufficiently rich, as point identification of $f(\nu)$ and $g(\omega)$ can only occur at values of $\nu$ and $\omega$ that belong, respectively, to intervals $[\nu^*,\nu^{**}]\subseteq [0,\bar\nu]$ and $[\omega^*,\omega^{**}]\subseteq[0,\bar\omega]$ satisfying the next assumption.

assumption[Independent variation in $\mathbf{x}$] Let $[\nu^*,\nu^{**}]\subseteq [0,\bar\nu]$ and $[\omega^*,\omega^{**}]\subseteq[0,\bar\omega]$ be intervals such that, for some $\epsilon>0$, the random vector $\mathbf{x}=(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ has strictly positive density on the sets $\mathcal{S}^1_\epsilon(\nu^*,\nu^{**})\subset \mathbb{R}^2$ and $\mathcal{S}^0_\epsilon(\omega^*,\omega^{**})\subset \mathbb{R}^2$, with \begin{align*} \mathcal{S}^1_\epsilon(\nu^*,\nu^{**})&=\left\{\mathsf{B}_\epsilon(\mathbf{X}^1(\nu)), \nu\in[\nu^*,\nu^{**}]\right\},\\ \mathcal{S}^0_\epsilon(\omega^*,\omega^{**})&=\left\{\mathsf{B}_\epsilon(\mathbf{X}^0(\omega)), \omega\in[\omega^*,\omega^{**}]\right\}. \end{align*} where $\mathsf{B}_a(c)$ denotes a ball in $\mathbb{R}^2$ of radius $a$ centered at $c$.

Assumption (ref) guarantees that for each $\nu\in[\nu^*,\nu^{**}]$ there are values of $\mathbf{x}$ such that $\mathbf{X}^1(\nu)$ is non-empty and that there is an $\epsilon$-neighborhood around $\mathbf{X}^1(\nu)$ with positive density (and similarly for $\mathbf{X}^0(\omega)$ and all $\omega\in[\omega^*,\omega^{**}]$). This yields sufficient observed variation in $\mathbf{x}$ to identify the functionals that we are after. We illustrate the notion of distinct contexts and independent variation in $\mathbf{x}$ via Figure (ref), which depicts $\mathbf{X}^0(\omega)$ and $\mathbf{X}^1(\nu)$ drawn for different pairs of $\mu$'s. First, $\mathbf{X}^0(\omega)$ and $\mathbf{X}^1(\nu)$ intersect only at a single point.\footnote{In our empirical model described in Section (ref), this intersection point corresponds to $\nu=0$ and $\omega=1$, i.e., respectively, no risk aversion and no probability distortions. } Second, these curves are both monotone. We present them with a scatterplot of unconditional data from our empirical application in the background, to highlight the fact that even when variation in $\mathbf{x}$ does not cover the entire $\mathbb{R}^2_{+}$, identification is attainable since Assumption (ref) requires variation in $\mathbf{x}$ only to cover respective neighborhoods of $\mathbf{X}^0(\omega)$ and $\mathbf{X}^1(\nu)$.

figure[figure omitted — 1,853 chars of source]

We next explain why, under full consideration, our assumptions suffice for identification of the share of preference types and the distributions of the respective random coefficients. Fix a value of $\nu\in[\nu^*,\nu^{**}]$ at which one wants to learn $f(\nu)$. Under Assumption (ref), $\mathbf{X}^1(\nu)$ is non-empty and there is an $\epsilon$-ball of positive density around it. Along with Assumption (ref), this implies that there is a vector $(\mathbf{x}^{\texttt{I} \prime},\mathbf{x}^{\texttt{II}\prime})\in \mathsf{B}_\epsilon(\mathbf{X}^1(\nu))$ such that $\nu=\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^{\texttt{I} \prime})<\mathcal{V}^{1,1}_{1,2}(\mathbf{x}^{\texttt{II} \prime})$ and $\mathcal{W}^{1,1}_{2,1}(\mathbf{x}^{\texttt{I} \prime})>\mathcal{W}^{1,1}_{1,2}(\mathbf{x}^{\texttt{II} \prime})$. Then, as shown in Figure (ref), under Assumptions (ref) and (ref),\footnote{Recall that these assumptions, jointly, imply that any agent who draws $\nu<\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^{\texttt{I} \prime})<\mathcal{V}^{1,1}_{1,2}(\mathbf{x}^{\texttt{II} \prime})$ unambiguously prefers alternative $\ell^{1\texttt{I}}$ to all other alternatives in $\mathcal{D}^\texttt{I}$, unambiguously prefers alternative $\ell^{1\texttt{II}}$ to all other alternatives in $\mathcal{D}^\texttt{II}$, and therefore unambiguously prefers bundle $\mathcal{I}_{1,1}$ to any other bundle in $\mathcal{D}$.}

align[align omitted — 232 chars of source]

In turn, owing to the fact that $\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^{\texttt{I} \prime})$ depends on $\mathbf{x}^\texttt{I}$ but $\mathcal{W}^{1,1}_{1,2}(\mathbf{x}^{\texttt{II} \prime})$ does not, this yields

align[align omitted — 270 chars of source]

where the term $\frac{\partial\mathcal{V}^{1,1}_{2,1}(\mathbf{x}^{\texttt{I} \prime})}{\partial \mathbf{x}^{\texttt{I}}}$ is a known function of $\mathbf{x}^\texttt{I}$ and is different from zero due to Assumption (ref) (where cutoff functions are assumed to be strictly monotone in $\mathbf{x}$). If $[\nu^*,\nu^{**}]=[0,\bar\nu]$, one can repeat the above argument for all $\nu$ on the support and then use the fact that $f(\nu)$ integrates to one to learn $\alpha$. One can similarly learn $g(\omega)$, $\omega\in[\omega^*,\omega^{**}]$.

Restrictions on the consideration set formation mechanism

In the presence of limited consideration, the above argument does not directly apply, as one needs to account for all possible consideration sets in which bundle $\mathcal{I}_{1,1}$ is included. We therefore need to introduce additional notation and some restrictions.

For any $\mathcal{K}_1,\mathcal{K}_2\subseteq \mathcal{D}$, $\mathcal{K}_1\cap\mathcal{K}_2=\emptyset$, denote the probability that all elements of $\mathcal{K}_1$ are included in the consideration set while all elements of $\mathcal{K}_2$ are excluded from it, by

align*[align* omitted — 318 chars of source]

and define $\mathcal{O}_{0}(\mathcal{K}_1;\mathcal{K}_2)$ similarly, where $\mathcal{Q}_{t}(\mathcal{K})$, $t=0,1$, was introduced in Assumption (ref).

Denote by $\mathbb{B}(\mathcal{I}_{\ell,q},\mathbf{x};\zeta)$ the collection of bundles that, at a given value of $\zeta$, strictly dominate bundle $\mathcal{I}_{\ell,q}$, with $\zeta=\nu_i$ for agents of type $t_i=1$, and $\zeta=\omega_i$ for $t_i=0$:

align*[align* omitted — 186 chars of source]

Then, for a given value of $\mathbf{x}$, any bundle $\mathcal{I}_{\ell,q}\in\mathcal{D}$ is chosen if and only if it is considered and every bundle that dominates it is not:\footnote{Equivalently, bundle $\mathcal{I}_{\ell,q}$ is chosen if and only if it is the first best among the ones considered:

align*[align* omitted — 575 chars of source]

}

align[align omitted — 298 chars of source]

Eq. (ref) with $(\ell,q)=(1,1)$ shows that $\mathcal{I}_{1,1}$ is chosen when it is the bundle in $C_i$ with the highest certainty equivalent, i.e., no bundle that yields a higher certainty equivalent (those in $\mathbb{B}(\mathcal{I}_{1,1},\mathbf{x};\cdot)$) is considered. Hence, an agent choosing $\mathcal{I}_{1,1}$ switches to or from a different bundle $\mathcal{I}_{k,r}$ if and only if (i) they are indifferent between $\mathcal{I}_{1,1}$ and $\mathcal{I}_{k,r}$; and (ii) they do not consider any bundle in $\mathcal{D}$ that dominates $\mathcal{I}_{1,1}$ and $\mathcal{I}_{k,r}$. As the indifference cutoffs involving bundle $\mathcal{I}_{1,1}$ are unique, differentiating Eq. (ref) we have

align[align omitted — 663 chars of source]

The summation in Eq. (ref) collects all relevant consideration sets across preference types and indifference points (cutoffs), weighted by the density function at these indifference points and taking into account how the change in $\mathbf{x}^\texttt{I}$ affects the indifference points themselves.\footnote{For $\frac{\partial\Pr(\mathcal{I}^*=\mathcal{I}_{1,1}|\mathbf{x})}{\partial \mathbf{x}^{\texttt{II}}}$, the right-hand-side of Eq. (ref) remains as is, with $\partial \mathbf{x}^\texttt{II}$ replacing $\partial \mathbf{x}^{\texttt{I}}$.}

We impose the following restrictions on the consideration set formation mechanism:

assumption[Minimally informative consideration] One of the following holds: \begin{enumerate}[label=(\Roman*)] • $\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{2,1}\};\emptyset)=\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{1,2}\};\emptyset)>0.$$\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{2,1}\};\emptyset)-\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{1,2}\};\emptyset)\neq 0$, and • $\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\};\emptyset)=\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\}; \{\mathcal{I}_{2,2},\mathcal{I}_{1,2}\})$.\footnote{Alternatively, $\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{1,2}\};\emptyset)=\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{1,2}\}; \{\mathcal{I}_{2,2},\mathcal{I}_{2,1}\})$ can replace the last condition in Assumption (ref)-(ref). In our application this alternative restriction is satisfied because bundle $\mathcal{I}_{1,2}$ (which is the deductible bundle $\{\$1000,\$500\}$) is chosen with probability zero, and hence both probabilities are zero.} \end{enumerate} One of these two restrictions also holds with $\mathcal{O}_{0}$ replacing $\mathcal{O}_{1}$.

Assumption (ref)-(ref) requires symmetry in the probability with which the triplets $(\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{1,2})$ and $(\mathcal{I}_{1,1},\mathcal{I}_{2,2},\mathcal{I}_{2,1})$ are included in the consideration set, and that each probability is strictly positive, so that information can be extracted through the differentiation in Eq. (ref). Assumption (ref)-(ref) requires that if such symmetry is absent, then alternatives $\mathcal{I}_{1,1}$ and $\mathcal{I}_{2,1}$ can only be considered together when neither $\mathcal{I}_{1,2}$ nor $\mathcal{I}_{2,2}$ are considered (a trivial case that would guarantee this condition is that $\mathcal{I}_{2,1}$ is never considered when $\mathcal{I}_{1,1}$ is). The conditions in Assumption (ref) are sufficient (together with the other assumptions listed above) for our identification results. However, they can be replaced by technical yet verifiable assumptions on the behavior of the cutoffs involving comparisons of alternatives $\mathcal{I}_{1,1},\mathcal{I}_{2,1},\mathcal{I}_{1,2},\mathcal{I}_{2,2}$.\footnote{These conditions are available from the authors upon request, and require that $\partial\mathcal{V}^{1,1}_{1,2}(\mathbf{x})/\partial \mathbf{x}^{\texttt{II}}$ does not equal a specific linear function of $\partial\mathcal{V}^{1,1}_{2,1}(\mathbf{x})/\partial \mathbf{x}^{\texttt{I}}$.}

Point identification results

We next state our main identification results, whose proofs are in the Appendix.

theoremLet Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref), (ref) hold. Then \begin{enumerate} • $f(\cdot)$ is identified up to scale on any interval $[\nu^*,\nu^{**}]$ satisfying Assumption (ref). • $g(\cdot)$ is identified up to scale on any interval $[\omega^*,\omega^{**}]$ satisfying Assumption (ref). • If $[\nu^*,\nu^{**}]=[0,\bar\nu]$ and $[\omega^*,\omega^{**}]=[0,\bar \omega]$, then $f(\cdot)$ and $g(\cdot)$ are identified. \end{enumerate}

Theorem (ref) shows that under limited consideration, despite the lack of independent variation in observed covariates across alternatives (within a single context), it is nonetheless possible to identify the distribution of the random coefficient for each preference type without relying on identification at infinity arguments.\footnote{If one had variation in $\mathbf{x}^j$ across alternatives and unbounded support, letting the observed covariate (say, price) for a given alternative go to infinity would be akin to assuming that one observes agents repeated choices in context $j$ while facing feasible sets that include/exclude each single alternative.} While to pin down the entire distribution of preferences large support is required, our approach identifies (up to scale) the density function of each random coefficient conditional on a given interval. Let $\bar{V}$ (respectively, $\bar{W}$) denote the union of all intervals $[\nu^*,\nu^{**}]$ (respectively, $[\omega^*,\omega^{**}]$) satisfying Assumption (ref). If $\bar{V}$ is a proper subset of $[0,\bar\nu]$ (respectively, $\bar{W}$ is a proper subset of $[0,\bar\omega]$), partial identification of the entire distribution of preferences is still possible, by collecting the probability distribution functions that have density equal to $f(\nu)$ for all $\nu\in\bar{V}$ (respectively, $g(\omega)$ for all $\omega\in\bar{W}$). For a general treatment of partial identification of preferences in discrete choice models with limited consideration, see BCMT21.

One can point identify the shares of preference types under a mild additional restriction, where the probability of including one specific pair of bundles in the consideration set and excluding another specific bundle (or pair of bundles) is independent of preference type.

corollary$\alpha$ is identified if all Assumptions of Theorem (ref) hold, and either: \begin{enumerate} • Assumption (ref)-(ref) holds for both agents with preference types $t_i=1$ and $t_i=0$, and \[\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\};\emptyset)-\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\}; \{\mathcal{I}_{2,2},\mathcal{I}_{1,2}\})=\mathcal{O}_{0}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\};\emptyset)-\mathcal{O}_{0}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,1}\}; \{\mathcal{I}_{2,2},\mathcal{I}_{1,2}\}).\] • Assumption (ref)-(ref) holds for both agents with preference types $t_i=1$ and $t_i=0$, and \[\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2}\}; \mathcal{I}_{2,1})-\mathcal{O}_{1}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2}\}; \mathcal{I}_{1,2})= \mathcal{O}_{0}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2}\}; \mathcal{I}_{2,1})-\mathcal{O}_{0}(\{\mathcal{I}_{1,1},\mathcal{I}_{2,2}\}; \mathcal{I}_{1,2}).\] \end{enumerate}

Given the distributions of the random coefficients, $F(\cdot)$ and $G(\cdot)$, the system of equations defined in Eq. (ref) ($L\times M$ equations for a given $\mathbf{x}$) is linear in the consideration probabilities across the two types, weighted by their respective shares $\alpha$ and $1-\alpha$. This in turn implies that we have a continuum of $L\times M$ linear equations to pin down $2^{L\times M+1}$ parameters. In general, with sufficient variation in $\mathbf{x}$, these parameters are over-identified, subject to standard non-redundancy assumptions.\footnote{For example, if for type $t_i=1$ alternative $\mathcal{I}_{\ell,k}$ dominates alternative $\mathcal{I}_{q,r}$, $\mathcal{Q}_1(\{\mathcal{I}_{\ell,k},\mathcal{I}_{q,r}\})$ cannot be separately identified from $\mathcal{Q}_1(\{\mathcal{I}_{\ell,k}\})$.} However, depending on the specific models of preferences assumed, and on the richness of variation in the data observed, it may not be possible to identify some parts of the distribution of consideration sets. Nevertheless, for a specific model, given the data, one can test whether a full rank system of equations results across observed values of $\mathbf{x}$ chen:fang19.

More broadly, our limited consideration model has several testable implications. We highlight two: one specific to our broad consideration case, the other more general. First, suppose Assumption (ref) holds. Then under full or narrow consideration, the marginal distribution of choices in context $\texttt{I}$ is invariant to changes in $\mathbf{x}^\texttt{II}$ and vice versa. Under broad consideration this is not the case, as can be seen through a simple example where $\mathrm{card}(\mathcal{D}^j)=2$ for both $j=\texttt{I}$ and $j=\texttt{II}$, and a positive share of agents consider only the two bundles $\{\mathcal{I}_{1,1}, \mathcal{I}_{2,2}\}$. Hence, one can test for violations of a narrow consideration model by checking whether the marginal distribution of choices in context $\texttt{I}$ (respectively, $\texttt{II}$) responds to changes in $\mathbf{x}^\texttt{II}$ (respectively, $\mathbf{x}^\texttt{I}$). A second testable implication of the model is obtained as follows. Recall that our identification argument focuses on the cheapest bundle, $\mathcal{I}_{1,1}$, and is built by looking at how its share responds to changes in $\mathbf{x}^\texttt{I}$ and $\mathbf{x}^\texttt{II}$. An identical argument can be constructed by focusing on the most expensive bundle, $\mathcal{I}_{M^\texttt{I},M^\texttt{II}}$. Hence, the density functions $f(\nu)$ and $g(\omega)$ can be recovered through two different channels. If they do not coincide, this implies that at least one modeling assumption is violated.

We conclude by comparing the amount of variation in $\mathbf{x}=(\mathbf{x}^\texttt{I},\mathbf{x}^\texttt{II})$ that we require for our point identification results, with that required in the closely related prior work of BaMoTh21 to obtain semi-nonparametric point identification of a model with a single preference type. BaMoTh21's results are derived for an environment where agents are observed making choices only in a single context and with a single source of independent data variation, say context $\texttt{I}$ with variation in $\mathbf{x}^\texttt{I}$. The covariate $\mathbf{x}^\texttt{I}$ is assumed to vary independently across agents; however, for a given agent there is no requirement of independent variation in $\mathbf{x}^\texttt{I}$ across alternatives in $\mathcal{D}^\texttt{I}$ (similarly to this paper). Due to the less rich choice environment observed, to recover the conditional distribution of preferences, BaMoTh21 impose stronger restrictions than we do here on the consideration set formation mechanism.\footnote{For example, BaMoTh21 require that whenever $\ell^{1\texttt{I}}$ is considered, $\ell^{2\texttt{I}}$ is also considered. They do so because there is not a one-to-one mapping between $\partial\Pr(\mathcal{I}^*=\mathcal{I}_{1}|\mathbf{x})/\partial \mathbf{x}^{\texttt{I}}$ and the (up-to-scale) density function evaluated at a single point. Rather, $\partial\Pr(\mathcal{I}^*=\mathcal{I}_{1}|\mathbf{x})/\partial \mathbf{x}^{\texttt{I}}$ maps into a linear combination of the density function evaluated at cutoffs $\mathcal{V}^{1}_{k}(\mathbf{x}^{\texttt{I}}), k>1$. In contrast, here by properly utilizing variation in $\mathbf{x}^{\texttt{II}}$ we are able to create such a mapping even though there can be multiple preference types.}

Model & Data on Choices in Automobile Insurance

Empirical model

As introduced in Section (ref), we model agents' choices in two contexts of insurance coverage, where each coverage provides full insurance against covered losses in excess of a deductible chosen by the agent. In our data, the decision maker is a household; hence, we refer to agents as households. As a reminder, $\mu_i^j$ denotes the probability of household $i$ experiencing a claim in context $j$; for each coverage $j\in \{\texttt{I},\texttt{II}\}$, household $i$ faces a menu of premium-deductible pairs, $\mathcal{M}_i^j\equiv\{(\texttt{d}^{\ell j},\mathbf{x}_i^{\ell j}):\ell\in\mathcal{D}^j\}$, where $\mathbf{x}_i^{\ell j}$ is the household-specific premium associated with deductible $\texttt{d}^{\ell j}$ and $\mathcal{D}^j$ is the set of deductible options offered in context $j$. As discussed in Section (ref), for each context $j\in \{\texttt{I},\texttt{II}\}$ the ratio of the price of deductible $\texttt{d}^{\ell j}$ to the price of deductible $\texttt{d}^{k j}$ is constant across households for all $\texttt{d}^{\ell j},\texttt{d}^{k j}\in\mathcal{D}^j$.

We make assumptions, that are widespread in the literature on property insurance, related to filing claims and their probabilities:

assumption[Restrictions Related to Claim Probabilities] {\color{white}line} \begin{enumerate}[label=(\Roman*)] • Households disregard the possibility of experiencing more than one claim during the policy period. • Any claim exceeds the highest available deductible; payment of the deductible is the only cost associated with a claim; the household's deductible choice does not influence its claim probability. \end{enumerate}

We assume that the two types of preferences described in Section (ref) result from either Expected Utility Theory (EU) or Yaari's Yaari1987 Dual Theory (DT). Within EU, a single-context lottery is evaluated through

align[align omitted — 208 chars of source]

where $w_{i}$ is the household's wealth and $u_{i}(\cdot)$ is its Bernoulli utility function, which under Assumption (ref) is the same for each context. In the EU model, utility is linear in the probabilities and aversion to risk is driven by the shape of the utility function $u_{i}(\cdot)$.

Yaari's Yaari1987 DT model aims at decoupling the decision maker's attitude towards risk from her attitude towards wealth. Within DT, a single-context lottery is evaluated through

align[align omitted — 224 chars of source]

where $\Omega_{i}(\cdot)$ is the household's probability distortion function, which under Assumption (ref) is the same for each context. In the DT model, utility is linear in the outcomes and aversion to risk is driven by the shape of the probability distortion function $\Omega_{i}(\cdot)$.\footnote{Probability distortions are featured also in, e.g., prospect theory Kahneman1979,Tversky1992, rank-dependent expected utility theory Quiggin1982, Gul1991 disappointment aversion theory, and Koszegi2006,Koszegi2007 reference-dependent utility theory.} We remark that in our setting (as well as in many others where subjective beliefs data are not collected and the analysis relies on an often implicit rational expectations assumption), the DT model is indistinguishable from one in which agents' subjective loss probabilities systematically deviate through the $\Omega_{i}(\cdot)$ function from the objective ones.

To strike a balance between model generality and its empirical tractability, we impose shape restrictions on $u_{i}(\cdot)$ and $\Omega_{i}(\cdot)$, respectively. We assume $u_{i}(\cdot)$ exhibits constant absolute risk aversion (CARA):

assumption[CARA] $u_i(y)=\frac{1-\exp(-\nu_i y)}{\nu_i}$ for $\nu_i\neq 0$ and $u_i(y)=y$ for $\nu_i=0$.

Assuming CARA has two key virtues. First, $u_{i}(\cdot)$ is fully characterized by a single parameter: the Arrow-Pratt coefficient of absolute risk aversion, $\nu_{i}\equiv-u_{i}^{\prime\prime}(w_{i})/u_{i}^{\prime}(w_{i})$. Second, $\nu_{i}$ is a constant function of $w_{i}$, and hence we need not observe wealth to estimate $u_{i}(\cdot)$.

To keep the EU model and the DT model on “equal footing," we need $\Omega_{i}(\cdot)$ to be as parsimonious as $u_{i}(\cdot)$. This suggests a single-parameter specification. The literature contains many examples, and we run our analysis with the following one due to Prelec1998:

assumption[Prelec1998's $\Omega(\cdot)$ function] $\Omega_i(\mu)=\exp(-(-\ln\mu)^{\omega_i})$, $\omega_i>0$.

We also carry out our analysis using other utility functions for the EU type (one proposed by Cohen2007 and one by Barseghyan2013) and other probability distortion functions for the DT type (one put forward by Tversky1992 and one by BMT16). The results confirm the main takeaways reported here, and are available from the authors upon request.\footnote{Vuong tests comparing the various models confirm the good fit of our preferred specification.}

The EU and DT models are true alternative theories of decision making under risk.\footnote{Except when both degenerate into net present value calculations with $\nu_i=0$ and $\omega_i=1$.} Neither model is a special case of the other. DT preferences depart from EU preferences in two key ways. First, risk averse behavior is driven by distortions of probabilities for households with DT preferences, but by nonlinear evaluation of wealth for households with EU preferences. Second, narrow bracketing has behavioral implications for households with DT preferences, but not for households with EU preferences. In our framework, where the lotteries are independent across the brackets,\footnote{Independence results from the assumption that claims follow a Poisson distribution, which is imposed in estimating the probability of a claim Barseghyan2013,BTX18.} the choices of a household with EU preferences and CARA utility are independent of the scope of bracketing Rabin2009. The well-known reason is the absence of wealth effects with CARA utility. In contrast, the choices of a household with DT preferences are not independent of the scope of bracketing, because of the rank-dependent nature of how probability distortions are applied.

Within context $j$, the resulting utility function is

equation[equation omitted — 422 chars of source]

While we obtain conditions for nonparametric point identification of $F(\cdot)$ and $G(\cdot)$, for tractability we estimate a fully parametric model via Maximum Likelihood.\footnote{Inspection of Eqs. (ref)-(ref)-(ref) in the Appendix shows that under Assumption (ref), $f(\cdot)$ and $g(\cdot)$ are identified, provided the intervals $[\nu^*,\nu^{**}]$ and $[\omega^*,\omega^{**}]$ in Assumption (ref) are not singletons.}

assumption[Heterogeneity Restrictions] {\color{white}a} \begin{enumerate}[label=(\Roman*),topsep=1ex,itemsep=0pt] • Conditional on $t_i=1$, $\nu_i$ follows a Beta distribution on $[0,0.025]$ with parameter vector $(\gamma_{\nu 1},\gamma_{\nu 2})$ and is independent of $[(\mu_i^j,\mathbf{x}_i^j),j=\texttt{I},\texttt{II}]$. • Conditional on $t_i=0$, $\omega_i$ follows a Beta distribution on $[0,1]$ with parameter vector $(\gamma_{\omega 1},\gamma_{\omega 2})$ and is independent of $[(\mu_i^j,\mathbf{x}_i^j),j=\texttt{I},\texttt{II}]$. \end{enumerate}

Assumption (ref) specifies that the distributions of $\nu$ and $\omega$ are Beta distributions. The main attraction of the Beta distribution is its flexibility Ghosal2001. Its bounded support is a plus given our setting. A lower bound of zero rules out risk-loving preferences and seems appropriate for insurance markets that exist primarily because of risk aversion. Imposing an upper bound enables us to rule out absurd levels of risk aversion. The choice of 0.025 for CARA is conservative both as a theoretical matter and in light of prior empirical estimates in similar settings Cohen2007,Sydnor2010,Barseghyan2011,Barseghyan2013,BMT16. Similarly, for the probability distortion function, the upper bound of 1 insures over-weighting of probabilities; the lower bound of 0 insures that it is a well-behaved function. None of these constraints is binding in our analysis.

We close the empirical model by restricting how $C_i\subseteq\mathcal{D}=\mathcal{D}^\texttt{I}\times\mathcal{D}^\texttt{II}$ is drawn:

assumption[(Broad) Alternative-Specific Consideration] Household $i$ draws a consideration set $C_i\subseteq\mathcal{D}$ s.t. \begin{align*} \Pr(C_i = G) = \prod_{\mathcal{I}\in G}\phi_\mathcal{I} \prod_{\tilde{\mathcal{I}}\notin G} (1-\phi_{\tilde{\mathcal{I}}}), \forall G\subseteq\mathcal{D}, \end{align*} where $\phi_\mathcal{I}\equiv\Pr(\mathcal{I}\in C_i)=\Pr(\mathcal{I}\in C_i|t_i)\ge0,~\mathcal{I}\in\mathcal{D}$, and $\phi_{\mathcal{I}_{1,1}}=1$.

Assumption (ref) strengthens Assumption (ref) by requiring consideration to be independent of type (in addition to being independent of households' preferences given type). This is not needed to establish identification, but we think it prudent to impose it in our application because, as further discussed below, $|\mathcal{D}|=30$ and allowing for type-dependent consideration would add 60 rather than 30 consideration parameters to the model. Assumption (ref) also adapts the Alternative-specific Random Consideration (ARC) model first proposed by Manski1977 and later axiomatized by man:mar14, to hold over bundles of insurance deductibles across contexts. Each bundle $\mathcal{I}\in\mathcal{D}$ appears in the consideration set with probability $\phi_\mathcal{I}$ independently of other bundles. To avoid empty consideration sets, following Manski1977, we assume that one bundle is always considered, and further impose that the always-considered bundle is the cheapest one.\footnote{Alternatively, we could assume that if the realized consideration set is empty, agents choose one of the alternatives in $\mathcal{D}$ uniformly at random. Our estimation results are robust to this modeling assumption.} Once the consideration set is drawn, the household chooses the best alternative according to its preferences as in Eq. (ref).

Data Description

We obtained the data from a large U.S. property and casualty insurance company. The company offers several lines of insurance, including auto. As explained in Section (ref), we focus on deductible choices in auto collision and auto comprehensive. Our analysis uses a sample of 7,736 households who purchased their auto and home policies for the first time between 2003 and 2007 and within six months of each other (this is the same sample used by BaMoTh21).\footnote{As explained in BaMoTh21, the dataset is an updated version of the one used in Barseghyan2013. It contains information for an additional year of data and puts stricter restrictions on the timing of purchases across different lines. These restrictions are meant to minimize potential biases stemming from non-active choices, such as policy renewals, and temporal changes in socioeconomic conditions.} We observe households' deductible choices in auto collision and auto comprehensive, and the premiums they paid for these coverages. We also observe the household-coverage specific menus of deductible-premium combinations---i.e., the pricing menus---that were available to the households when they made their deductible choices.

We refer to Section (ref) for a discussion of how households' pricing menus are determined by the company in each context. As explained there, in each context the premium $\mathbf{x}_i^{\ell j}$ associated to deductible $\texttt{d}^\ell,\ell\in\mathcal{D}^j$, is a household-invariant affine function of a household-specific base price $\mathbf{x}_i^j$, and the company determines this base price applying a coverage-specific rating function to household $i$'s coverage-relevant characteristics. Naturally, the base prices $\mathbf{x}_i^\texttt{I}$ and $\mathbf{x}_i^\texttt{II}$ may exhibit substantial correlation due to common factors entering the rating function (this correlation equals 0.74 in our data), highlighting the importance of our weak requirement on variation in $\mathbf{x}$ stated in Assumption (ref) -- which in particular can hold when $\mathbf{x}^\texttt{I}$ and $\mathbf{x}^\texttt{II}$ are strongly correlated (see Figure (ref) and its discussion).

table[table omitted — 1,528 chars of source]

Table (ref) reports the deductible choices of the households in our sample. In each context, the modal choice is \$500. Interestingly, virtually no household purchases a comprehensive deductible larger than their collision deductible. As we discuss in more detail below, this choice pattern cannot be rationalized by standard discrete choice models under the assumption of full consideration, but can easily be explained once one allows for limited consideration.

The top panel of Table (ref) shows that base premiums vary dramatically in our sample. The ninety-ninth percentile of the \$500 deductible is more than ten times the corresponding first percentile in each line of coverage. While not reported in the table, here we summarize the pricing menus. The cost of decreasing the deductible from \$500 to \$250 is on average \$56 in collision and \$31 in comprehensive. The saving from increasing the deductible from \$500 to \$1,000 is on average \$42 in collision and \$23 in comprehensive.

The claim probabilities $\mu_i^j$ stem from BTX18, who estimated them using coverage-by-coverage Poisson-Gamma Bayesian credibility models applied to a large auxiliary panel of more than one million observations. We treat estimated claim probabilities as if they were observed data. Predicted claim probabilities (summarized in the bottom panel of Table (ref)) exhibit substantial variation: the ninety-ninth percentile claim probability in collision (comprehensive) is 4.3 (12) times higher than the corresponding first percentile. Finally, the correlation between claim probabilities and premiums for the \$500 deductible is 0.38 for collision and 0.15 for comprehensive. Hence, there is independent variation in both (although our identification results only require independent variation in premiums).

table[table omitted — 1,098 chars of source]

Evidence in support of unobserved heterogeneity in $C_i$

As discussed in, e.g., BMT16,BCMT21,BaMoTh21, standard models of risk preferences fail to rationalize some salient data patterns. First, in our data the pricing rule in collision coverage is such that (virtually) no household, regardless of their preference type and random coefficient, should choose the $\$200$ deductible under full consideration. The reason is that for agents with lower risk aversion (probability distortions) it is dominated by the $\$250$ deductible, and for agents with higher risk aversion (probability distortions) it is dominated by the $\$100$ deductible.\footnote{An analogous fact can be established even if an i.i.d., type-specific, noise term were added to the utility function in Eq. (ref) at the coverage level or, more broadly, for any model that abides a notion of generalized dominance formally defined in BaMoTh21. } A limited consideration model, even in the case where the consideration set forms narrowly (i.e., with $C_i^\texttt{I}$ drawn independently from $C_i^\texttt{II}$ and $C_i=C_i^\texttt{I}\times C_i^\texttt{II}$) has no problems explaining such a pattern, because it allows for the $\$200$ deductible to be considered without either $\$100$ or $\$250$. Under Assumption (ref) (consideration sets drawn at the bundle level), that is not necessary, because utility comparisons are at the bundle level.

Second, the joint probability mass function of choices across contexts (see Table (ref)) exhibits a striking pattern where virtually none of the 7,736 households purchase a deductible in comprehensive that exceeds the deductible they purchase in collision. Unless prices (and claim probabilities) exhibit strong negative correlation, a feature that does not occur in our data, standard models (e.g., a Mixed Logit with full consideration) under the assumption of context invariant preferences will struggle to replicate this pattern.

A final note pertains to modeling limited consideration as operating at the bundle level, rather than independently across contexts. A model where limited consideration operates independently across contexts may be successful in matching the marginal distribution of choices within each context, but not the joint BaMoTh19. The limited consideration model studied in this paper, by operating on the bundles, does have the capacity to match the joint distribution of choices. By doing so, it also resolves the preference stability debate discussed in, e.g., Barseghyan2011,Einav2012,BMT16. This debate is centered around the fact that while households’ risk aversion relative to their peers is correlated across lines of coverage, implying that households preferences have a stable component, analyses based on revealed preference reject the standard models: under full consideration, for the vast majority of households one cannot find a level of (household-specific) risk aversion that justifies their choices simultaneously across all contexts. Limited consideration allows the model to match the observed joint distribution of choices, and hence their rank correlations. Under limited consideration, testing for preference stability amounts to asking whether one can find a consideration set and a random coefficient (preference parameter) which jointly rationalize an agent's choice, which is inherently weaker then asking whether one can find preferences that rationalize the agent's choice under full consideration BCMT21.

Estimation Results

table[table omitted — 1,879 chars of source]

We begin our discussion of the estimates that we obtain through MLE by focusing on the type of limited consideration that we uncover, and its role in the results one obtains when estimating preferences. Table (ref) reports the estimated consideration probabilities for each bundle (these are the $\phi_\mathcal{I}$ coefficients in Assumption (ref)), along with 95% confidence intervals obtained by subsampling.\footnote{We use subsampling because the parameter vector is on the boundary of the parameter space.} The estimated model is very far from a full consideration one. Bundles where the collision deductible is strictly lower than the comprehensive one are almost never considered (the probability that the bundle $(\$200,\$1000$) is considered is 1/100, and all others are zero).\footnote{Given the choice patterns in the data discussed in Section (ref), this is not surprising, as MLE sets the consideration probability of never-chosen bundles to zero.} The cheapest bundles, excluding the one where the collision deductible is lower than the comprehensive one, are considered most often (the consideration probabilities for $(\$500,\$500)$ and $(\$1000,\$500)$ are, respectively, 0.83 and 0.47).\footnote{Recall that we assume that $(\$1000,\$1000)$ is considered with probability one.}

The presence of limited consideration alters inference about preference types and about the distribution of the random coefficient within each type in essentially every possible way. To illustrate these effects, we estimate preferences in a pure random coefficients model under three scenarios for the consideration set formation mechanism: limited consideration as in Assumption (ref) (our proposed model); triangular consideration, where for $\mathcal{I}=[\ell^\texttt{I},q^\texttt{II}]$, $\phi_\mathcal{I}=0$ when $\ell^\texttt{I}<q^\texttt{II}$ and $\phi_\mathcal{I}=1$ when $\ell^\texttt{I}\ge q^\texttt{II}$; and full consideration, where $\phi_\mathcal{I}=1$ for all $\mathcal{I}\in\mathcal{D}$. In all cases, we estimate a model where households choose their optimal bundle according to Eq. (ref) with the utility function in Eq. (ref).\footnote{Under full consideration, the likelihood of observing non-zero shares of never-the-first-best alternatives is zero. Due to this, in estimation we set the consideration probability of each bundle to 0.99 instead of 1.00.}

Figure (ref) depicts the resulting Prelec distortion function in Assumption (ref) when $\omega_i$ equals the mean, median, 25th and 75th quantile of the distribution $G(\omega)$ estimated in the limited consideration model (left panel), in the triangular consideration model (center panel), and in the full consideration model (right panel), each with a mixture of types. As the figure illustrates, there is substantial variation in the function across these different values of $\omega$, and all functions are substantially far from the $45^o$ line, indicating substantial over-weighting of small probabilities. Of notice is the fact that the over-weighting is larger in the limited consideration model than in the triangular or in the full consideration model.

figure[figure omitted — 355 chars of source]
figure[figure omitted — 279 chars of source]

Figure (ref) depicts the cumulative distribution function $F(\cdot)$ in our limited consideration model (left panel), in the triangular consideration model (middle panel), and in the full consideration model (right panel). Each panel depicts $F(\cdot)$ for a model that assumes that all households are of the EU type (blue line), for our model with a mixture of EU and DT types (red line), and, for the mixture model, also the implied cumulative distribution function for the entire population, where the $(1-\alpha)$ share of DT households has $\nu=0$. The important feature to notice is that in all panels of Figure (ref), the risk aversion displayed is much higher for the EU households in the mixture model than in the single-type model, and the discrepancy grows from the limited to the triangular to the full consideration model.

In Table (ref) we analyze the same interplay between consideration and preferences from a different angle. We report the estimated excess willingness to pay (WTP) of households in our sample to avoid a lottery where with probability 10% the household loses \$500 (hence, the total WTP equals $\$50$ plus the values reported in the table).

table[table omitted — 2,579 chars of source]

A first feature to notice is that the estimated share of EU types is much higher when the model allows for limited consideration than in models that assume triangular or full consideration (almost a half versus 30% and 20% respectively). The implied degree of aversion to risk changes for households of both preference types, but in opposite directions. The top left panel of Table (ref) shows that if one disregards limited consideration, one infers that the risk aversion of EU types is much higher (more than 40% according to our metric) than under limited consideration, but the aversion to risk of DT types is about one third lower under full consideration (and similarly for triangular consideration). The cumulative effect of limited consideration in the overall population results in a near 12 percent higher willingness to pay to avoid the simple lottery relative to a model that imposes full consideration.\footnote{These results are sensitive to the choice of the simple lottery to benchmark willingness to pay. Changing the stakes will induce a non-linear response by the EU types but a linear one by the DT types. Changing the loss probability will induce a non-linear response by the DT types but a linear one by the EU types. }

We conclude by observing that both the full and the triangular consideration model cannot rationalize the choices of a substantial fraction of households in our data and in general deliver a poor fit, as shown in Figure (ref). Even adding an Extreme Value Type I error term to the utility function in Eq. (ref) and estimating a Mixed Logit model does not remedy this problem. Indeed, the Mixed Logits do not fit our data well, while our limited consideration model essentially replicates the observed shares.

For completeness, in the figure we also display the fit of a limited consideration model where consideration is narrow and choice follows from Eq. (ref). While this model fits the data well relative to the Mixed Logit models with full or triangular consideration (compare the third panel to the top two panels in Figure (ref)), it falls short of our benchmark model. This is not surprising: by construction, this model is restrictive in how bundles enter the consideration sets. As a result, it cannot, e.g., set the shares of bundles with $\texttt{d}^\texttt{I}<\texttt{d}^\texttt{II}$ to zero, or match certain features of the joint distribution of chosen alternatives in the two contexts, such as rank correlations of choices across the two different coverages.\footnote{The narrow consideration model implies a rank correlation of .42 while in the data and under the broad consideration model this coefficient equals .61 and .62, respectively. In comparison, in the Mixed Logit model with full consideration this correlation is .45, while with lower triangular consideration it is .65.}

figure[figure omitted — 363 chars of source]

Implications for Welfare Analysis

In our setting, there are three channels for potential welfare losses. First, limited consideration may prevent agents from choosing their first best. Second, if the probability distortions are capturing a mismatch between subjective and objective beliefs about loss probabilities,\footnote{See, e.g., the model with imperfect information in gua:sin23.} agents may not choose their objective first best, even if they consider it. Third, non-expected utility maximizing households (the DT type in our model) may be open to nudging, whereby modifications of market features that leave the behavior (and welfare) of EU households mostly unchanged may trigger large changes in behavior (and welfare) of DT households.

We therefore conduct two welfare exercises aimed at assessing the impact of each of these channels on the welfare of households purchasing auto deductible insurance. In the first exercise, we estimate the impact on welfare of all households having full consideration. To do so, we take the preferences estimated using our limited consideration model, predict each household's optimal choice from the entire menu $\mathcal{D}$, and compute each household's utility gain (in certainty equivalent terms). To carry out this exercise, we need to take a stand on how does the household value alternatives. For the EU type, we use their choice utility (also called decision utility), i.e., the CARA utility function (with $\nu$ distributed according to our estimate of the distribution $F$). For the DT types, we report results both for their choice utility, i.e., using the Prelec distortion function in Eq. (ref) (with $\omega$ distributed according to our estimate of the distribution $G$); and for the case where the probability distortion function is completely removed, so that $\Omega(\mu)=\mu$ and the household values alternatives based on their net present value (NPV). This also allows one to think about the effect of eliminating the mismatch between subjective and objective beliefs about loss probabilities, if this is what the probability distortion function captures.

In the second exercise, we propose a restructuring of the auto insurance market where collision and comprehensive coverage are offered as a single auto insurance product with

align*[align* omitted — 192 chars of source]

where $\mu^{auto}$ is the probability of experiencing a claim in either collision or comprehensive (we disregard the probability that a claim occurs in both contexts within the policy period as this probability is extremely low in our data) and $\mathbf{x}^{\ell\,auto}$ is the premium charged for an auto coverage that offers the same deductible in collision and comprehensive when firms operate under perfect competition or if they use a constant markup rule.

Again, we take the preferences estimated using our limited consideration model, predict each household's optimal choice, and compute the household's utility gain/loss (in certainty equivalent terms). However, to carry out the exercise not only do we need to take a stand on how does the household value alternatives, but, importantly, also on how does the household draw its consideration set after the intervention. For the former, we proceed as in our first welfare exercise, and report results where the EU types value alternatives based on their choice utility, and DT types based on both their choice utility and on the alternatives' NPV. For the latter, we report our results under several scenarios, detailed below. This exercise may help inform the debate on the need to “simplify insurance choice," and clarify the role of limited consideration in mediating nudging effects.

Before presenting the results of these two exercises, we explain why EU and DT households may respond differently to an intervention that combines collision and comprehensive into a single coverage. A defining feature of the DT model is that it is non-linear in probabilities. Hence, offering insurance as a bundle or as a single product may have a first order impact on DT households' choices and welfare. To see why, suppose the probability distortion function is strictly sub-additive (as is the case in our estimated model). Then, under the maintained assumption of narrow bracketing (Assumption (ref)), the agent's willingness to pay to avoid a $\$500$ loss which occurs with a 10 percent chance, is strictly lower than twice their willingness to pay to avoid the same loss with 5 percent chance. Put differently, a single insurance product against two (mutually exclusive) identical losses, instead of a bundle of two products, reduces the degree of over-weighting of loss probabilities. At the same time, combining insurance products into one line of insurance limits choice, and may eliminate the first best alternative. Ceteris paribus, for a fully rational agent making choices according to the EU model, this can only be welfare reducing. Interestingly, there are examples of insurance products that are indeed sold both as a single coverage and as a bundle, such as single limit liability coverage versus bodily injury and property damage in auto insurance.

table[table omitted — 1,470 chars of source]

In summary, our first welfare exercise addresses the question: what is the (average) welfare cost associated with limited consideration? Our second welfare exercise addresses the question: what are the welfare implications of combining collision and comprehensive into a single product, and how does the presence of limited consideration alter these implications?

The top panel of Table (ref) reports our estimates of the welfare losses due to limited consideration. Using the choice utility for each preference type, the welfare losses are about \$30, or 12.7% of the average price of the cheapest bundle. The effect is smaller (\$18 or 7.6%) if for DT types we use the alternatives' NPV as their value (i.e., we shut down the probability distortion). This is expected, since all utilities and utility differences decrease.

The bottom panel of Table (ref) reports estimated welfare changes associated with combining collision and comprehensive insurance into a single product. We carry out the exercise for three different ways in which consideration sets may be drawn after the market intervention. In the worst case scenario, in the sense that consideration is lowest, the probability that deductible $\texttt{d}$ is considered equals the estimated consideration probability for bundle $(\texttt{d},\texttt{d}),\texttt{d}\in\mathcal{D}^{auto}$. In this case, the impact of the intervention is negative, although the magnitude of the effect depends substantially on how the welfare of DT types is evaluated. This is because under choice utility, following the intervention, DT types overweight the overall loss probability to a lesser degree than they did with separate coverages, and this effect attenuates substantially the welfare reduction from not being able to choose from a larger menu. On the other hand, when welfare of DT types is evaluated according to NPV, although the overweighting of loss probabilities affects choice, it does not enter the welfare calculations.

Under full consideration, the best case scenario, the welfare gains for both evaluation approaches are positive and large. Relative to the worst case scenario, this is, of course, expected. What is more interesting is that the welfare gains are higher than those obtained in the counterfactual of full consideration that maintains the status-quo separation between collision and comprehensive insurance. This is because under full consideration, the EU types are worse off when the collision and comprehensive are combined into a single product (for them, the choice set is being reduced without any associated benefit); however, the DT types, despite facing a smaller choice set, benefit from such a reduction because in making choices they overweight losses by a smaller degree. The latter effect dominates, more so when welfare is computed based on choice utility rather than on NPV.

For completeness we also report welfare changes for a case that we label “middle consideration,” in which each deductible in the combined single coverage is considered with a probability equal to the sum of the probability that it is considered either as collision or comprehensive deductible (or with probability one if the sum exceeds one). The results are reported in the middle row of the bottom panel of Table (ref). Even with this intermediate consideration level, the welfare gains are substantial.

Based on these welfare exercises, we argue that the interplay between features of the decision making process at the utility evaluation level and of the consideration mechanism cannot be ignored when analyzing possible market interventions. In the second welfare exercise carried out above, reducing the feasible set may lead to unambiguous welfare gains, provided consideration increases. However, if consideration does not increase, the same intervention can lead to welfare losses that exceed the gains stemming from nudging the non-expected utility maximizers in the population.

Discussion

This paper provides semi-nonparametric point identification results for a model of discrete choice under risk that allows for unobserved heterogeneity in preference types, unobserved heterogeneity within each type, and unobserved heterogeneity in consideration sets, while confronting the fact that the covariates $\mathbf{x}$ characterizing products do not exhibit independent variation across alternatives within a context, but only across contexts. We apply our method to study demand for deductible insurance in two lines of property insurance, and to analyze the welfare implications of an hypothetical market intervention where the two lines of insurance are combined into a single one. Our findings provide evidence of the importance of allowing for the rich amount of unobserved heterogeneity that our model features.

The choice environment that we study in this paper is similar to that studied in BaMoTh21. They offer a comprehensive analysis of the implications of the Spence-Mirlees single crossing property for semi-nonparametric identification of a model of discrete choice under risk that features a single preference type and unobserved heterogeneity in consideration sets. They also illustrate the tradeoff between the common exclusion restrictions and the restrictions on consideration set formation required for semi-nonparametric point identification. Their work is the closest to ours. However, in our model consideration sets are formed at the bundle level (i.e., across contexts), and hence the single crossing property that both BaMoTh21 and we assume to hold within a context, may not necessarily hold across tuples of alternatives. This is because bundles may not be monotonically ranked (with respect to preference parameters) against each other. Hence, the results in BaMoTh21 do not apply and in this paper we develop a new approach to obtain point identification of the distribution of preferences, of the shares of preferences types, and of features of the distribution of consideration sets given type.\footnote{As we allow for multiple preference types, our analysis extends that of BaMoTh21 even in the simplified framework where consideration is independent across contexts.} In BM23 we show that in a richer data environment where the researcher observes a characteristic for each alternative that displays independent variation both across agents and across alternatives, and that affects utility but not consideration, semi-nonparametric point identification holds for a flexible pure random coefficients model with unrestricted dependence between the random coefficients and the consideration set formation mechanism.

The challenges posed to identification of discrete choice models by unobserved heterogeneity in consideration sets have long been recognized Manski1977.\footnote{Many important papers in the theory literature---including papers on revealed preference analysis under limited attention, limited consideration, rational inattention, and other forms of bounded rationality that manifest in unobserved heterogeneity in consideration sets---also grapple with the identification problem masatioglu_12,man:mar14,Caplin2015,Lleras2017,Cattaneo2019. However, these papers generally assume rich datasets---e.g., observed choices from every possible subset of the feasible set---that often are not available in applied work, especially outside of the laboratory. } It is not uncommon for the problem to be ignored, as a textbook assumption is that agents pick an alternative to maximize their utility over the entire feasible set. When heterogeneity in consideration sets is allowed for, point identification of the model often relies on the availability of auxiliary information about the composition or distribution of agents' consideration sets, or on two-way exclusion restrictions, whereby certain variables impact consideration but not preferences and vice versa. A third approach relies primarily on restrictions to the consideration set formation process.\footnote{Examples for the first approach include DelosSantos2012, Conlon2013,HonkaRAND2017,HonkaMS2017; for the second, Goeree2008,vanNierop2010,Gaynor2016,Heiss2016. Recent examples for the third approach include Abaluck2019,Crawford2019,Lu2018.}

When such assumptions may not be credible and one does not have access to auxiliary data or valid exclusion restrictions, BCMT21 provide a method to obtain informative sharp identification regions for the parameters of discrete choice models, even when preferences and consideration sets may depend on each other, under the assumption that agents' consideration sets include at least two alternatives. Cattaneo2019,cat:che:ma:mas21 provide revealed preference theory, testable implications, and partial identification results for preference orderings and attention frequency, in very general models of limited consideration but without heterogeneity in preferences, under the assumption that one observes agents repeated choices (in a single context) while facing varying choice sets.