Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
60,313 characters · 16 sections · 0 citation commands
Identification with Posterior-Separable Latent Information Costs
\frame
Seemingly mistaken decisions are sometimes rational. Often times, choice situations involve uncertainty about the payoffs to different alternatives. In principle individuals can gather information to learn about outcomes. Regardless of the abundance of available information, the extent to which uncertainty can be reduced is constrained when information acquisition is costly. Information costs can arise due to a variety of reasons, such as the cognitive effort to pay attention and the opportunity cost of the time needed to attend to the environment. Rationally inattention (RI) models incorporate attention as a scarce resource. There is a trade-off between accuracy of information and attention costs: on the one hand, the decision maker (DM) wants to learn as much as she can so as to make better informed deicisons; on the other hand, learning is costly in terms of attention effort. Rationally inattentive behavior involves a two-step optimization problem: first, the individual decides how to allocate her attention (i.e. how much and what to learn about the true state of the world -the true payoffs); then, she decides her choice given what was learned.
Recent evidence suggests attention costs play a key role in rejecting the standard random utility model at describing aggregate behavior (Aguiar, Boccardi, Kashaev, and Kim, 2023). Moreover, RI models have been shown to have empirical content (Mat\v{e}jka and McKay, 2015; Caplin, Dean, and Leahy, 2019) and nonparametric test for rationally inattentive behavior have been provided (Caplin and Dean, 2015; Caplin, Dean, and Leahy, 2022). Recent papers have proposed ways to estimate demand that account for costly information acquisition when the characteristics of the different alternatives are complex (Brown and Jeon, 2020) or not directly observable (Joo, 2023), reaching to conclusions, such that limiting the number of options available increases welfare, that contradict standard discrete-choice models. In this paper, I aim at contributing to bridge the gap between the theory of rational inattention and empirical work.
Taking RI models to the data is challenging. First, most of the literature is restricted to individual behavior, not accounting for heterogeneity. Notable exceptions that include heterogeneity are recent papers proposing ways to estimate these models (Brown and Jeon, 2023; Joo, 2023; Liao, 2024). Second, the standard cost function used in the literature, the Shannon entropy, is too empirically restrictive. As shown in Mat\v{e}jka and McKay (2015), in RI models with Shannon entropy, choice probabilities follow a multinomial logit form. Albeit conveniently tractable, this modeling choice comes at the cost of imposing the condition of independence of irrelevant alternatives (IIA) on observable behavior, hence resulting in unrealistic substitution patterns (i.e. the ratio of choice frequencies between two alternatives does not change when adding or substracting a third item from the menu). Recent efforts have been made to keep tractability without harming empirical relevance. Caplin and Dean (2015) use a generic cost function. Fosgerau, Melo, De Palma, and Shum (2020) use Bregmann information costs, a generalized entropy that allows for more realistic substitution patterns. Caplin et al. (2022) propose two classes of posterior-separable cost functions that allow for the use of standard Lagrangian methods to solve the model, while relaxing the restrictions placed by the Shannon entropy. To prevent undesirable behavioral restrictions, I present a theoretical model that generalizes Caplin et al. (2022) by including additively separable latent heterogeneity in both preferences and attention costs. By introducing unobservable heterogeneity, I am able to describe the population of rationally inattentive DMs. To my best knowkledge, this is the first paper using RI for identifying demand that does not make use of an entropy cost.
Third, most features of RI models cannot be directly observable. For instance, attention allocations are modeled either as information structures (e.g. Caplin and Dean, 2015) or Bayes-consistent mixtures of posterior distributions (Caplin et al., 2022). For this reason, the literature has provided an equivalent empirical counterpart: a state-dependent stochastic choice (SDSC) model subject to information costs (Mat\v{e}jka and McKay, 2015; Caplin et al., 2019; Caplin et al., 2022). SDSC data consists of choice frequencies at each possible (discrete) state of the world. I present a constrained SDSC model with additively separable heterogeneity and prove it is equivalent to my theoretical model.
Fourth, most papers in the literature assume that the utilities are known. Moreover, in these models, it is usually not possible to separately identify preferences and attention costs. I address this issue by adding observable covariates (i.e. alternative-specific attributes) in the alternative-specific utility indices. Due to additive separability of latent heterogeneity and an independence assumption, I show that my model admits a representative agent and belongs to the class of perturbed utility models (PUMs), thus allowing me to use the identification results in Allen and Rehbeck (2019). Using conditional probabilities of choice given observed covariates and states, I recover utility indices, (a measure of) welfare, factual changes in welfare, and counterfactual bounds on choice probabilities.
Furthermore, by assuming that alternative-specific utility indices are additively separable in states, I am able to extend the identification results to stochastic choice data, that is, assuming the econometrician observes conditional probabilities of choice given observed covariates only. This latter result is of interest to practitioners of demand estimation since the requirement on observables is only market-level data.
More generally, my paper is related to models of costly information acquisition. A strand of this literature studies models of motivated cognition, where decision makers form beliefs and make choices under uncertainty in the presence of a trade-off between accuracy and desirability of information. Individuals with motivated cognition typically either have preferences over beliefs or derive anticipatory utility from the flow of expected returns that investment in beliefs yield (B�nabou and Tirole, 2016). Within the theory of motivated cognition, models of wishful thinking predict that individuals manipulate their beliefs to maximize their subjective expected utility net of belief-distorsion costs (Caplin and Leahy, 2019; Kovach, 2020). The types of behavior explained by wishful thinking span procrastination, confirmation bias, and polarization. One of the most relevant applications of these models is the market for assets as the behavior predicted by the theory explains the occurrence of price bubbles.
The RI model with heterogeneity I provide is of general interest since it might be extended to cover models of costly information acquisition other than rational inattention, such as wishful thinking. Moreover, it could be used to develop a statistical test for the null hypothesis that a dataset of conditional distributions was generated by a population of rationally inattentive consumers.\\
The rest of the paper proceeds as follows. Section 2 introduces my RI model. Section 3 shows equivalence between this model and an empirical counterpart. Section 4 covers the properties of the latter, which are exploited for establishing identification with conditional mean state-dependent stochastic choice data in Section 5. Section 6 imposes additional structure to the model that enables identification with conditional mean stochastic choice data. Section 7 concludes and discusses future research.
Consider a population of decision makers (DMs) who choose among a finite set of alternatives $A$ whose payoffs vary with the occurrence of different states of the world. The set of conceivable states of the world $\Omega$ is finite and known to them, and at the moment of making a decision, it is uncertain what the actual realization of the state is. $\gamma\in\Delta(\Omega)$ is a belief about the true state, where $\Delta(\Omega)$ denotes the set of probability distributions on $\Omega$. DMs want to pick the alternative that gives the highest perturbed expected utility,
\[ \gamma\cdot u_{a}(\omega)+\mathbf{E}(a), \] where $u_{a}(\omega)$ is a $|\Omega|$-dimensional vector whose $j^{th}$ component is the utility index $u_{a}(\omega_{j})$ representing the desirability of item $a$ in state $\omega_{j}$, and the disturbance function $\mathbf{E}:A\mapsto\mathbb{R}\cup\{-\infty\}$ denotes unobservable heterogeneity in preferences, independent of the state.\footnote{Throughout the paper, I use $\textbf{boldface}$ to refer to random objects.} The distrubance function can be interpreted as heterogeneity across individuals in the population and across choice instances, i.e. a preference shock. Importantly, at the moment of decision-making, the realization of this disturbance function is known to the individual.
In this environment, DMs are endowed with a prior belief $\mu\in\Delta(\Omega)$, which they can update by gathering information about the state. A DM is said to learn something when the posterior belief $\gamma$ she forms differs from the prior. Learning, however, comes at a cost. Let $\mathbf{T}\in\mathcal{T}$ be the posterior-specific attention cost function, where $\mathcal{T}$ is the space of convex functions $T:\Delta(\Omega)\mapsto\mathbb{R}_{+}\cup\{+\infty\}$ that satisfy $T(\mu)=0$. Intuitively, learning nothing is costless and learning something is weakly costly. This function represents latent heterogeneity in disutility from attention effort. One may think of these attention costs as cognitive effort from paying attention or the opportunity cost of the time spent on learning. The realization of the attention costs is also known to the individual at the moment of deciding. Define \[ N^{a}(\gamma;\omega,\mathbf{E},\mathbf{T}):=\gamma\cdot u_{a}(\omega)+\mathbf{E}(a)-\mathbf{T}(\gamma), \] the net expected utility of choosing $a$ at $\gamma$. Note the trade-off between accuracy of information and attention effort. Whereas in principle, individuals would like to learn as much as possible about the true state to make better informed decisions that lead to higher payoffs, acquiring information is costly.
DMs face a two-step decision problem. The first step is deciding the attention allocation: how much to learn from the environment. An attention allocation $Q\in\Delta\left(\Delta(\Omega)\right)$ is a probability distribution over beliefs, where $Q(\gamma)$ denotes the probability of a posterior $\gamma$. Intuitively, the extent to which the formed posterior belief $\gamma$ differs from the prior $\mu$ indicates how much is learned. Let
\[ \mathcal{Q}:=\left\{ Q\in\Delta\left(\Delta(\Omega)\right)\Big|\sum_{\gamma\in Supp(Q)}Q(\gamma)\gamma=\mu\right\} \] be the set of feasible attention allocations, that is, the set of distributions in $\Delta\left(\Delta(\Omega)\right)$ that satisfy Bayes' rule.\footnote{The notation $SuppQ$ refers to the support of distribution $Q$.} The second step is selecting a stochastic choice function $q:Supp(Q)\mapsto\Delta(A)$ which, for each realized posterior, gives a probability distribution over alternatives in the choice set. Let \[ \Lambda:=\left\{ Q\in\mathcal{Q},q:Supp(Q)\mapsto\Delta(A)\right\} \] be the set of feasible posterior-based policies.
The perturbed expected utility of a given feasible posterior-based policy $(Q,q)\in\Lambda$ is given by
\[ \sum_{\gamma\in Supp(Q)}\sum_{a\in A}Q(\gamma)q(a|\gamma)\left(\gamma\cdot u_{a}(\omega)+\mathbf{E}(a)\right). \] The expression
\[ \sum_{\gamma\in Supp(Q)}Q(\gamma)\mathbf{T}(\gamma) \] denotes the attention cost at a Bayes-consistent attention allocation $Q\in\mathcal{Q}$.
A DM is said to be rationally inattentive if she chooses a pair $(\mathbf{Q},\mathbf{q})$ in the subset of feasible posterior-based policies $\boldsymbol{\Lambda}\subset\Lambda$ that majorize the perturbed expected utility net of attention costs,
Write ((ref)) as a two-step optimization problem:
\[ \sup_{Q\in\mathcal{Q}}\sum_{\gamma\in Supp(Q)}Q(\gamma)\sup_{q(\gamma)\in\Delta(A)}\sum_{a\in A}q(a|\gamma)N^{a}(\gamma;\omega,\mathbf{E},\mathbf{T}). \] Fix any $\gamma\in SuppQ$. Then the second-stage problem satisfies:
\[ \sup_{q(\gamma)\in\Delta(A)}\sum_{a\in A}q(a|\gamma)N^{a}(\gamma,\omega,\mathbf{E},\mathbf{T})=N^{a}(\gamma;\omega,\mathbf{E},\mathbf{T}), \] for each $a\in A$ with $\mathbf{q}(a|\gamma)>0$. Define the value of a posterior $\gamma$ as the maximized net expected utility at $\gamma$,
\[ N(\gamma;\omega,\mathbf{E},\mathbf{T}):=\max_{a\in A}\left\{ N^{a}(\gamma;\omega,\mathbf{E},\mathbf{T})\right\} . \] Note that the maximized value attained in $(1)$ can then be written as the optimization of posterior-specific values over Bayes-consistent attention policies,
Before stating my first assumption, I introduce some concepts that are necessary to characterize the solution to the model. Fix any $(E,T)$ in the support. In the interest of compactness in the notation, below I will not explicitly write $(E,T)$ in the functions parametrized by this tuple (e.g. $N(\gamma)\equiv N(\gamma;\omega,E,T)$). Define the supporting function of the hypograph of $N$ in the direction $\lambda\in\mathbb{R}^{|\Omega|}$, \[ \delta^{*}(\lambda|\text{hyp}N):=\sup_{\gamma\in\Delta(\Omega),r\leq N(\gamma)}\lambda_{1}r+\lambda_{2}\cdot\gamma, \] and denote by $\varGamma(\lambda)$ the set of posteriors supported by the supporting hyperplane in the direction $\lambda$,
\[ \varGamma(\lambda):=\left\{ \gamma\in\Delta(\Omega)\Big|\lambda_{1}N(\gamma)-\lambda_{2}\cdot\gamma=\delta^{*}(\lambda|\text{hyp}N)\right\} , \] which generates the matrix
\[ \Gamma_{\lambda}:=[\gamma^{(1)}\;...\;\gamma^{(N)}],\;\text{where }\left\{ \gamma^{(1)}\;...\;\gamma^{(N)}\right\} =\varGamma(\lambda), \] with arbitrary order $n=1,...,N$.\footnote{See Section 13 in Rockafellar (1970) for a more thorough explanation of the support function.}
In Assumption 1, condition $(i)$ means that the model in ((ref)) is the data-generating process. Condition $(ii)$ ensures that the disturbance function is such that there is always at least one alternative $a$ whose perturbed expected utility is a real number. Since the attention cost function evaluated at the prior is $0$, then the feasible strategy $(Q,q)$ with $Q(\mu)=1$ and $q(a|\mu)=1$ gives always a real value. Therefore, the model has at least one maximizer. Condition $(iii)$ implies that for any direction, the number of posteriors on the supporting hyperplane is at most as large as the number of possible states. This condition prevents the model from having mutiple optimal attention allocations. Finally, condition $(iv)$, imposes that for each pair of alternatives $a$ and $b$, their corresponding net expected utilities $N^{a}(\gamma)$ and $N^{b}(\gamma)$ have different slopes at each posterior belief $\gamma$. This condition precludes optimal stochastic choice functions from giving nondegenerate distributions at any posterior supported by the optimal attention allocation. In other words, indifference among alternatives is ruled out at each optimal posterior, thus resulting in a unique optimal choice function. Overall, the conditions stated in this assumption suffice for an optimizer of the model to exist and to be unique.
For any $(D,T)$ in the support, an optimal posterior $\hat{\gamma}$ is said to be uniquely associated with an alternative $a$ in the choice set if at $\hat{\gamma}$, $a$ is chosen with probability $1$ and $a$ is not chosen with positive probability at any other optimal posterior.
Due to functions $\left(N^{a}\right){}_{a\in A}$ being strictly concave in $\gamma$ and since at each optimal posterior, there are no ties among alternatives, then it follows that the supporting hyperplane that characterizes optimal posteriors can be tangent to any function $N^{a}$ at most at one point. As a result, each optimal posterior is uniquely associated with one alternative in the choice set, i.e. if $(\hat{Q},\hat{q})$ is optimal, then it follows that for each $\hat{\gamma}\in Supp(\hat{Q})$ there is $a\in A$ such that $\hat{q}(a|\hat{\gamma})=1$ and $\hat{q}(a|\hat{\gamma}')=0$ for each $\hat{\gamma}'\in Supp(\hat{Q})\backslash\{\hat{\gamma}\}$. I formalize this result below.
Define the set of posteriors at which, the optimal choice is $a$,
\[ \gamma^{a}:=\left\{ \gamma\in Supp(Q)|q(a|\gamma)=1\right\} . \]
If an optimal posterior $\hat{\gamma}$ is uniquely associated with an alternative $a$ in the choice set, it is known that when $\hat{\gamma}$ is realized, $a$ is picked with probability $1$, and that the only optimal posterior at which $a$ is chosen is $\hat{\gamma}$. This property will be crucial to map the model to observables in the next section.
Finally, note that ((ref)) can be further reexpressed as
In this section, I explore the link between my model with heterogeneity and state-dependent stochastic choice data, the empirical primitive in the literature. Such datasets comprise choice frequencies at each possible state.
Define $\mathcal{P}:=\{P:\Omega\mapsto\Delta(A)\}$, the set of state-dependent stochastic choice (SDSC) functions. While attention allocations and choice functions are not directly observable, SDSC functions might be in principle observed by the analyst. I then study how, according to my model, decisions by rationally inattentive individuals relate to SDSC data.
Uniqueness of the maximizer in the rational inattention model entails that, for any given realization of heterogeneity $(D,T)$, the pair consisting of the optimal attention allocation and the optimal choice function $(Q,q)$ unequivocally induces a unique SDSC function $P$. I formalize this result below.
Next, I show what plausibly observable SDSC data reveals from rationally inattentive behavior and provide a state-dependent stochastic choice model subject to attention costs with latent heterogeneity.
Now consider an alternative model where the choice set, uncertainty about payoffs, prior beliefs, and latent heterogeneity, all remain the same as in ((ref)), but where, instead of solving a two-step problem, DMs make a single-step decision. In other words, given a choice set $A$ and an endowment of utility indices $u$ and prior belief $\mu$, individuals drawing latent functions $\mathbf{E}$ and $\mathbf{T}$ choose a SDSC function $\rho\in\mathcal{P}$ that gives a probability distribution over alternatives in $A$ at each state $\omega\in\Omega$. In this model, the perturbed expected utility of an item $a$ is
\[ \sum_{j=1}^{J}\frac{\rho(a|\omega_{j})\mu(\omega_{j})}{\rho(a)}u_{a}(\omega_{j})+\mathbf{E}(a). \] Let $f$ be a vector-valued function $f:\mathcal{P}\times A\mapsto\Delta(\Omega)$ where for $\rho\in\mathcal{P}$, $a\in A$, the $j^{th}$ entry of $f(\rho;a)$ is defined as $f_{j}\left(\rho;a\right):=\frac{\rho(a|\omega_{j})\mu(\omega_{j})}{\rho(a)}$ for $a\in Supp(\rho)$, $\omega_{j}\in\Omega$, and $f\left(\rho;a\right)=0_{|\Omega|-1}$ for $a\in A\backslash\{Supp(\rho)\}$. Interpret the function $f$ evaluated at $(\rho;a)$ as the posterior belief resulting from $\rho(a)$. The net expected utility of $a$ at $f(\rho;a)$ is
\[ N^{a}\left(f(\rho;a);\omega,\mathbf{E},\mathbf{T}\right):=f(\rho;a)\cdot u_{a}(\omega)+\mathbf{E}(a)-\mathbf{T}\left(f(\rho;a)\right). \]
$\boldsymbol{\rho}\in\mathcal{P}$ is chosen to solve
Note that by definition of $N^{a}$, the model can be equivalently expressed as
\[ \sup_{\rho\in\mathcal{P}}\sum_{a\in A}\rho(a)N^{a}\left(f(\rho;a);\omega,\mathbf{E},\mathbf{T}\right). \] Notice further that under Assumption 1 $(iv)$, $a\in Supp(\rho)$ implies that for any $(E,T)$ in the support,
\[ N^{a}\left(f(\rho;a);\omega,E,T\right)>N^{b}\left(f(\rho;a);\omega,E,T\right),\;\text{for each }b\in A. \] Hence, by definition of $N$, the model can be further reexpressed as
Under my assumptions, this model too has a unique maximizer. As a result, fixing any latent functions $(E,T)$ in the support, revealed posteriors, revealed attention and revealed choice functions may be recovered from SDSC data.
Uniqueness in both models allows me to enunciate the main result of the paper.
Under Assumption 1, the optimizer of each version of the model satisfies the one-to-one mapping property with respect to the other. If a pair of an attention allocation and a choice function is the optimizer ((ref)), then it follows that its empirical counterpart, its generated SDSC function, is the maximizer of ((ref)). On the other hand, if a SDSC function is the maximizer ((ref)), then its theoretical counterpart, the pair consisting of its revealed attention and revealed choice function, is the optimizer of ((ref)).
One of the reasons why this result is critical to learn about aggregate demand of a population of rationally inattentive individuals is because it implies the theoretical model and its empirical counterpart, the constrained SDSC model, can be used interchangeably. This is useful as mean SDSC data might in principle be observable.
The SDSC model presented before is a perturbed utility model (PUM). In the interest of identifying structural and counterfactual parameters related to changes in prices and other attributes, I extend the model to incorporate covariates. I then impose assumptions sufficient to derive desirable aggregation properties of the model and features of its optimization structure that enable identification with conditional mean SDSC data.
For each alternative $a\in A$, let $\mathbf{x}_{a}\in\mathcal{X}_{a}\subseteq\mathbb{R}^{L_{a}}$ be a random vector listing attributes of $a$. For each $a\in A$, $\omega_{j}\in\Omega$, let $u_{a}(.,\omega_{j}):\mathcal{X}_{a}\mapsto\mathbb{R}$ be a utility index that represents how attributes affect the desirability of item $a$ at state $\omega$. Denote $u_{a}(x_{a},\omega):=\left(u_{a}(x_{a},\omega_{1}),\;...,\;u_{a,}(x_{a},\omega_{J})\right)^{T}$, the vector of utilities of item $a$ evaluated at $x_{a}$ at each possible state $\omega_{1},...,\omega_{J}\in\Omega$.
For the purpose of identification of structural and counterfactual parameters, the SDSC model is useful, not only due to being the observable counterpart of the theoretical model, but also because it keeps the additively separability property of latent heterogeneity. This condition together with the independence assumption allow me to link my rational inattention model with latent heterogeneity to the identification framework in Allen and Rehbeck (2019).
Concretely, additively separable latent heterogeneity and independence are the key conditions that suffice for the model to admit a representative agent, which is key for establishing identification under the assumption that conditional mean SDSC data is observable.
This aggregation result renders identification of the distribution of unobservable heterogeneity unnecessary to identify utility indices or mean indirect utility, as well as specifies the data requirements for identification. Specifically, it is sufficient to observe conditional mean SDSC functions $\left(\mathbb{E}[\mathbf{P}|\omega,\mathbf{x}=x]\right)_{\omega\in\Omega}$. Note that in both the problem of the representative agent and the average indirect utility $V$, optimization is over function in $\mathcal{P}$. According to the original theorem in Allen and Rehbeck (2019), optimization in both cases is over the convex hull of the set $\mathcal{P}$. However, the convex hull of $\mathcal{P}$ is the set $\mathcal{P}$ itself.
As the aggregation theorem requires the convex hull of the feasibility set, using the SDSC model is necessary for measure-theoretic reasons. If one considers the theoretical model, the convex hull of $\Lambda$ is $\Lambda$ itself. It is clear that, in contrast to $\mathcal{P}$, $\mathcal{Q}$ is an infinite-dimensional set, and hence, does not fit in the theorem proposed by Allen and Rehbeck (2019).
Building upon the aggregation result, I am able to exploit the optimization structure of the model for identification. Concretely, I apply an envelope theorem to the representative agent model and leverage asymmetries of cross-parital derivatives of the conditional mean SDSC function.
Below, I derive the properties that follow from the optimization structure of the model and leave the identification results for the next section.
The fact that the model admits a representative agent indicates that observing the conditional mean SDSC function is the only data requirement for identification. This is to say, it is enough for the analyst to observe the conditional probability distribution of choice at each possible state.
Using assymmetry of the cross-partial derivatives of the conditional mean SDSC function and an exclusion restriction (i.e. covariates are exclusive to one alternative), I identify the utility indices up to location and scale without having to recover the distribution of latent heterogeneity.
By Lemma 3,
\[ \frac{\partial}{\partial x_{a,p}}V\left(u(x,\omega)\right)=\mathbb{E}[\mathbf{P}(a|\omega_{j})|\mathbf{x}=x]\mu(\omega_{j})\frac{\partial}{\partial x_{a,p}}u_{a}(x_{a},\omega_{j}) \]
By Lemma 6,
\[ \dfrac{\frac{\partial}{\partial x_{a,p}}\mathbb{E}[\mathbf{P}(b|\omega)|\mathbf{x}=x]}{\frac{\partial}{\partial x_{b,q}}\mathbb{E}[\mathbf{P}(a|\omega)|\mathbf{x}=x]}=\dfrac{\frac{\partial}{\partial x_{a,p}}u_{a}(x_{a},\omega_{j})}{\frac{\partial}{\partial x_{b,q}}u_{b}(x_{b},\omega_{j})}. \]
Then,
\[ \frac{\partial}{\partial x_{a,p}}V\left(u(x,\omega)\right)=\mathbb{E}[\mathbf{P}(a|\omega_{j})|\mathbf{x}=x]\mu(\omega_{j})\dfrac{\frac{\partial}{\partial x_{a,p}}\mathbb{E}[\mathbf{P}(b|\omega_{j})|\mathbf{x}=x]}{\frac{\partial}{\partial x_{b,q}}\mathbb{E}[\mathbf{P}(a|\omega_{j})|\mathbf{x}=x]}\frac{\partial}{\partial x_{b,q}}u_{b}(x_{b},\omega_{j}) \]
Consider $x',x''\in Supp(\mathbf{x})$. Let $x''$ be identical to $x'$ except for the $(a,p)^{th}$ component, then
\[
\]
Recall $\frac{\partial}{\partial x_{b,q}}u_{b}(x_{b},\omega_{j})$ is the scale term at a fixed value $x_{b}$. If $\mathbf{x}_{b}$ is the price of good $b$ and $\frac{\partial}{\partial x_{b,q}}u_{b}(x_{b},\omega_{j})=-1$, then $V\left(u(x',\omega)\right)-V\left(u(x,\omega)\right)$ can be interpreted as the change in mean indirect utility in terms of dollars.
Define $u_{a}(x_{a},\omega_{j}):=u_{a}(x_{a})+G^{a}(\omega_{j})$, where $G^{a}:\Omega\mapsto\mathbb{R}$. Treat $\boldsymbol{\omega}$ as a latent variable with known support $\Omega$. Importantly, both the support of $\boldsymbol{\omega}$ and the functions $\{G^{a}\}_{a\in A}$ are known to the decision maker at the moment of deciding.
Consider the alternative model where $\boldsymbol{\rho}\in\mathcal{P}$ is chosen to satisfy
Additive separability in states and the independence of covariates and unobservable heterogeneity ensure the model with latent states too admits a representative agent.
This aggregation property of the model indicates that only conditional mean stochastic choice data $\mathbb{E}[\mathbf{P}|\mathbf{x}=x]$ is the only data requirement for identification. The fact that the actual realization of the state needs not to be observed suggests that market-level data might be used, thus broadening the scope of the empirical applications of the model.
I enunciate the envelope theorem below and leave the identification results for the Appendix.
I present a theoretical model with additively separable latent heterogeneity that describes the behavior of a population of rationally inattentive decision makers. Under some regularity conditions, I show that this model is observationally equivalent to a state-dependent stochastic choice model constrained by attention costs. In the interest of learning how demand responds to changes in attributes, I include covariates as arguments in the utility indices. Assuming regressors and unobservable heterogeneity are independent, I show that my model admits a representative agent. This aggregation property, together with the structure of the model, allows me to to identify structural and counterfactual parameters when the conditional mean SDSC function is observable, that is, when the econometrician observes conditional probabilities given states and covariates. In particular, I identify how attributes shift the desirability of different goods, (a measure of) welfare, factual changes in welfare, and bounds on counterfactual probabilities of choice. Further assuming utility indices are additively separable in (latent) states, I establish identification with conditional mean stochastic choice data. In the latter case, as the analyst does not need to observe the realization of the state, the model can be used for empirical applications using market-level data.
I close by mentioning extensions to the paper that I will pursue in further research:
$(i)$ Some applications require the states to be continuous. I will generalize the aggregation theorem so as to accomodate for infinite states.
$(ii)$ Although utility indices are identified up to location and scale and unobservable heterogeneity is recovered, the components of the latter, that is, latent preferences and latent attention costs are not separately identified. I will place bounds on both components of latent heterogeneity by deriving the empirical content of the model, under the assumption that the econometrician observes exogenous variation in choice sets. This strategy might also serve for placing bounds on choice frequencies after a counterfactual change in choice sets.
$(iii)$ So far, I assumed a known, homogeneous prior. I will let the prior be a latent random variable. Using again the empirical content of the model, I will build a finite mixture model based on equivalence classes, i.e. partitions of the space of subjective priors that are equivalent in terms of revealed preference relations.
$(iv)$ I will generalize my model to include other types of costly information acquisition, such as wishful thinking.
Aguiar, V. H., Boccardi, M. J., Kashaev, N., & Kim, J. (2023). Random utility and limited consideration. Quantitative Economics, 14(1), 71-116.
Allen, R., & Rehbeck, J. (2019). Identification with additively separable heterogeneity. Econometrica, 87(3), 1021-1054.
B�nabou, R., & Tirole, J. (2016). Mindful economics: The production, consumption, and value of beliefs. Journal of Economic Perspectives, 30(3), 141-164.
Brown, Z. Y., & Jeon, J. (2020). Endogenous information and simplifying insurance choice. Mimeo, University of Michigan.
Caplin, A., & Dean, M. (2015). Revealed preference, rational inattention, and costly information acquisition. American Economic Review, 105(7), 2183-2203.
Caplin, A., Dean, M., & Leahy, J. (2019). Rational inattention, optimal consideration sets, and stochastic choice. The Review of Economic Studies, 86(3), 1061-1094.
Caplin, A., Dean, M., & Leahy, J. (2022). Rationally inattentive behavior: Characterizing and generalizing Shannon entropy. Journal of Political Economy, 130(6), 1676-1715.
Caplin, A., & Leahy, J. V. (2019). Wishful thinking (No. w25707). National Bureau of Economic Research.
Fosgerau, M., Melo, E., De Palma, A., & Shum, M. (2020). Discrete choice and rational inattention: A general equivalence result. International economic review, 61(4), 1569-1589.
Joo, J. (2023). Rational inattention as an empirical framework for discrete choice and consumer-welfare evaluation. Journal of Marketing Research, 60(2), 278-298.
Kovach, M. (2020). Twisting the truth: Foundations of wishful thinking. Theoretical Economics, 15(3), 989-1022.
Liao, M. (2024). Identification of a rational inattention discrete choice model. Journal of Econometrics, 240(1), 105670.
Ma\'{c}kowiak, B., Mat\v{e}jka, F., & Wiederholt, M. (2023). Rational inattention: A review. Journal of Economic Literature, 61(1), 226-273.
Mat\v{e}jka, F., & McKay, A. (2015). Rational inattention to discrete choices: A new foundation for the multinomial logit model. American Economic Review, 105(1), 272-298.
Rockafellar, R. T. (1970). Convex Analysis. Princeton Math. Series, 28.