Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
156,174 characters · 0 sections · 65 citation commands
Revealed Price Preference: Theory and Empirical Analysis
\email{[email removed]}
\email{[email removed]}
\email{[email removed]}
\address{$^{\between}$University of Toronto, $^\dagger$Yale University, $^\ddagger$Johns Hopkins University, $^\star$Cornell University.} \email{[email removed]}
\makeatletter
\defcitealias{kitamura2018}{KS} \defcitealias{McFadden1991}{MR} \defcitealias{attanasio2020}{AP}
\makeatother \newtheorem{cor}{Corollary} \newtheorem{theorem}{Theorem} \newtheorem*{afriattheorem*}{Afriat's Theorem} \newtheorem*{mcrichtheorem*}{\MR Theorem} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{observation}{Observation} \theoremstyle{definition} \newtheorem{definition}{Definition}[section] \newtheorem{example}{Example} \newtheorem{claim}{Claim} \newtheorem{assumption}{Assumption} \newtheorem{condition}{Condition}
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {Introduction}
A central question in economic analysis is the determination of the welfare effect of price changes. As an example, suppose we observe a consumer's purchases of two goods, gasoline and food, from two separate trips to a grocery store with an on site gasoline retailer. In the first instance $t$, the prices are $p^t=(2,2)$ of gasoline and food respectively and she buys a bundle $x^t=(6,3)$. In her second trip $t'$, the prices are $p^{t'}=(3,1)$ and she purchases $x^{t'}=(5,4)$. The most basic welfare question one can ask here is whether the consumer is better off at the prices prevailing at $t$ or at $t'$ (keeping fixed the prices of all other goods she consumes)? In this paper, we introduce a theoretical framework based on revealed preference, along with a nonparametric econometric technique, that would allow us to answer questions of this type.
A typical approach to this problem is to model the consumer as having a quasilinear utility function $\widetilde{U}(x)-p\cdot x$ since, in particular, this allows for simple “sufficient statistics” analysis of welfare gains or losses using a Harberger formula (see chetty2009 and most recently kleven2021 for an overview of this approach). Of course, the second term ($-p\cdot x$) in the quasilinear utility function captures the fact that the goods being analyzed (food and gasoline in our simple example) do {\em not} constitute the universe of the consumer's consumption; expenditure lowers utility because it reduces the consumption of an outside (numeraire) good.
The point of departure of our analysis is the following simple observation. Without having to model the consumer's preference as quasilinear (or taking any other specific functional form) we can still conclude that she is better off at $t$ compared to $t'$. This is because $p^{t'}\cdot x^{t'}=19$ whereas $p^{t}\cdot x^{t'}=18$. In other words, if the prices prevailing at $t'$ were $p^t$ instead of $p^{t'}$, the consumer would be better off since purchasing the same bundle $x^{t'}$ would cost less, leaving the consumer with more money to buy other goods (outside the set of goods analyzed).\footnote{Another way of seeing this is the following. Suppose $t'$ is a supermarket where the prices are $p^{t'}$ and we observe the bundle $x^{t'}$ being bought by a consumer. If at supermarket $t$, the prices are $p^t$, then we know that the consumer would prefer this supermarket, since the same purchases at $t'$ would cost less at $t$.} More generally, the consumer has a {\em preference over prices} that an analyst could at least partially discern from the data: if at observations $t$ and $t'$, we find that $p^{t}\cdot x^{t'}\leq (<) p^{t'}\cdot x^{t'}$, then
Welfare comparisons made in this way will only be consistent if the revealed preference relation over prices is free of cycles, a property we call the {\em generalized axiom of price preference} (GAPP). This leads inevitably to the following question: precisely what does GAPP mean for consumer behavior?
{\bf Augmented Utility Functions:}\, To answer this question, we assume that the analyst collects a data set $\mathcal{D}=\{p^t,x^t\}_{t=1}^T$ from a consumer; each observation $t$ consists of the prices $p^t\in\mathbb{R}^L_{++}$ of $L$ goods (representing some but not all the goods she consumes) and the consumer's demand $x^t\in\mathbb{R}^L_{+}$ at those prices. We show that GAPP (on $\mathcal{D}$) is both necessary and sufficient for the existence of a strictly increasing function $U:\mathbb{R}^L_+\times \mathbb{R}_{-}\to \mathbb{R}$ that {\em rationalizes $\mathcal{D}$} in the following sense: $$x^t\in\operatornamewithlimits{argmax}_{x\in\mathbb{R}^L_+}U(x,-p^t\cdot x)\:\mbox{ for all $t=1,2,...,T$.}$$ The function $U$ should be interpreted as an {\em expenditure-augmented utility function}, where $U(x,-e)$ is the consumer's utility when she acquires $x$ at the cost of $e$. It recognizes that the consumer's expenditure on the observed goods is endogenous and dependent on prices: she could in principle spend more than what she actually spent (note that she optimizes over $x\in \mathbb{R}^L_{+}$) but the trade-off is the dis-utility of greater expenditure. Note that the quasilinear utility function $U(x,-p\cdot x)=\widetilde U(x)-p\cdot x$ is a special case of an augmented utility function.
The augmented utility model has a number of features that makes it widely applicable and easy to use. We highlight a few of them.
(1)\, Being more general than the quasilinear model, it does not have some of its overly strong implications on the structure of consumer demand (see \hyperref[sec:AUQL]{Section (ref)}). In particular, it is broad enough to accommodate phenomena emphasized in the behavioral economics literature, such as reference dependence, mental budgeting and inattention to prices. We briefly describe the first of these here, a more detailed discussion can be found in \hyperref[sec:behavioral]{Section (ref)}. koszegi2006 and heidhues2008 argue that consumption decisions can depend not just on the actual prices but also on the prices the consumer expected to pay. Specifically, the disutility from spending is greater if the expected price was lower than the sticker price and vice versa. A simple way they propose of capturing this phenomenon is the following function
The first two terms capture standard quasilinear preferences whereas the third term captures a general form of reference dependence.\footnote{For a related model of reference prices leading to a similar functional form, see sakovics2011.} In Koszegi and Rabin's terminology, the consumer gets “gain-loss utility” by comparing the expenditure $p\cdot x$ she incurs on a bundle $x$ against the expenditure $\tilde{p} \cdot x$ she expected to incur, where $\tilde{p}$ are her reference prices.\footnote{A common choice for $F$ is $F(p\cdot x - \tilde{p}\cdot x)=\max\{\bar{k}(p\cdot x - \tilde{p}\cdot x),0\}+\min\{\underline{k}(p\cdot x - \tilde{p}\cdot x),0\}$ where it is typically assumed that $\bar{k}>\underline{k}>0$ or that the consumer feels losses relative to the reference point more severely than commensurate gains.}
(2)\, In this model, a consumer's utility at prices $p$ is given by $\max_{x\in\mathbb{R}^L_+}U(x,-p\cdot x)$, which obviously leads to a ranking or preference on prices. Going further, it is possible to develop notions analogous to compensating and equivalent variations, which gives us a {\em quantitative} sense of how much one set of prices is ranked above another and could form the basis for interpersonal comparisons (see \hyperref[sec:compensation]{Section (ref)}).
(3)\, Readers familiar with \hyperref[thm:Afriat]{Afriat's Theorem} afriat1967 will no doubt have already noticed that we are working in a similar framework. That theorem characterizes a data set $\mathcal{D}=\{p^t,x^t\}_{t=1}^T$ that could be rationalized in the following sense: there is $\widetilde U:\mathbb{R}^L_+\to\mathbb{R}$ such that $\widetilde U(x^t)\geq \widetilde U(x)$ for all $x\in\mathbb{R}^L_+$ that satisfy $p^t\cdot x\leq p^t\cdot x^t$. The notion of rationalization in our model is distinct from that in \hyperref[thm:Afriat]{Afriat's Theorem} (even the utility functions have different domains) and there are data sets that could be rationalized in one sense but not the other. We explain these differences in \hyperref[sec:Afrsection]{Sections (ref)} and (ref).
Empirical researchers who apply \hyperref[thm:Afriat]{Afriat's Theorem} must contend with cases where a data set is not exactly rationalizable. They have developed an easily interpretable way of measuring how close a data set is to being rationalized known as the {\em critical cost efficiency index}. In \hyperref[sec:index]{Section (ref)} we develop a similarly intuitive index that should facilitate empirical applications of the augmented utility model.
(4)\, Our notion of revealed preference over prices is not simply applicable to a Euclidean consumption space. It applies even when goods can only be consumed in discrete quantities (as is often the case in empirical IO models) or when they are represented by characteristics. Furthermore, when prices are nonlinear, it is still possible to compare price systems by asking if an agent could replicate the purchases under one system in another price system. Requiring non-cycling comparisons in this case leads to a natural extension of GAPP and the augmented utility model, which we explain in \hyperref[sec:nonlinearGAPP]{Section (ref)}.
{\bf Random Augmented Utility Model (RAUM):}\, In the second part of the paper, we develop the random version of the augmented utility model, in order to study the demand distribution of a population of consumers drawn from repeated cross-sectional data. We first devise a test to check if the data are consistent with the RAUM. We then develop a procedure to estimate the proportion of consumers who are made better or worse off by a given change in prices; welfare analysis of this kind under general preference heterogeneity is a challenging empirical issue, and has attracted considerable recent research (see, for example, hausman2016 and its references).
Unlike the case of data collected from a single individual, it is worth noting that, in this case, both model testing and welfare analysis are statistical since we need to account for sampling error inherent in repeated cross sectional data. Our RAUM test uses existing (though recently developed) econometric methods. On the other hand, to carry out the welfare analysis, we develop new theoretical econometric results; it is worth stressing that this is a stand alone contribution that has applications beyond this paper.
For reasons we shall now explain, testing the RAUM on actual repeated cross-sectional data (such as household survey data) turns out to be a lot more straightforward than testing the random version of the standard budget-constrained utility model where the population is required to be rational in the sense of \hyperref[thm:Afriat]{Afriat's Theorem} (defined earlier). We refer to the latter model as RUM (random utility model) for short. The test for RUM is broadly set out in McFadden1991, but two challenges must first be overcome. First, McFadden1991 do not account for finite sample issues as they assume that the econometrician observes the population distributions of demand; this hurdle was recently overcome by kitamura2018 who develop a testing procedure which incorporates sampling error. Second, the test suggested by McFadden1991 requires the observation of large samples of consumers who face not only the same prices but also make identical total expenditures. This feature is not true of any real observational data where a consumer's demand (and thus total expenditure) on a set of observed goods will typically be price dependent. Thus to implement their test, kitamura2018 need to first estimate demand distributions at a fixed level of (median) expenditure, which requires the use an instrumental variable technique (with all its attendant assumptions) to adjust for the endogeneity of observed total expenditure.
In contrast, the RAUM can be tested directly on household survey data, even when the demand distribution at a given price vector implies {\em heterogenous levels of total expenditure across consumers}.\footnote{A bit more formally, it is possible for two demand bundles $x$ and $x'$ in the support of the demand distribution when prices are $p^t$ to satisfy $p^t\cdot x\neq p^t\cdot x'$.} This allows us to estimate the demand distribution by simply using sample frequencies and we can avoid the above-mentioned additional layer of demand estimation needed for testing RUM.
The reason for this remarkable simplification is somewhat ironic: we show that a data set is consistent with the RAUM if, and only if, a converted version of the data set (which results in identical expenditures at each price) of the type envisaged by McFadden1991 passes the RUM test suggested by them. In other words, we apply the test suggested by McFadden1991, but not for the model they have in mind. This trick also means that we can use, and in a more straightforward way, the econometric techniques in kitamura2018.
Assuming that a data set is consistent with our model, we can then evaluate the welfare impact of an observed change in prices. Indeed, if we observe the true distribution of demand at each price, it is possible to impose bounds (based on theory) on the proportion of the population who are revealed better off or worse off following an observed change in prices. Of course, when samples are finite, these bounds instead have to be estimated. To do so, we develop new econometric techniques that allow us to form confidence intervals on the proportion of consumers who are better or worse off; these techniques build on the econometric theory in kitamura2018 but are distinct from it.
We emphasize that these new econometric techniques can be more generally applied to linear hypothesis testing of parameter vectors that are partially identified, even in models that are unrelated to demand theory (see, for example, lazzati2018). They provide a new method for estimation and inference in nonparametric counterfactual analysis and, since the evaluation of counterfactuals is an important goal of empirical research, they are potentially very useful to practitioners.
{\bf Empirical Applications:}\, We use separate data sets to demonstrate how welfare analysis can be done using both the deterministic and random versions of our model. First, we use the deterministic augmented utility model to analyze panel data from the Mexican conditional cash transfer program Progresa. Recently, attanasio2020 showed that sellers responded to these transfers by altering the nonlinear prices they charge for staples. We focus our analysis on the untreated households that did not receive cash transfers; we show, via revealed preference over the nonlinear price systems, that these households have tended to benefit from the price changes that occurred during the observation period. This is consistent with the finding in attanasio2020 that the change in the wealth distribution induced by Progresa led to larger quantity discounts (which favored the untreated households because they are usually better-off and consumed more).
Finally, we show how the RAUM can be used to estimate the welfare impact of the changes in observed prices in repeated cross-sectional data. Specifically, we take the model to two separate national household expenditure data sets from Canada and the U.K. and show that we can meaningfully estimate bounds on the percentage of households who are better and worse off. Even though these bounds are typically only partially identified, the estimated bounds are almost always narrower than ten percentage points and often substantially narrower than that. This demonstrates how to operationalize our novel econometric methodology to conduct inference for counterfactuals.
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {The Deterministic Model}
We consider an econometrician who is studying a consumer's demand for $L$ goods. We assume an idealized environment suitable for partial equilibrium analysis, where the consumer's demand for these goods at different prices are observed, while the consumer's wealth and the prices of other goods are held fixed.\footnote{Under fairly standard (but strong) assumptions, changes to the external environment can be precisely justified by deflating the prices of the $L$ goods (see \hyperref[sec:deflprices]{Section (ref)}).}
Specifically, the econometrician collects a data set with a finite number of observations; each observation $t$ can be represented as $(p^t,x^t)$, where $p^t\in \mathbb{R}^L_{++}$ are the prices of the $L$ goods and $x^t\in \mathbb{R}^L_+$ is the bundle of those goods purchased by the consumer.\footnote{We postpone the discussion of discrete consumption spaces and nonlinear pricing to \hyperref[sec:nonlinearGAPP]{Section (ref)}.} We denote the data set by $\mathcal{D}:= \{(p^t,x^t)\}_{t=1}^T$. (We shall slightly abuse notation and use $T$ to refer both to the (finite) number of observations and to the set $\{1,\dots,T\}$; similarly, $L$ could denote both the number, and the set, of commodities.)
We begin with a basic question: given ${\mathcal D}$, can the econometrician sign the welfare impact of a price change from $p^t$ to $p^{t'}$? Perhaps the most intuitive welfare comparison that can be made in this setting is as follows: if at prices $p^{t'}$, the econometrician finds that $p^{t'}\cdot x^t< p^t\cdot x^t$ then he may conclude that the agent is better off at the price vector $p^{t'}$ compared to $p^t$. This is because, at the price $p^{t'}$ the consumer can, if she wishes, buy the bundle bought at $p^t$ and she would still have money left over to buy other things, so she must be strictly better off at $p^{t'}$. This ranking is eminently sensible, but can it lead to inconsistencies?
This example shows that for an econometrician to be able to consistently compare the consumer's welfare at different prices, some restriction has to be imposed on the data set. To be precise, define the binary relations $\succeq_p$ and $\succ_p$ on ${\mathcal P}:=\{p^t\}_{t\in T}$, that is, the set of price vectors observed in $\mathcal D$, in the following manner: $$p^{t'} \succeq_p (\succ_p) p^t \text{ if } p^{t'}\cdot x^{t}\leq (<)p^{t}\cdot x^{t}.$$ We say that price $p^{t'}$ is {\em directly (strictly) revealed preferred} to $p^t$ if $p^{t'} \succeq_p (\succ_p) p^t$, that is, whenever the bundle $x^t$ is (strictly) cheaper at prices $p^{t'}$ than at prices $p^t$. We denote the transitive closure of $\succeq_p$ by $\succeq_p^*$, that is, for $p^{t'}$ and $p^t$ in $\mathcal P$, we have $p^{t'}\succeq_p^* p^t$ if there are $t_1$, $t_2$,...,$t_N$ in $T$ such that $p^{t'}\succeq_p p^{t_1}$, $p^{t_1}\succeq_p p^{t_2}$,..., $p^{t_{N-1}}\succeq_p p^{t_N}$, and $p^{t_N}\succeq_p p^{t}$; in this case we say that $p^{t'}$ is {\em revealed preferred} to $p^t$. If anywhere along this sequence, it is possible to replace $\succeq_p$ with $\succ_p$ then we say that $p^{t'}$ is {\em revealed strictly preferred} to $p^t$ and denote that relation by $p^{t'}\succ^*_p p^{t}$.\footnote{Notice that it makes sense to write $\hat p\succeq_p p^t$ even if $\hat p$ is not in $\mathcal P$, since the demand at $\hat p$ is not needed in the definition revealed preference. Similarly, it is possible to define $\hat p\succ_p p^t$ and the transitive extensions $\hat p\succeq_p^* p^t$ and $\hat p\succ_p^* p^t$. This observation is useful later on, in Sections (ref) and (ref). .} The following restriction, which excludes circularity in the econometrician's assessment of the consumer's wellbeing at different prices, is a bare minimum condition to impose on $\mathcal{D}$.
This in turn leads naturally to the following question: if a consumer's observed demand behavior obeys GAPP, what could we say about her decision making procedure?
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{The Expenditure-Augmented Utility Model}
An expenditure-augmented utility function (or simply, an {\em augmented utility function}) is a function $U:\mathbb{R}^L_{+}\times \mathbb{R}_{-}\to\mathbb{R}$, where $U(x,-e)$ is the consumer's utility when she spends $e$ to purchase the bundle $x$. We require that $U(x,-e)$ is strictly increasing in the last argument (in other words, utility is strictly decreasing in expenditure), which captures the tradeoff the consumer faces between consuming $x$ and consuming other goods (outside the set $L$).
At a given price $p$, the consumer chooses a bundle $x$ to maximize $U(x,-p\cdot x)$. We denote the {\em indirect utility at price $p$} by
If the consumer's augmented utility maximization problem has a solution at every price vector $p\in\mathbb{R}_{++}^L$, then $V$ is also defined at those prices and this induces a reflexive, transitive, and complete preference over prices in $\mathbb{R}_{++}^L$.
A data set $\mathcal{D}=\{(p^t,x^t)\}_{t=1}^T$ is rationalized by an augmented utility function if there exists such a function $U:\mathbb{R}_+^{L}\times \mathbb{R}_{-}\to \mathbb{R}$ with
It is straightforward to see that GAPP is necessary for a data set to be rationalized by an augmented utility function. First, notice that if $p^{t'} \succeq_p p^t$, then $p^{t'} \cdot x^{t}\leq p^{t}\cdot x^t$, and so $$V(p^{t'})\geq U(x^t,-p^{t'}\cdot x^t)\geq U(x^t,-p^{t}\cdot x^t)=V(p^t).$$ Furthermore, $U(x^t,-p^{t'}\cdot x^t)> U(x^t,-p^{t}\cdot x^t)$ if $p^{t'} \succ_p p^t$, and in that case $V(p^{t'})> V(p^t)$. Suppose GAPP were not satisfied and there were two observations $t,t'\in T$ such that $p^{t'}\succeq^*_p p^{t}$ and $p^{t}\succ^*_p p^{t'}$. Then there would exist $t_1,t_2,\dots,t_N\in T$ such that $$V(p^{t'})\geq V(p^{t_1})\geq\cdots \geq V(p^{t_N}) \geq V(p^t) > V(p^{t'})$$ which is impossible.
Our main theoretical result, which we state next, also establishes the sufficiency of GAPP for rationalization. Moreover, the result states that whenever $\mathcal{D}$ can be rationalized, it can be rationalized by an augmented utility function $U$ with a list of properties that make it convenient for analysis.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{\hyperref[thm:Afriat]{Afriat's Theorem} and Proof of \hyperref[thm:GAPP]{Theorem (ref)}}
Before presenting the proof of \hyperref[thm:GAPP]{Theorem (ref)}, it is worth providing a short description of the standard theory of revealed preference and, \hyperref[thm:Afriat]{Afriat's Theorem}, its central result. This will be useful not just because we will invoke the result several times but also since it will serve as an important point of contrast for our axiom and results.
The standard theory due to afriat1967 is built formally on the same primitives as our model: a finite data set of prices and corresponding consumption bundles. Unlike our model however, it is assumed that the observed goods correspond to the universe of the consumer's consumption. Formally, a data set $\mathcal{D}$ is said to be rationalized by a utility function if there exists a locally nonsatiated\footnote{This means that at any bundle $x$ and open neighborhood of $x$, there is a bundle $x'$ in the neighborhood with strictly higher utility.} utility function $\widetilde{U}:\mathbb{R}_+^L\to \mathbb{R}$ such that
In words, this criterion asks whether there is a utility function defined over the $L$ observed goods such that the consumer is utility maximizing at every observation $t$ over the fixed budget $p^t\cdot x^t$ corresponding to the observed expenditure.
Of course, data sets (outside of laboratory data) almost never contain the universe of consumed goods and the consumer's true budget set is not observed, especially when one takes into account the possibility of borrowing and saving. Given this, when checking if a data set $\mathcal{D}$ can be rationalized in the sense of ((ref)), we are effectively testing whether the consumer is maximizing a sub-utility function $\widetilde U:\mathbb{R}^L_+\to\mathbb{R}$ defined specifically on those $L$ goods (or equivalently, has weakly separable preferences).
It should be clear that rationalization in the sense of ((ref)) is distinct from rationalization by an augmented utility function. The augmented utility model specifically takes into account the impact of the prices of these $L$ goods on the consumption of other goods; it is {\em necessarily} a partial equilibrium model, and designed for partial equilibrium welfare analysis of the type carried out in empirical industrial organization or public economics. An example is the study of the welfare impact of a sales tax levied on a subset of goods.
It is possible that a data set $\mathcal{D}$ can be rationalized in both senses, but that does not hold in general. The precise conditions needed for rationalization by a utility function are given by \hyperref[thm:Afriat]{Afriat's Theorem}, which we now describe.
{\em Revealed preference} in Afriat's setting is captured by two binary relations, $\succeq_x$ and $\succ_x$ which are defined on the set of chosen bundles observed in $\mathcal D$, that is, the set ${\mathcal X}:=\{x^t\}_{t\in T}$, as follows: $$x^{t'} \succeq_x (\succ_x)\, x^t \text{ if } p^{t'}\cdot x^{t'} \geq (>)\, p^{t'}\cdot x^{t}.$$ We say that the bundle $x^{t'}$ is {\em directly revealed (strictly) preferred} to $x^t$ if $x^{t'} \succeq_x (\succ_x)\, x^t$, that is, whenever the bundle $x^t$ is (strictly) cheaper at prices $p^{t'}$ than the bundle $x^{t'}$. This terminology is intuitive: if the agent is maximizing some locally nonsatiated utility function $\widetilde U:\mathbb{R}_+^L\to\mathbb{R}$, then $x^{t'} \succeq_x x^t$ ($x^{t'}\succ_x x^t$) must imply that $\widetilde U(x^{t'})\geq (>)\, \widetilde U(x^t)$.
We denote the transitive closure of $\succeq_x$ by $\succeq_x^*$, that is, for $x^{t'}$ and $x^t$ in $\mathcal X$, we have $x^{t'}\succeq_x^* x^t$ if there are $t_1$, $t_2$,...,$t_N$ in $T$ such that $x^{t'}\succeq_x x^{t_1}$, $x^{t_1}\succeq_x x^{t_2},\dots,x^{t_{N-1}}\succeq_x x^{t_N} \succeq_x x^{t}$, and $x^{t_N}\succeq_x x^{t}$; in this case, we say that $x^{t'}$ is {\em revealed preferred} to $x^t$. If anywhere along this sequence, it is possible to replace $\succeq_x$ with $\succ_x$ then we say that $x^{t'}$ is {\em revealed strictly preferred} to $x^t$ and denote that relation by $x^{t'}\succ^*_x x^{t}$. Clearly, if $\mathcal D$ is rationalizable by some locally nonsatiated utility function $\widetilde U$, then $x^{t'}\succeq^*_x (\succ^*_x)\, x^t$ implies that $\widetilde U(x^{t'})\geq (>)\, \widetilde U(x^t)$. This observation in turn implies that a necessary condition for rationalization by a utility function is that the revealed preference relation has no cycles.
The main insight of \hyperref[thm:Afriat]{Afriat's Theorem} is to show that this condition is also sufficient (the formal statement can be found in the online \hyperref[sec:afriat]{Appendix (ref)}).
Having described \hyperref[thm:Afriat]{Afriat's Theorem}, we are now in a position to prove \hyperref[thm:GAPP]{Theorem (ref)}.
{\sc Proof of \hyperref[thm:GAPP]{Theorem (ref)}.} We will show that $(2)\implies (3)$. We have already argued that $(1)\implies (2)$ and $(3)\implies (1)$ by definition.
Choose a number $M>\max_{t} p^t\cdot x^{t}$ and define the augmented data set $\widetilde{\mathcal{D}}=\{(p^t,1),(x^t,M-p^t\cdot x^t)\}_{t=1}^T$. This data set augments $\mathcal D$ since we have introduced an $L+1^{\text{th}}$ good, which we have priced at 1 across all observations, with the demand for this good equal to $M-p^t\cdot x^t$.
The crucial observation to make here is that $$(p^t,1) (x^t,M-p^t\cdot x^t) \geq (p^t,1) (x^{t'},M-p^{t'}\cdot x^{t'}) \text{ if and only if } p^{t'}\cdot x^{t'} \geq p^t\cdot x^{t'},$$ which means that $$(x^t,M-p^t\cdot x^t) \succeq_x (x^{t'},M-p^{t'}\cdot x^{t'}) \text{ if and only if } p^t\succeq_p p^{t'}.$$ Similarly, $$(p^t,1) (x^t,M-p^t\cdot x^t) > (p^t,1) (x^{t'},M-p^{t'}\cdot x^{t'}) \text{ if and only if } p^{t'}\cdot x^{t'} > p^t\cdot x^{t'},$$ and so $$(x^t,M-p^t\cdot x^t) \succ_x (x^{t'},M-p^{t'}\cdot x^{t'}) \text{ if and only if } p^t \succ_p p^{t'}.$$ Consequently, $\mathcal{D}$ satisfies GAPP if and only if $\widetilde{\mathcal{D}}$ satisfies GARP. Applying \hyperref[thm:Afriat]{Afriat's Theorem} when $\widetilde{\mathcal{D}}$ satisfies GARP, there is $\widetilde{U}:\mathbb{R}^{L+1}\rightarrow \mathbb{R}$ (notice that $\widetilde U$ is defined on $\mathbb{R}^{L+1}$ and not just $\mathbb{R}^{L+1}_{+}$; see Remark 3 in \hyperref[sec:afriat]{Appendix (ref)}) such that
The function $\widetilde U$ can be chosen to be strictly increasing, continuous, and concave, and the lower envelope of a finite set of affine functions. Clearly, the augmented utility function $\widebar U:\mathbb{R}^L_+\times \mathbb{R}_{-}\rightarrow \mathbb{R}$ defined by $\widebar U(x,-e):=\widetilde{U}(x,M-e)$ is strictly increasing in $(x,-e)$, continuous, concave and rationalizes $\mathcal{D}$.
Define $\widehat U:\mathbb{R}_+^L\times \mathbb{R}_{-}\rightarrow \mathbb{R}$ by
where $h:\mathbb{R}_+\rightarrow \mathbb{R}$ is a differentiable function satisfying $h(0)=0$, $h'(k)>0$, $h''(k)\geq 0$ for $k\in \mathbb{R}_+$, and $\lim_{k\rightarrow \infty} h'(k)=\infty$. (For example, $h(k)=k^3$.) Like $\widebar U$, the function $\widehat U$ is strictly increasing in $(x,-e)$, continuous and concave and $x^t$ solves $\max_{x\in \mathbb{R}^L_+} \widehat U(x,-p^t x)$ (because $\widehat U(x,-e)\leq \widebar U(x,-e)$ for all $(x,-e)$, and $\widehat U(x^t,-p^t\cdot x^t)=\widebar U(x^t,-p^t\cdot x^t)$). Furthermore, for every $p\in\mathbb{R}^L_{++}$, $\operatornamewithlimits{argmax}_{x\in X} \widehat U(x,-p\cdot x)$ is nonempty.\footnote{Choose a sequence $x^n\in \mathbb{R}^L_+$ such that $\widehat U(x^n,-p\cdot x^n)$ tends to $\sup_{x\in \mathbb{R}^L_+} \widehat U(x,-p\cdot x)$ (which we allow to be infinity). It is impossible for $p\cdot x^n\rightarrow \infty$ because the piecewise linearity of $U(x,-e)$ in $x$ and the assumption that $\lim_{k\to \infty}h'(k)\to\infty$ implies that $\widehat U(x^{n},-p\cdot x^n)\rightarrow -\infty$. So the sequence $p\cdot x^n$ is bounded, which in turn means that there is a subsequence of $x^n$ that converges to $x^\star\in \mathbb{R}^L_+$. By the continuity of $\widehat U$, we obtain $\widehat U(x^\star,-p\cdot x^\star)= \sup_{x\in \mathbb{R}^L_+} \widehat U(x,-p\cdot x)$.} $\blacksquare$
We end this section by noting that GARP imposes testable restriction distinct from GAPP. This is immediate from \hyperref[eg:GARPnotGAPP]{Example (ref)} and can be seen from \hyperref[fig:GARPnotGAPP_A]{Figure (ref)} which plots not just the observed consumption bundles but also the corresponding budget sets (derived from the observed prices and expenditures).
As we argued, GAPP does not hold in this example but, since the budget sets do not even cross, it is immediate to conclude that GARP does. We defer the description of the exact relation between the two criteria to \hyperref[sec:GAPPvsGARP]{Section (ref)}.
From this point onwards, when we refer to `rationalization' without additional qualifiers, we shall mean rationalization by an augmented utility function, that is, in the sense given by ((ref)) rather than in the sense given by ((ref)).
Up to now, we have motivated our model by showing that it is the utility representation of a basic axiom requiring consistent price comparisons. In the next two subsections, we provide direct motivation for the augmented utility function itself by arguing that it contains, as special cases, several distinct (standard and behavioral) preference-modeling approaches.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{`Standard' consumer theory and the augmented utility function}
Perhaps the clearest motivation for our model is to think of it as a generalization of the quasilinear utility model, in which the consumer derives utility $\widetilde{U}(x)$ from the bundle $x$ and maximizes utility net of expenditure, that is, she chooses $x$ to maximize
where $e=p\cdot x$. There is a familiar textbook way of justifying this objective function by fitting it within the constrained optimization model of standard consumer theory. This is to think of the consumer as having a utility function $\widebar U$ defined over $L+1$ goods, with the last `outside' good entering additively and linearly into the utility function, so that $\widebar U(x,z)=\widetilde U(x)+z$. Assuming that the consumer has a total wealth of $W$, the utility of purchasing a bundle $x\in\mathbb{R}^L_+$ is then $$\widebar U(x,W-p\cdot x)=\widetilde U(x)-p\cdot x+W.$$ Ignoring boundary issues, the consumer is effectively maximizing ((ref)).
Even though the quasilinear model is widely used in partial equilibrium analysis, it is well known that the complete absence of income effects makes it unsuitable for certain empirical applications. For this reason, it is also common to remove the linear structure on $\widebar U$ while retaining the assumption that all outside consumption opportunities can be represented by a {\em single} outside good; this is true, for example, in the literature on modeling the demand for differentiated goods.\footnote{For example, in berry1995 and in nevo2000, $\widebar U$ is additively separable between the $L$ goods and the outside good; in the former, the utility of consuming $y$ units of the outside good is $\alpha \ln y$, for some $\alpha>0$, whereas in the latter it is $\alpha y$ (in other words, $\widebar U$ is quasilinear). In bhattacharya2015, $\widebar U$ is allowed to be a general function defined on $L+1$ goods. In models of differentiated goods, the consumption space is typically assumed to be discrete rather than $\mathbb{R}^L_+$, but the augmented utility model is still applicable in that context (see \hyperref[sec:nonlinearGAPP]{Section (ref)}). } In this case, the utility of purchasing a bundle $x\in\mathbb{R}^L_+$ is $\widebar U(x,W-p\cdot x)$; notice that, provided $W$ is fixed, we can think of the consumer as maximizing an augmented utility function: simply let $U(x,-e)=\widebar U(x,W-e)$.
Obviously, a consumer's outside consumption opportunities would in reality involve more than one good, and the prices of those outside goods could change as well. Within the familiar constrained-optimal model of consumer theory, there are known conditions that justify the representation of those consumption opportunities by a representative good (with its corresponding price index). This is explained in detail in \hyperref[sec:deflprices]{Section (ref)}.
Finally, it is worth mentioning that the augmented utility function could also capture, as a special case, quasilinear utility maximization subject to certain constraints. One such example is consumption with a subsistence constraint, which we describe in more detail in the empirical application in \hyperref[sec:progresa]{Section (ref)}. Loosely speaking, we can capture constraints on $(x,-e)$ with an augmented utility $U(x,-e)$ that assigns very low values at $(x,-e)$ that violate the constraint.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Behavioral preferences captured by the augmented-utility model}
The central feature of the augmented utility model is that consumers experience disutility from expenditure. As we explained in the previous subsection, this disutility could be interpreted in a purely opportunity cost sense -- more expenditure on the consumed goods imply less money available for other goods. In this understanding, the augmented utility function is a reduced form of a broader 'true' utility function defined on all goods.
However, it is also reasonable to think of the augmented utility function in another way: that the consumer has -- directly -- a preference over bundles of the observed $L$ goods and their associated expenditure, which she has developed as a way of guiding her purchasing decisions. Thus it is the basic object of analysis and not the reduced form of something more fundamental. This understanding of choice behavior is exploited in the behavioral economics literature and the following quote from prelec1998 is effectively a description of the augmented utility function:
In this understanding, the disutility of expenditure is still related to opportunity cost, but the relationship is more flexible than what is permitted in a classical framework.\footnote{Another paper that spells out remarkably clearly this approach to modelling consumer decisions is friedman2015, though the authors primarily have in mind the quasilinear utility model.}
In the Introduction, we described one example of behavioral preferences (reference-dependent preferences) that could be captured by an augmented utility function. In the remainder of this section, we describe how our model relates to two other prominent themes in the behavioral literature.
The public economics literature following chettylooney2009 has observed that consumers often misperceive prices: in their context, shoppers at grocery stores do not internalize the price effect of taxes. This literature is summarized in a recent survey by gabaix2019 who argues that many behavioral biases often take the form of inattention. Our model naturally captures a version of the inattention to prices discussed in bordalo2013 and gabaix2014. Here a consumer faced with a price $p$ perceives the expenditure associated with a bundle $x$ as $f(x,p\cdot x)$, where $f$ is increasing in the true expenditure $p\cdot x$ and could potentially depend on $x$. With this misperception, and assuming that the consumer has a quasilinear preference, she then chooses $x\in\mathbb{R}^L_+$ to maximize
A special case of this model is where a consumer has a default price $p^d$ and misperceives the actual price $p$ to be $a p + (1-a) p^d$ where $a\in[0,1]$ is the `attention parameter.' The perceived expenditure is then $f(x,p\cdot x)=ap\cdot x+(1-a)p^d\cdot x$. More generally, the model accommodates $f(x,p\cdot x)=a(x)p\cdot x+(1-a(x))p^d\cdot x$, where the attention parameter $a(x)\in [0,1]$ varies across bundles.\footnote{This formulation of perceived expenditure is more general than gabaix2014 in that it allows the attention parameter $a$ to depend on $x$ but is less general in that the parameter does not vary across goods.} This is a natural extension since, among other things, it allows a consumer to be more attentive to her actual expenditure if she is purchasing large bundles compared to small ones (so that $a(x)$ tends to 1 when $x$ is large). Yet another possibility is that the consumer is not completely sensitive to every dollar increase in expenditure but pays more attention only when certain thresholds are crossed; this would correspond to the case where $f$ depends only on the expenditure $e=p\cdot x$ and has the shape of a step function of expenditure.
Clearly, inattention as modeled by ((ref)) is an instance where the agent has an augmented utility function, even though it will typically not be quasilinear (in actual expenditure).
It is also worth mentioning that using an augmented utility function (such as (ref)) to capture price inattention is particularly apt because, as gabaix2014 notes, the numeraire serves as “the shock absorber that adjusts to the budget constraint.” The alternative is to model the consumer as having {\em both} price misperception over a given set of goods {\em and} a budget on those goods that must be satisfied, which inevitably leads to the added complication of modelling how the agent adjusts her intended demand when she realizes it violates (because prices are misperceived) the budget constraint at the true prices.\footnote{gabaix2014 proposes one way to deal with this issue.}
As we discussed in \hyperref[sec:AUQL]{Section (ref)}, a common approach to partial equilibrium analysis is to add a numeraire as an additional good and assume that the agent has a (standard) utility function and budget set defined on the $L+1$ goods, with price and income information used to determine the level of the numeraire consumed. Of course such an approach could only work when income information is available and that is not always the case.\footnote{Several widely used data sets, such as supermarket scanner panel data, that contain rich information on purchases, do not have accurate measures of income. Here, income information is typically the category (income ranges) that households self report when applying for loyalty cards (and so the information becomes out of date).} Even when this information is available, it is strictly speaking not the right value to use as the global budget if the consumer can save and borrow to a significant degree (as acknowledged, for example, in hausman2016). More generally, figuring out what really constitutes `the budget' is not always straightforward, even in a classical setting.
Regularities highlighted by behavioral economists add a further wrinkle to the concept of a budget. It has been widely observed that households do not always treat money as fungible and instead create separate accounts for various categories of goods thaler1999. This is not only true for consumption decisions (see, for instance, hastings2013, hastings2018) but also for savings decisions, which is why consumers often save more when they have access to commitment savings options (important theoretical and empirical contributions are amador2006 and feldman2010, dupas2013 respectively).
Now consider a researcher who is trying to model the demand for a set of $L$ goods which form a subset of all the goods consumed by an agent. If mental accounting effects are important, the researcher will have to allow for the fact that he cannot observe how goods are categorized by the agent, nor does he know what really constitutes the mental budget from which the agent is drawing her expenditure (on the $L$ observed goods and their perceived alternatives). In this situation, the augmented utility framework provides a natural way to model the demand for those $L$ goods: it is consistent with constrained utility maximization incorporating an outside good (see \hyperref[sec:AUQL]{Section (ref)}) but {\em does not require} the researcher to take a stand on the (unobserved, mental) budget from which the agent is drawing her expenditure.\footnote{Here we are assuming that the data $\mathcal{D}=\{(p^t,x^t)\}_{t\in T}$ are collected over a period where the mental budget for the $L$ observed goods and their alternatives is stable. Changing mental budgets would manifest itself as violations of GAPP (see \hyperref[ex:mentalbudget]{Example (ref)}).}
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {Properties of the augmented utility model}
In this section, we explore various aspects of the augmented utility model, beginning with a discussion of the relationship between GAPP and GARP. We then go on to discuss welfare analysis in the augmented utility framework. Since one would not expect data sets to be completely consistent with the augmented utility model, we discuss how departures from GAPP could be measured. Lastly, we discuss how prices could be deflated in this model to account for general changes in the price level.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Comparing GAPP and GARP}
Recall that \hyperref[eg:GARPnotGAPP]{Example (ref)} in \hyperref[sec:det_model]{Section (ref)} is an example of a data set that obeys GARP but fails GAPP. We now present an example of a data set that satisfies GAPP but fails GARP.
While GAPP and GARP are not in general the same conditions, they coincide in any data set where $p^t\cdot x^t=1$ for all $t\in T$. This is because $x^t\succeq_x(\succ_x) x^{t'}$ if and only if $p^t\succeq_p (\succ_p) p^{t'}$ since both conditions are equivalent to $1\geq(>)\, p^t\cdot x^{t'}$. Given a data set ${\mathcal D}=\{(p^t,x^t)\}_{t=1}^T$, we define the {\em iso-expenditure version of $\mathcal D$} as another data set $\breve {\mathcal D}:=\left\{(p^t,\breve{x}^t)\right\}_{t=1}^T$, such that $\breve{x}^t=x^t/(p^t\cdot x^t)$. This new data set has the feature that $p^t\cdot \breve{x}^t=1$ for all $t\in T$. Notice that the revealed price preference relations $\succeq_p$, $\succ_p$ remain unchanged when consumption bundles are scaled. Thus a data set obeys GAPP if and only if its iso-expenditure version obeys GAPP, which in this case is equivalent to GARP.\footnote{There is an analogous `GARP-version' of \hyperref[prop:GAPPscaling]{Proposition (ref)} and that this observation (or some close variation of it) has been exploited before in the literature (see, for example, sakai1977). Suppose ${\mathcal D}=\{(p^t,x^t)\}_{t=1}^T$ obeys GARP. Then GARP holds even if each observed price vector $p^t$ is arbitrarily scaled. In particular, $\mathcal D$ obeys GARP if and only if $\hat{\mathcal D}=\{(\hat p^t,x^t)\}_{t\in T}$, where $\hat {p}^t=p^t/(p^t\cdot x^t)$, obeys GARP (equivalently, GAPP) since $\hat p^t\cdot x^t=1$ for all $t\in T$. The latter perspective is useful because it highlights the possibility of applying \hyperref[thm:Afriat]{Afriat's Theorem} on $\hat{\mathcal D}$, {\em in the space of prices} (in other words, with the roles of prices and bundles reversed). This immediately gives us a different, `dual' rationalization of $\mathcal D$ in terms of indirect utility, that is, there is a continuous, strictly decreasing, and convex function $\widetilde{V}:\mathbb{R}_{++}^L\to\mathbb{R}$ such that $\hat p^t\in {\arg\min}_{\{p\in\mathbb{R}_{++}^L\,:\,p\cdot x^t\geq 1\}}\widetilde{V}(p)$. For an application of this observation, see brown2000.} The next proposition gives a more detailed statement of these observations.
{\bf Proof.}\, Notice that $$p^t\cdot \frac{x^t}{p^t\cdot x^t} \geq p^t\cdot \frac{x^{t'}}{p^{t'}\cdot x^{t'}} \, \iff \, p^{t'}\cdot x^{t'} \geq p^t\cdot x^{t'}.$$ The left side of the equivalence says that $\breve{x}^t\succeq_x \breve{x}^{t'}$ while the right side says that $p^t\succeq_p p^{t'}$. This implies (1) since $\succeq_p^*$ and $\succeq_x^*$ are the transitive closures of $\succeq_p$ and $\succeq_x$ respectively. Similarly, it follows from $$p^t \cdot \frac{x^t}{p^t \cdot x^t} > p^t \cdot \frac{x^{t'}}{p^{t'}\cdot x^{t'}} \, \iff \, p^{t'}\cdot x^{t'} > p^t\cdot x^{t'}$$ that $\breve{x}^t\succ_x \breve{x}^{t'}$ if and only if $p^t\succ_p p^{t'}$, which leads to (2). The claims (1) and (2) together guarantee that there is a sequence of observations in $\mathcal D$ that lead to a GAPP violation if and only if the analogous sequence in $\breve{\mathcal D}$ lead to a GARP violation. $\blacksquare$
As an illustration, compare the data sets in \hyperref[fig:GARPnotGAPP_A]{Figure (ref)} and \hyperref[fig:GAPPnotGARP_A]{Figure (ref)} to the iso-expenditure data sets in \hyperref[fig:GARPnotGAPP_B]{Figure (ref)} and \hyperref[fig:GAPPnotGARP_B]{Figure (ref)}. It can be clearly observed that the iso-expenditure data in \hyperref[fig:GARPnotGAPP_B]{Figure (ref)} contains a GARP violation (which implies it does not satisfy GAPP) whereas the data in \hyperref[fig:GAPPnotGARP_B]{Figure (ref)} does not violate GARP (and, hence, satisfies GAPP).
A consequence of \hyperref[prop:GAPPscaling]{Proposition (ref)} is that the augmented utility model can be tested in two ways: we can either test GAPP directly or we can test GARP on its iso-expenditure version. If we are simply interested in testing GAPP on a single-agent data set $\mathcal D$, normalization brings no advantage: the test is computationally straightforward in either case and involves the construction of their (respective) revealed preference relations and checking for acyclicity. However, as we shall see in \hyperref[sec:RAUMX]{Section (ref)}, iso-expenditure scaling plays an important role in the test we develop (on repeated cross-sectional demand data) for the random utility version of the augmented utility model.
While GARP and GAPP are distinct properties, they are not mutually exclusive and it is possible for a data set to satisfy both. For example, if $\mathcal{D}=\{(p^t,x^t)\}_{t=1}^T$ is collected from a consumer who is maximizing a quasilinear augmented utility function, then it will satisfy both GAPP and GARP.\footnote{When $U$ has the form ((ref)), $x^t$ maximizes $U(x,-p^t\cdot x)$ only if $x^t$ maximizes $\widetilde U(x)$ in $\{x\in\mathbb{R}^L_+:p^t\cdot x\leq p^t\cdot x^t\}$. Thus $\mathcal{D}$ must also obey GARP. A broader class of augmented utility functions that satisfy both GAPP and GARP is given in \hyperref[sec:modGPGP]{Section (ref)}.} When both properties are satisfied, then an analyst could make use of either property when making predictions of demand at an out-of-sample price; the two properties will then typically lead to different set predictions. We discuss this in greater detail in \hyperref[sec:GAPPGARP]{Section (ref)} of the online appendix, which also contains more discussion of the relationship between revealed preferences under GAPP and under GARP.
In light of \hyperref[prop:GAPPscaling]{Proposition (ref)} and the fact that the revealed price preference relation is not affected by scaling consumption bundles, it is natural to wonder about the relationship between the testable implications of the augmented-utility model and the constrained-optimization model (as in ((ref))) restricted to homothetic preferences. A data set that can be rationalized in the latter sense\footnote{For the precise characterization, see varian1983.} will have the feature that it must satisfy GARP for any arbitrary scaling of consumption bundles and thus will satisfy GAPP. By contrast, a data set that satisfies GAPP must only satisfy GARP for the particular scaling that equalizes expenditure across observations. In other words, GAPP is a less stringent property; that it is {\em strictly} less stringent is clear from \hyperref[eg:GAPPnotGARP]{Example (ref)}, which satisfies GAPP but violates GARP and therefore cannot be rationalized in Afriat's sense (as given by ((ref))) for any locally nonsatiated preference, let alone a homothetic preference.\footnote{\hyperref[ex:yat]{Example (ref)} in the online appendix contrasts demand predictions using the augmented utility model and the constrained-optimization model (both with and without imposing homotheticity on the preference).}
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Preference over Prices}
We know from \hyperref[thm:GAPP]{Theorem (ref)} that if $\mathcal D$ obeys GAPP then it can be rationalized by an augmented utility function with an indirect utility that is defined at all price vectors in $\mathbb{R}_{++}^L$. It is straightforward to check that any indirect utility function $V$ as defined by ((ref)) has the following two properties:
Any rationalizable data set $\mathcal D$ could potentially be rationalized by many augmented utility functions, with each one leading to a different indirect utility function. We denote this set of indirect utility functions by $\mathbf{V}({\mathcal D})$. We have already observed that if $p^{t}\succeq_p^* (\succ_p^*)\, p^{t'}$ then $V(p^{t})\geq (>)\, V(p^{t'})$ for any $V\in\mathbf{V}({\mathcal D})$; in other words, the conclusion that the consumer prefers the prices $p^t$ to $p^{t'}$ is {\em nonparametric} in the sense that it is independent of the precise augmented utility function used to rationalize $\mathcal D$. The next result (proved in \hyperref[sec:pricewelfare]{Appendix (ref)}) says that, without further information on the augmented utility function, this is {\em all} the information on the consumer's preference over prices in $\mathcal P$ that we can glean from the data. Thus, in our nonparametric setting, the revealed price preference relation contains the most detailed information for welfare comparisons.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Compensation for a price change}
In standard consumer theory, the compensating and equivalent variations are two ways of quantifying the welfare impact of a price change (see mascolell1995, Chapter 3.I). We now argue that analogues exist for the augmented utility model and that bounds for them can be recovered from the data.
Let $U$ be the consumer's augmented utility function. Suppose that the initial price is $p^{t_1}$ and it changes to $p^{t_2}$, leading to a change in consumption from $x^{t_1}$ to $x^{t_2}$. Then we can find $\mu_c$ such that
Note that $\mu_c$ is unique since $U$ is strictly increasing in the last argument. We could think of $\mu_c$ as the lump sum transferred {\em from} the consumer (if it is positive) or {\em to} the consumer (if it is negative) after the price change that will make her just indifferent between the situation before and after the change.
Suppose we interpret $U$ as arising from an overall utility function $\widetilde{U}(x,z)$ (that depends on the observed goods $x$ and the level $z$ of an outside good), given the consumer's wealth of $M$, so that $U(x,-e)=\widetilde{U}(x,M-e)$. Since $\mu_c$ solves ((ref)), it will also satisfy $$\max\nolimits_{\{x\in \mathbb{R}_+^L:\; p^{t_2}\cdot x\leq M-\mu_c\}} \widetilde{U}(x,(M-\mu_c)-p^{t_2}\cdot x)\; = \; \widetilde{U}(x^{t_1},M-p^{t_1}\cdot x^{t_1}).$$ In other words, $\mu_c$ is the reduction in total wealth that will leave the consumer's overall utility at $p^{t_2}$ the same as it was at $p^{t_1}$. Thus, with this particular interpretation of the augmented utility function, $\mu_c$ coincides with what is called the {\em compensating variation} in standard consumer theory. For this reason, we shall also refer to $\mu_c$, defined by ((ref)), as the compensating variation.
Pushing the analogy further, it is possible to use the compensating variation in our model in the same way it is typically used. For example, the price change from $p^{t_1}$ to $p^{t_2}$ may benefit some consumers while hurting others. The Kaldor criterion would deem this change an overall improvement if the sum of the compensating variations across all consumers is positive since it guarantees that those who benefit from the price change could, in principle, compensate the losers and still be better off.
In a similar way, we can define the {\em equivalent variation} as the value $\mu_e$ that solves
If $U(x,-e)=\widetilde{U}(x,M-e)$ then $\mu_e$ also solves $$\max\nolimits_{\{x\in \mathbb{R}_+^L: \; p^{t_2}\cdot x\leq M+\mu_e\}} \widetilde{U}(x,(M+\mu_e)-p^{t_1}\cdot x)=\widetilde{U}(x^{t_2},M-p^{t_2}\cdot x^{t_2}).$$ In other words, $\mu_e$ coincides with the equivalent variation as it is usually defined.
Now suppose a data set $\mathcal{D}$ obeys GAPP and contains the observation $(p^{t_1},x^{t_1})$. What can we say about the compensating variation of a price change from $p^{t_1}$ to $p^{t_2}$ (where the latter may or may not be a price observed in $\mathcal{D}$)? There will typically be a range of these values since there is more than one augmented utility function that rationalizes $\mathcal{D}$. Nonetheless, it is possible to obtain a tight lower bound for the set of possible compensating variation values. Formally, this is given by $$\inf\{\mu_c:\mbox{$\mu_c$ solves (\ref{CV}) for some augmented utility function $U$ that rationalizes $\mathcal{D}$}\}.$$ Abusing terminology somewhat, we shall denote this term simply by $\inf({\mu}_{c})$.
We now describe how to compute this bound.\footnote{We leave the reader to carry out the analogous exercise for the equivalent variation.} Let $S\subset T$ be the set of observations such that $s\in S$ if $p^{s}\succeq^*_p p^{t_1}$. This set is nonempty since it contains $p^{t_1}$ itself. For each $s\in S$, there is $m_c^s$ such that
We claim that for any $U$ that rationalizes $\mathcal{D}$, the compensating variation $\mu_c\geq m_c^s$. This is because if $m<m_c^s$, then $m\neq \mu_c$ for any utility function rationalizing $\mathcal{D}$. Indeed,
Thus $\inf(\mu_c)\geq m_c^s$ for all $s\in S$. In fact, it is possible to obtain a stronger conclusion:
Since the right side of this equation can be easily computed from the data, we have found a practical way of calculating $\inf(\mu_c)$.
Notice that if $p^{t_2}$ is revealed preferred to $p^{t_1}$ (equivalently, that there is $s'\in S$ such that $m_c^{s'}\geq 0$),\footnote{Recall that $p^{t_2}\succeq_p^* p^{t_1}$ makes sense even if $p^{t_2}$ is not observed in the data set; see footnote (ref).} then $\inf(\mu_c)\geq 0$; in other words, at $p=p^{t_2}$, a lump sum {\em tax} of $\inf(\mu_c)$ will leave the agent no worse off than at $t_1$ and potentially better off. On the other hand, if $p^{t_2}$ is {\em not} revealed preferred to $p^{t_1}$, that is, for every $s\in S$, we have $m_c^s< 0$, then $\inf(\mu_c)< 0$; in other words, at $p=p^{t_2}$, a lump sum {\em transfer} of $\inf(\mu_c)$ to the agent will leave the agent no worse off than at $t_1$ and potentially better off.
We provide a fuller discussion on the compensating variation, including a proof of ((ref)), in \hyperref[sec:morecomvar]{Appendix (ref)}.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Measuring departures from rationality}
Empirical studies that apply \hyperref[thm:Afriat]{Afriat's Theorem} frequently find GARP violations. A common way of measuring the {\em extent} of such violations is to compute the {\em critical cost efficiency index} afriat1973. This refers to the largest $e\in (0,1]$ such that $\mathcal{D}$ can be rationalized in the following sense: there is a locally-nonsatiated utility function $\widetilde{U}$ such that $\widetilde{U} (x^t)\geq \widetilde{U}(x)$ for all $x$ in the `shrunken' budget set $B^t_{e}=\left\{x\in\mathbb{R}_+^L:\; p^t \cdot x\leq ep^t \cdot x^t\right\}$. Rationality is imperfect if $e<1$ since the consumer behaves as though she ignores bundles $x'$ that satisfy $e p^{t}\cdot x^{ t}<p^{t}\cdot x'\leq p^{t}\cdot x^{t}$ and, there could be some observation $\hat t$ and bundle $x'$ in this range for which $\widetilde{U}(x')> \widetilde{U}(x^{\hat t})$. Importantly, the calculation of the critical cost efficiency index is straightforward and is facilitated by a modified version of GARP.
There is a similar way of measuring the extent to which a data set $\mathcal{D}$ fails to be rationalized by an augmented utility function. For a given $\vartheta\in (0,1]$, there is a weaker version of the GAPP test that allows us to determine whether there is an expenditure-augmented utility $U:\mathbb{R}_+^{L}\times \mathbb{R}_{-}\to \mathbb{R}$ such that, at each observation $t$, $$U(x^t,-p^t\cdot x^t)\geq U(x,-\vartheta^{-1}p^t\cdot x)\:\mbox{ for all $x\in\mathbb{R}_+^L$.}$$ If there is, we say that $\mathcal{D}$ is $\vartheta$-rationalized by an augmented utility function. Notice that if $\mathcal{D}$ can be $\vartheta$-rationalized then it can be $\vartheta'$-rationalized for any $\vartheta'<\vartheta$, since $U$ is strictly decreasing in expenditure. The consumer who is $\vartheta$-rational (for $\vartheta<1$) may have only limited or bounded rationality in the sense that there could be a bundle $x'$ and an observation $\hat t$ such that $$U(x^{\hat t},-p^{\hat t}\cdot x')>U(x^{\hat t},-p^{\hat t}\cdot x^{\hat t})\geq U(x',-\vartheta^{-1}p^{\hat t}\cdot x').$$ In other words, the consumer fails to recognize that bundle $x'$ is superior to $x^{\hat t}$ at $t=\hat t$ because she has inflated (by $\vartheta^{-1}$) the expenditure of purchasing $x'$. Any data set can be $\vartheta$-rationalized for some $\vartheta\in (0,1]$ and the supremum $\vartheta^*$ over these values provides a natural measure of rationality which we shall refer to as the {\em rationality index}.
The following proposition establishes a connection of our rationality index with the critical cost efficiency index.
A consequence of this result is that the rationality index inherits the ease of computation of the critical cost efficiency index. In \hyperref[sec:rat-indices]{Appendix (ref)}, we provide instructions on this computation including in more general environments with nonlinear prices. The proof of \hyperref[prop:GAPP-GARP]{Proposition (ref)} can be found in \hyperref[sec:thetaGPGP]{Appendix (ref)}.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Deflating prices}
When the data set $\mathcal{D}=\{(p^t,x^t)\}_{t=1}^T$ is collected over an extended period, it is possible that there are changes in the prices of all goods, including goods outside the ones observed. Thus the nominal value of expenditure may no longer be an accurate measure of the opportunity cost of expenditure. A simple way of taking this into account is to deflate the prices of the $L$ goods with a general price index. In other words, one could check if $\widetilde{\mathcal{D}}=\{(p^t/k^t,x^t)\}_{t=1}^T$ obeys GAPP, where $k^t\in \mathbb{R}_{++}$ is an index of the general price level. If it does, it would mean that there is an augmented utility function $U$ that rationalizes the data after deflation; in other words, $$x^t\in\operatornamewithlimits{argmax}_{x\in\mathbb{R}^L_+}U\left(x,-\frac{p^t\cdot x}{k^t}\right)\:\mbox{ for all $t\in T$.}$$
This simple way of accounting for general price changes could be precisely justified when the augmented utility function is the reduced form of a larger constrained optimization problem. Indeed, suppose that the consumer is maximizing an overall utility $\widetilde{U}(x,y)$ that depends both on the observed bundle $x$ and on a bundle $y$ of other goods, subject to a global budget of $M$. Formally, the consumer maximizes $\widetilde{U}(x,y)$ subject to $p\cdot x+q\cdot y\leq M$, where $q$ are the prices of goods $y$. Keeping $q$ and $M$ fixed, $U(x,-e)$ is defined as the greatest overall utility the consumer can achieve by choosing $y$ optimally, subject to expenditure $M-e$ and conditional on consuming $x$, that is,
At the prices $p^t$ for the observed goods and $q$ for the outside goods, the consumer chooses a bundle $(x^t,y^t)$ to maximize $\widetilde{U}$ subject to $p^t\cdot x+q\cdot y\leq M$. Then $\mathcal{D}=\{(p^t,x^t)\}_{t=1}^T$ will obey GAPP, since $x^t$ maximizes $U(x,-p^t\cdot x)$, with $U$ as defined by ((ref)).
Now suppose that the prices of the other goods are changing. Consider the simplest case where these prices move up or down proportionately, so they are $k^t q$ at observation $t$, for some scalar $k^t>0$. Furthermore, assume that the agent's global budget at $t$ also increases by a factor $k^t$, which means that the consumer's nominal wealth is keeping pace with price inflation. Then at observation $t$, the consumer maximizes $\widetilde{U}(x,y)$ subject to $(x,y)$ obeying $$p^t\cdot x+k^t q\cdot y\leq k^t M.$$ Dividing this inequality by $k^t$, we see that the consumer's choice is identical to the case where the price of the observed goods is $p^t/k^t$, with external prices and total wealth constant at $q$ and $M$ respectively. Therefore, the data set with deflated prices, $\widetilde{\mathcal{D}}=\{(p^t/k^t,x^t)\}_{t=1}^T$ obeys GAPP.
In the case where the relative prices of the outside goods change, it is still possible to derive a price index which ensures that GAPP holds after deflating $p^t$, but this requires stronger assumptions on $\widetilde{U}$. We discuss this in detail in \hyperref[sec:indices]{Appendix (ref)}.
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {General consumption spaces and nonlinear pricing}
So far we have assumed that the consumption space is $\mathbb{R}^L_+$ and that prices are linear, but in fact neither feature is crucial to our main result. In this section, we assume that the space from which the consumer chooses her consumption is $X\subseteq \mathbb{R}_+^L$. We define a {\em price system} as a map $\psi: X\to\mathbb{R}_+$, where $\psi (x)$ is the cost of purchasing $x\in X$. Of course, a special case of a price system is $\psi (x)=p\cdot x$ but the more general formulation with $\psi$ allows for quantity discounts, bundle pricing and other pricing features that can be important in certain contexts (such as our empirical application in \hyperref[sec:progresa]{Section (ref)}).
We assume that both the price system and the bundle chosen by the consumer are observed. Formally, a data set is a collection $\mathcal{D}=\{(\psi^t,x^t)\}_{t=1}^T$. This data set is rationalized by an augmented utility function $U:X\times\mathbb{R}_{-}\to \mathbb{R}$ if
The notion of revealed preference over prices can be extended to a revealed preference over price systems. We say that $\psi^{t'}$ is {\em directly revealed preferred} ({\em directly revealed strictly preferred}) to $\psi^t$ if $\psi^{t'} (x^{t})\leq (<)\psi^{t}(x^{t})$; we denote this by $\psi^{t'} \succeq_p (\succ_p) \psi^t$. We denote the transitive closure of $\succeq_p$ by $\succeq_p^*$, that is, $\psi^{t'}\succeq_p^* \psi^t$ if there are $t_1$, $t_2,\ldots,t_N$ in $T$ such that $\psi^{t'}\succeq_p \psi^{t_1}$, $\psi^{t_1}\succeq_p \psi^{t_2},\ldots,\psi^{t_{N-1}}\succeq_p \psi^{t_N}$, and $\psi^{t_N}\succeq_p \psi^{t}$; in this case we say that $\psi^{t'}$ is {\em revealed preferred} to $\psi^t$. If anywhere along this sequence, it is possible to replace $\succeq_p$ with $\succ_p$ then we denote that relation by $\psi^{t'}\succ^*_p \psi^{t}$ and say that $\psi^{t'}$ is {\em strictly revealed preferred} to $\psi^t$. It is straightforward to check that if, $\mathcal{D}$ can be rationalized by an augmented utility function, then it obeys the following generalization of GAPP to price systems:
The following theorem asserts that the converse is also true and that, under further conditions, we can guarantee that the data can be rationalized by an augmented utility function with additional properties. The proof of this result (in fact of a more general result allowing for errors) is in \hyperref[sec:ultimate]{Appendix (ref)}.
Remarks:\, (1)\, Note that condition (ii) is a weak assumption requiring that there be no arbitrarily large bundles with a bounded price. (2) By definition, an augmented utility function is strictly decreasing in expenditure, but in certain cases it may be natural to require $U$ to be strictly increasing in $x_K$ for some set $K$ (which can be empty). The theorem says that this is possible, so long as the price systems are also strictly increasing in $x_K$. (3) Lastly, the theorem guarantees that the domain of the augmented utility function can be larger than $X$. For reasons which will be clear later in this section, this is natural in certain applications. However when we say that $U$ rationalizes the data, we mean that ((ref)) holds and, in particular, $x^t$ need not be optimal in $Y$.
The literature on mental accounting has emphasized the possibility of actors in the economy manipulating the mental budgets of agents. The following example shows how a nonlinear GAPP test can be used to detect such phenomena.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Discrete consumption spaces}
Below are four instances where \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} could be applied.
(1)\, Suppose that the consumption space consists of $L$ goods of which the first $K$ can only be consumed in discrete quantities (as in the model of polisson2013, for example). The consumption space is then $X=\mathbb{N}^{K}\times \mathbb{R}^{L-K}_+$, where $\mathbb{N}$ is the set of natural numbers. \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} is applicable whether or not prices are linear. Suppose that each good has a price $p_i>0$. Since $X$ is closed and the price system $\psi^t(x)=p^t\cdot x$ is strictly increasing in $x$, \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} guarantees that, if GAPP holds, then there is a continuous augmented utility function that is strictly increasing in $x\in X$ and rationalizes $\mathcal D$.
(2)\, Another natural choice environment is one where the consumer is deciding on buying a subset of objects from a set with $L$ items. Then each subset could be represented as an element of $X=\{0,1\}^L$. For $x\in X$, the $\ell^{\text{th}}$ entry $x_{\ell}$ equals 1 if and only if the $\ell^{\text{th}}$ object is chosen. If only certain subsets are permissible, as in the case of discrete choice, then $X$ would be a strict subset of $\{0,1\}^L$. The price system $\psi$ gives the price of different bundles of goods. Let $e_{\ell}$ denote the vector with 1 in the $\ell^{\text{th}}$ entry and zero everywhere else. Then $\psi(e_{\ell})$ is the price of purchasing good ${\ell}$ alone. The price system is nonlinear if $\psi (x)\neq \sum_{\ell=1}^L x_{\ell}\psi(e_{\ell})$ for some $x\in X$.
(3)\, In empirical models of demand for differentiated goods, it is common to model each good as embodying a set of characteristics (see nevo2000). For example, if each good is a type of breakfast cereal, then the characteristics could be the calories, fiber content etc. Suppose that there are $L$ characteristics and let $Y_{\ell}\subseteq\mathbb{R}_+$ be the set of values that characteristic $\ell$ can take. Then, the characteristics space is $Y=\times_{\ell=1}^L Y_{\ell}$.\footnote{If characteristic $1$ naturally takes on continuous values (such as calories) then we let $Y_{\ell}=\mathbb{R}_+$. Characteristic 2 could be the brand. Suppose there are two brands, then $Y_2=\{1,2\}$, and so on.} There are $I$ goods, with good $i$ having characteristics $x^i\in Y$. Assuming (as is common in these models) that a consumer purchases only one good, the consumption space is $X=\{x^i\}_{i=1}^I$ and a price system $\psi:X\to\mathbb{R}_{++}$ is just a list of prices for the different goods.
In this context, it is natural to model the consumer with an augmented utility function defined on characteristics and expenditures $Y\times \mathbb{R}_{-}$, even though the products available to her are only those in $X$. Furthermore, among the characteristics, there could be those where higher values are unambiguously better, in which case the researcher could be interested in guaranteeing that utility is strictly increasing in those characteristics. \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} allows for these considerations. If $\mathcal D$ obeys GAPP then it can be rationalized by a continuous augmented utility function $U:Y\times\mathbb{R}_{-}\to\mathbb{R}$. Additionally, for a set of characteristics $K\subseteq L$, one could guarantee that $U(y,-e)$ is strictly increasing in $y_K$ so long as $\psi^t(x)$ is strictly increasing in $x_K$, for all $t$.
In models of differentiated goods, it is also common to allow for the introduction of new goods and for changes to a product's characteristics.\footnote{These changes could be substantive (for example, a change to a breakfast cereal formula) or it could be a change in advertising expenditure that serves as a proxy for a change in a product's public profile.} Obviously, changes to a product's characteristics could potentially lead to a change in the product's utility which, unless taken into account by the test, could lead to a spurious rejection of augmented utility-maximization. In formal terms, these changes can be captured by allowing the set of alternatives to depend on $t$; in \hyperref[sec:differentiatedgoods]{Section (ref)}, we explain how it is possible to modify the GAPP test in \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} to account for changes of this type.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Characteristics models with continuous consumption spaces}
We assume that the space of characteristics is $Y=\mathbb{R}^L_+$, with each product $i$ represented by a vector of characteristics $x^i\in Y$. We allow these goods to be bought in bundles, so the consumption space is the convex cone $X$ generated by $\{x^i\}_{i=1}^I$.\footnote{For a GARP-based test of a model of this type, see blow2008.} We assume that the vectors $\{x^i\}_{i=1}^I$ are linearly independent; this guarantees that for each $\hat x\in X$, there is a {\em unique} bundle of goods, $\hat{\alpha}=(\hat{\alpha}_i)_{i=1}^I\in\mathbb{R}^I_+$ such that $\sum_{i=1}^I \hat{\alpha}_i x^i=\hat x$. We denote $\hat{\alpha}$ by $\alpha(\hat x)$. Let $p^t\in\mathbb{R}^I_{++}$ be the prices of the $I$ goods at observation $t$. To obtain the bundle $x\in X$, the consumer needs to spend $\psi^t(x)=p^t \cdot \alpha(x)$.
At observation $t$, the researcher observes $p^t$ and the consumer's purchases $\alpha^t\in\mathbb{R}^I_{+}$. We assume the researcher knows $\{x^i\}_{i=1}^I$ and so he can work out the consumption in characteristics space, $x^t=\sum_{i=1}^I \alpha^{t}_{i}x^i$, as well as the price system $\psi^t$. \hyperref[thm:GAPP_nonlinear]{Theorem (ref)} guarantees that if $\mathcal{D}=\{(\psi^t,x^t)\}_{t\in T}$ satisfies GAPP then it can be rationalized by a continuous augmented utility function $U:\mathbb{R}^L_+\times\mathbb{R}_{-}\to\mathbb{R}$. So long as $\psi^t(x)$ is strictly increasing in $x\in X$ for each $t$, we can also ensure that $U(y,-e)$ is increasing in the characteristics $y$.
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {The Random Augmented Utility Model}
In this section, we develop the random version of the expenditure-augmented utility model. We first describe our test procedure for this model via a simple example.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{An Illustrative Example}
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Rationalization by Random Augmented Utility}
The starting point of our analysis is a {\em repeated cross-sectional data set}, $\mathscr{D}:=\{(p^t,\mathring{\pi}^t)\}_{t=1}^T$, where each observation consists of the prevailing price $p^t$ and the distribution of demand in the population at that price, represented by a probability measure $\mathring{\pi}^t$ on $\mathbb{R}^L_+$. An example of $\mathscr{D}$ is the data set depicted in \hyperref[fig:Data]{Figure (ref)} where the probability measure corresponds to the empirical distribution of demand bundles. The following definition generalizes the notion of rationalization considered in that example.
In this definition, one could interpret $\Omega$ as the population of consumers and $\chi^t(\omega)$ as the demand of consumer type $\omega$ at observation $t$ (when the prevailing price is $p^t$); all consumer types in the support of $\mu$ are required to be consistent with the augmented utility model and the observed distribution of demand at each observation $t$, given by $\mathring{\pi}^t$, must coincide with that induced by the distribution $\mu$ over consumer types. Alternatively, the model can also be interpreted as one where each individual's augmented utility changes over time but in such a way that the population distribution is stationary.
In \hyperref[eg:StTesting]{Example (ref)}, the repeated cross-sectional data set has two observations, where the probability distributions are simply uniform distributions on finite support. A RAUM-rationalization involves matching observations in $t$ with those in $t'$, so that each pair obeys GAPP. In the general case with $T$ observations, the function $\chi$ solves a $T$-fold matching problem, where each group $\{\chi^t(\omega)\}_{t\in T}$ (along with the associated prices) satisfies GAPP and agrees with the observations (that is, (ref) is satisfied).\footnote{It is straightforward to check that, with two observations, finding a rationalization is equivalent to finding a zero-cost solution to the transportation problem (see galichon2011) where the cost of a pair of bundles is 0 if it obeys GAPP and 1 otherwise.}
We shall now explain the general procedure for deciding if a given repeated cross-sectional data set $\mathscr{D}$ can be RAUM-rationalized. This procedure mimics our solution to \hyperref[eg:StTesting]{Example (ref)}. For ease of exposition, we impose the following assumption on the data.
This assumption states that the probability of a bundle lying, after re-scaling, at the intersection of two budget planes is zero. The assumption is not required for any of our results but it is convenient because it simplifies the exposition.\footnote{If we allow for mass at budget intersections, then we would have to include them in our definition of patches. This is notationally cumbersome but once included our arguments (and \hyperref[thm:StGAPPTest]{Theorem (ref)}) remain correct.} It is always satisfied if $\mathring{\pi}^t$ is absolutely continuous with respect to the Lebesgue measure and is unlikely to be violated in any application with a continuous consumption space and linear prices.
Let $\{B^{1,t},\dots, B^{I_t,t}\}$ denote the collection of subsets, or patches, of $B^t$ where each subset has as its boundaries the intersection of $B^t$ with other budget sets and/or the boundary planes of the positive orthant. These are the higher-dimensional and multi-period analogs to the line segments in \hyperref[fig:Patches]{Figure (ref)}. Formally, for all $t\in T$ and $i_t\neq i'_t$, each set in $\{B^{1,t},\dots, B^{I_t,t}\}$ is closed and convex and satisfies the following conditions:
For the patch $B^{i_{t},t}$, we let
In words, $\pi^{i_t,t}$ is the probability that a bundle lies in $B^{i_t,t}$ after re-scaling. Note that, even though there may be $i_t$, $i'_t$ for which $B^{i_t,t}\cap B^{i'_t,t}$ is nonempty, \hyperref[ass:indif]{Assumption (ref)} guarantees that $\sum_{i_t=1}^{I_t}\pi^{i_t,t}=1$. We denote by $\pi^t$ the vector $(\pi^{i_t,t})_{i_t=1}^{I_t}$ and by $\pi$ the column vector $(\pi^1,\pi^2,\ldots,\pi^T)'$. We refer to $\pi$ as the {\em vector of observed patch probabilities}.
Consider a single-agent data set of the form $\mathcal{D}=\{(p^t,x^t)\}_{t=1}^T$. Given $\mathcal{D}$, we can define its {\em iso-expenditure version}, which is $\breve{\mathcal{D}}=\{(p^t,\breve{x}^t)\}_{t=1}^T$, where $\breve{x}^t=x^t/p^t\cdot x^t$ (so $p^t\cdot \breve{x}^t=1$ for all $t$). Suppose that $\breve{x}^t$ does not lie on the intersection of budget planes, that is, there is $i^t$ such that $\breve{x}^t\in \text{int}(B^{i^t,t})$. We make two important observations. First, \hyperref[prop:GAPPscaling]{Proposition (ref)} tells us that $\mathcal{D}$ satisfies GAPP if and only if $\breve{\mathcal{D}}$ satisfies GARP. Second, if $\mathcal{D}$ satisfies GAPP then so does $\mathcal{D}'=\{(p^t,y^t)\}_{t\in T}$ if $y^t$ has the property that its re-scaled version $\breve{y}^t$ satisfies $\tilde y^t\in\text{int}(B^{i^t,t})$; this is because the revealed preference relations (over the bundles $\tilde y^t$) are determined only by where $\breve{y}^t$ lies on the budget set relative to its intersection with another budget.
It follows from these observations that we may classify all single-agent data sets that obey GAPP according to the patch occupied by the scaled bundle $\breve{x}^t$ at each $B^t$. In formal terms, each $\mathcal{D}$ that obeys GAPP is associated with an iso-expenditure $\breve{\mathcal{D}}$ that obeys GARP, which is in turn associated with a vector $a = \left(a^{1,1},\dots, a^{I_T,T}\right)$ where
Thus, for the observed prices, we have partitioned the collection of all deterministic data sets obeying GAPP (of which there could be infinitely many) into a {\em finite} number of distinct classes or types, based on its associated vector $a$. We denote this set of vectors by $\mathcal{A}$. We use $A$ to denote the matrix whose columns consist of every $a\in\mathcal{A}$, arranged in an arbitrary order; we refer to $A$ as the {\em matrix of GARP-consistent types}.
In Example (ref), all the deterministic data sets that obey GAPP must correspond to one of three types (as depicted in \hyperref[fig:rational_types]{Figure (ref)}) and
(Each column in $A$ describes the types in \hyperref[fig:rational_types]{Figure (ref)}: from left to right the columns capture the types in Figures (ref), (ref) and (ref) respectively.)
Given a repeated cross-sectional data set $\mathscr{D}$, we can construct $\mathcal{A}$ and the matrix of GARP-consistent types $A$. Suppose that this data set can be rationalized by some distribution $\mu$. Let $\nu_a$ denote the mass of consumers of type $a$ in the population, that is $$\nu_a=\mu\left(\left\{\omega\in\Omega: \frac{\chi^t(\omega)}{p^t\cdot \chi^t(\omega)}\in B^{i_t,t}\:\mbox{ if $a^{i_t,t}=1$, for all $t\in T$}\right\}\right).$$ At a given observation $t$, let $\mathcal{A}^{i_{t},t}=\{a: a^{i_{t},t}=1\}$; this is the subset of GARP-consistent types that have their re-scaled demands in the patch $B^{i_{t},t}$ at observation $t$. The proportion of the population whose types belong to $\mathcal{A}^{i_{t},t}$ is $$\mu\left(\left\{\omega\in\Omega: \frac{\chi^t(\omega)}{p^t\cdot \chi^t(\omega)}\in B^{i_{t},t}\right\}\right)=\sum_{a\in \mathcal{A}^{i_t,t}}\nu_a=\sum_{a\in \mathcal{A}}\nu_a \, a^{i_t,t}.$$ Since $\mathscr{D}$ is rationalized by $\mu$, setting $Y=\{x\in\mathbb{R}^L_+: x/(p^t\cdot x)\in B^{i_{t},t}\}$ in ((ref)), we obtain
where $\pi^{i_{t},t}$ is defined by (ref). In other words, the observed probability of choices that land on $B^{i_t,t}$ after scaling must equal to the mass of GARP-consistent types implied by $\mu$. This condition must hold for all patches $B^{i_t,t}$, so (ref) can be more succinctly written as $A\nu=\pi$, where $\nu$ is the column vector $(\nu_a)_{a\in\mathcal{A}}$. (Recall that $\pi$ is the vector of observed patch probabilities.) In \hyperref[eg:StTesting]{Example (ref)}, $A$ is given by (ref), $\pi$ is given by (ref) and the solution $\nu$ by (ref).
To recap, given a data set $\mathscr{D}$, we calculate the matrix of GARP-consistent types $A$ and the vector of patch probabilities $\pi$. A necessary condition for $\mathscr{D}$ to be rationalized by RAUM is that there is $\nu\in\Delta^{|\mathcal{A}| -1}$ that solves $A\nu=\pi$. It turns out that this condition is also sufficient: if $\nu$ exists, then we can find a RAUM-rationalization of $\mathscr{D}$ where the proportion of the population with type $a$ is precisely $\nu_a$. The details of this final step are in the \hyperref[sec:appendix_RAUM]{Appendix (ref)}. The next result summarizes this discussion.
We end this section by contrasting the RAUM test with that of the {\em classic random utility model} (or RUM for short). The typical data environment for the latter is one where each observation consists of a distribution of choices on a given constraint set (which varies across observations). In that environment, McFadden1991 and McFadden2005 observe that the problem of testing RUM can be discretized. \citetalias{kitamura2018} operationalize this insight in the case where constraint sets are linear budget sets. In that context, requiring choices from the same constraint set simply means that $\mathscr{D}=\{(p^t,\mathring{\pi}^t)\}_{t=1}^T$ is {\em iso-expenditure}, in the sense that if $x$ is in the support of $\mathring{\pi}^t$ then $p^t\cdot x=1$. \citetalias{kitamura2018} demonstrates that an iso-expenditure data set $\mathscr{D}$ can be RUM-rationalized if and only if $A\nu=\pi$ for some $\nu\in\mathbb{R}^{|\mathcal{A}|}_+$ (where $A$ is the matrix of GARP-consistent types and $\pi$ is the vector of patch probabilities).
Notice that \hyperref[thm:StGAPPTest]{Theorem (ref)} recovers the result of \citetalias{kitamura2018} as a corollary. RAUM-rationalization guarantees the existence of a distribution over types that is consistent with the observations (that is, (ref) holds), with $\{(p^t,\chi^t(\omega))\}_{t\in T}$ satisfying GAPP almost surely. With the iso-expenditure condition, GAPP and GARP are equivalent properties, which means (by \hyperref[thm:Afriat]{Afriat's Theorem}) that there is is a strictly increasing utility function $\widetilde{U}_{\omega}:\mathbb{R}^L_+\to\mathbb{R}$ with $\widetilde{U}_{\omega}(\chi^t(\omega))\geq \widetilde{U}_{\omega}(x)$ for all $x\in B^t$; this is precisely what is needed for a RUM-rationalization. Of course, it is also clear from our proof of \hyperref[thm:StGAPPTest]{Theorem (ref)} that we are building on \citetalias{kitamura2018}, since our proof strategy involves (in effect) the following three steps: (i) converting $\mathscr{D}$ into an iso-expenditure data set $\breve{\mathscr{D}}$ (obtained from $\mathscr{D}$ simply by scaling demands); (ii) noticing that $\widetilde{\mathscr{D}}$ can be RAUM-rationalized if and only if $\breve{\mathscr{D}}$ can be RUM-rationalized; and (iii) then relying on the characterization of RUM-rationalization in \citetalias{kitamura2018}.
Since a population of heterogenous consumers typically do not have identical expenditures, an actual data set would not typically be iso-expenditure. In order to test RUM, \citetalias{kitamura2018} found it necessary to estimate an iso-expenditure data set $\breve{\mathscr{D}}$ from the true data set $\mathscr{D}$, which in turn requires an additional econometric procedure with all its attendant assumptions. In contrast, as we have established in \hyperref[thm:StGAPPTest]{Theorem (ref)}, the RAUM has the important empirical feature that it can be {\em directly tested} on data sets that are not iso-expenditure.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Welfare Comparisons}
Since the test for rationalizability involves finding a distribution $\nu$ over different types, it is possible to use this distribution for welfare analysis. To be specific, suppose that a government is contemplating a change in sales tax that could lead to prices changing from its current value $p^{t'}$ to $\hat p$. Relevant to the government's re-election prospects is the proportion of consumers who will be better off as a result of this price change.\footnote{We would like to thank an anonymous referee for suggesting this motivation.} Our methods allow us to obtain information on this proportion.
To be specific, suppose the analyst has access to a data set $\mathscr{D}$ that contains among its observations $(p^{t'},\mathring\pi^{t'})$, i.e., the prevailing prices and the demand distribution. To determine the welfare effect of a price change from $p^{t'}$ to $\hat p$, let $\mathbb{1}_{\hat p\succeq^*_p p^{t'}}$ denote the row vector with its length equal to the number of rational types ($|\mathcal{A}|$), such that the $j^{\text{th}}$ element is 1 if $\hat p\succeq^*_p p^{t'}$ for the rational type corresponding to column $j$ of $A$ and 0 otherwise.\footnote{Even though $\hat p$ is not among the observed prices, one could still define $\hat p\succeq^*_p p^{t'}$; see footnote (ref).} In words, $\mathbb{1}_{\hat p\succeq^*_p p^{t'}}$ enumerates the set of rational types for which $\hat p$ is revealed preferred to $p^{t'}$. For a rationalizable data set $\mathscr{D}$, \hyperref[thm:StGAPPTest]{Theorem (ref)} guarantees that
is the lower bound on the proportion of consumers who are revealed better off at prices $\hat p$ compared to $p^{t'}$, while the upper bound is
Since (ref) and (ref) are both linear programming problems (which have solutions if, and only if, $\mathscr{D}$ is rationalizable), they are easy to implement and computationally efficient. Suppose that the solutions are $\underline\nu$ and $\overline{\nu}$ respectively; then for any $\beta\in [0,1]$, $\beta\underline{\nu}+(1-\beta)\overline \nu$ is also a solution to $A\nu=\pi$ and, in this case, the proportion of consumers who are revealed better off at $\hat p$ compared to $p^{t'}$ is exactly $\beta\,\underline{\mathcal{N}}_{\hat p\succeq^*_p p^{t'}}+(1-\beta)\,\overline{\mathcal{N}}_{\hat p\succeq^*_p p^{t'}}$. In other words, the proportion of consumers who are revealed better off can take any value in the interval $[\underline{\mathcal{N}}_{\hat p\succeq^*_p p^{t'}}\, ,\, \overline{\mathcal{N}}_{\hat p\succeq^*_p p^{t'}}]$.
\hyperref[prop:p_welfare]{Proposition (ref)} tells us that the revealed preference relations are tight, in the sense that if, for some consumer, $\hat p$ is not revealed preferred to $p^{t'}$ then there exists an augmented utility function which rationalizes her consumption choices and for which she strictly prefers $p^{t'}$ to $\hat p$. Given this, we know that, amongst all rationalizations of $\mathscr D$, $\underline{\mathcal{N}}_{\hat p\succeq^*_p p^{t'}}$ is also the infimum on the proportion of consumers who are better off at $\hat p$ compared to $p^{t'}$.
The following proposition summarizes these observations.
It may be helpful to consider how \hyperref[prop:gocompare]{Proposition (ref)} applies in \hyperref[eg:StTesting]{Example (ref)}. In that case, there are three GAPP-consistent types with a {\em unique} $\nu$ that solves $A\nu=\pi$ (see ((ref))). Of the three types, $p^t\succeq^*_p p^{t'}$ holds only for type 2 (see \hyperref[fig:rational_types]{Figure (ref)}) and thus the proportion of consumers who are revealed better off at $p^t$ compared to $p^{t'}$ is $\nu_2=1/2$. Formally, we have $\mathbb{1}_{p^t\succeq^*_p p^{t'}}=(0,1,0)$, $\mathbb{1}_{p^t\succeq^*_p p^{t'}}\nu =1/2$, and $\underline{\mathcal{N}}_{ p^t\succeq^*_p p^{t'}}=\overline{\mathcal{N}}_{p^t\succeq^*_p p^{t'}}=1/2$.\footnote{Of the other two types in the population, type 3 (with $\nu_3=2/5$) are revealed better off at $p^{t'}$ compared to $p^t$, while type 1 consumers could be either better or worse at $p^t$ compared to $p^{t'}$. Therefore, across all rationalizations of that data set, the proportion of consumers who are better off (but not necessarily revealed better off) at $p^t$ compared to $p^{t'}$ can be as low as $1/2$ and as high as $1-2/5=3/5$.}
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {Statistical Test of RAUM, and Inference for Counterfactuals}
This section outlines our econometric methodologies. First, \hyperref[sec:econ_test]{Section (ref)} provides a statistical test of the RAUM (presented in \hyperref[sec:RAUMX]{Section (ref)}). Second, and more importantly, \hyperref[sec:econ_wel]{Section (ref)} develops a new methodology for obtaining asymptotically uniformly valid confidence intervals for counterfactual objects. This result applies to a general class of random utility models, including the RAUM. It can be used for statistical analyses of welfare comparisons and we use it for that purpose in our empirical study in \hyperref[sec:application_RAUM]{Section (ref)}.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Testing the Random Augmented Utility Model}
Recall from \hyperref[thm:StGAPPTest]{Theorem (ref)} that, given a set of prices and corresponding demand distributions $\mathscr{D}=\{(p^t,\mathring{\pi}^t)\}_{t=1}^T$ and an implied vector $\pi$ of choice probabilities on rescaled and discretized budgets, a test of the random augmented utility model is a test of
where $\Omega$ is a positive definite matrix and where the equivalence was noted and exploited in \citetalias{kitamura2018}.\footnote{The strategy to configure $H_0$ as a quadratic program also appears in dePaula2018, albeit for a different program and in a different context.}
In practice, we estimate $\pi$ by its sample analog $\hat{\pi}=(\hat{\pi}^1,\dots,\hat{\pi}^T)$ obtained by rescaling the empirical distribution of choices $\{x^t_{n_t}\}_{n_t=1}^{N_t}$ where $N_t$ is the number of observed choices in the data in period $t$. This gives rise to test statistic
where $N=\sum_{t=1}^T N_t$ denotes the total number of observations. Computing appropriate critical values for this test is delicate because the limiting distribution of $J_N$ depends discontinuously on nuisance parameters. We use the modified bootstrap procedure proposed by \citetalias{kitamura2018}.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Inference for Counterfactuals in a General Class of Random Utility Models}
A counterfactual quantity in a random utility model can be generally regarded as a function of the underlying distribution $\nu$ of individual preferences. This section focuses on the case where this mapping is linear, so that we are concerned with statistical inference for $\theta = \rho \cdot \nu$, where $\rho \in \mathbb{R}^{|\mathcal{A}|}$ is a known vector which varies with the counterfactual of interest. Our analysis of welfare comparisons in \hyperref[sec:st_welfare]{Section (ref)} falls into this framework, by letting $\theta$ be the proportion of consumers who are revealed better off at prices $\hat p$ compared to $p^{t'}$, with $\rho = \mathbb{1}_{{\hat p}\succeq^*_p p^{t'}}$. It is worth emphasizing that the methodology developed in this section has broad applicability: it can be used to study other random utility models (such as the model in KS19) and to investigate other objects of interest in random utility models; for example, lazzati2018 applies our technique to estimate the proportion of non-strategic players in a game.
Note that $\theta$ is partially identified as follows:
Our confidence interval inverts a test of
or equivalently, $$\min_{\nu\in \Delta^{|\mathcal{A}| -1}, \, \theta = \rho \cdot \nu}[\pi-A \nu]'\Omega[\pi-A \nu]=0.$$ The test statistic is a scaled sample analog
Once again, the naive bootstrap fails to deliver valid critical values for (ref), since its asymptotic distribution changes discontinuously, depending on the location of $\pi$ relative to the polytope $\mathcal S(\theta)$. This is akin to the nonstandard nature of the inference in (ref), though a simple application of the modified bootstrap algorithm in \citetalias{kitamura2018} does not work, as their method relies on, among other things, the polytope $\{A \nu:\nu \geq 0\}$ being a cone. This is not necessarily the case for counterfactual analysis, and we need to deal with $\mathcal S(\theta)$ without relying on conical properties. In this section we develop a new algorithm that guarantees asymptotic validity for inference concerning general counterfactuals.
That said, as in \citetalias{kitamura2018}, we do gain an insight from Weyl-Minkowski duality. In \hyperref[sec:appendix_metrics]{Appendix (ref)}, we show that there exist nonstochastic matrices $B$, $\tilde B$ and a nonstochastic vector-valued function $d(\theta)$ such that $\pi \in \mathcal S(\theta)$ if, and only if,
where $\bf 1$ is the $I$-vector of ones where $I=\sum_{t=1}^T I_t$ is the total number of patches. Thus, in principle this is a linear (in)equality testing problem. There is a rich literature on such problems. However, we cannot directly invoke that literature because we cannot compute $(B,\tilde B)$ in practice for a problem with a relevant scale.
While we therefore need to work with representation (ref), representation (ref) is useful. It illustrates that the inference problem is non-standard; in particular, the limiting distribution of the test statistic depends on how close to binding each of the constraints encoded in $(B,\tilde B,d(\theta))$ is. From analogy to the moment inequalities literature, it also pretty much implies that the constraints' slackness cannot be pre-estimated with sufficient accuracy; the reason being that it enters the test's asymptotic representation scaled by $\sqrt{N}$. However, we also know that certain existing procedures which shrink the estimated slack of all inequalities to zero before computing the distribution of $J_N$ will work. Our proposal is inspired by these but must implement the idea with the computationally feasible representation (ref) instead of (ref), which is only theoretically available. This means that we cannot calculate the empirical slack, which is explicit in (the empirical version of) representation (ref) but not in (ref), the very reason why a new method is called for.
Intuitively, we contract (or “tighten") the polytope $\mathcal S(\theta)$ toward a point in its relative interior, thereby effectively (but non-obviously) reducing the empirical slack in any inequality constraint. This forces all the constraints with small slacks to be binding after “tightening". Note that, unlike in \citetalias{kitamura2018}, we face substantial added complications because (i) we need to deal with a non-conical $\mathcal S(\theta)$, and (ii) the appropriate way to tighten the polytope $\mathcal S(\theta)$ varies with the value of $\theta$ through the dependence of $\mathcal S(\theta)$ on $\theta$. This leads to a restriction-dependent tightening approach which we now describe in broad strokes.
Choose a sequence $\tau_N$ such that $\tau _{N}\downarrow 0$ and $\sqrt{N}\tau _{N}\uparrow \infty $ (we make a specific proposal in the appendix) and define $$ \mathcal S_{\tau_N}(\theta) := \{A\nu\; | \; \rho \cdot \nu = \theta, \nu \in \mathcal V_{\tau_{N}}(\theta)\}, $$ where $\mathcal V_{\tau_{N}}(\theta)$ is obtained by appropriately constricting $\Delta^{|\mathcal{A}|-1}$; in particular, some components of $\nu$ are forced to be boundedly above $0$. Note that $\mathcal S_{\tau_N}(\theta)$ depends on $\theta$ through the equation $\rho \cdot \nu = \theta$ but also because, as the notation suggests, the construction of $\mathcal V_{\tau_{N}}(\theta)$ will change with $\theta$, a key feature of our algorithm. The definition of $\mathcal V_{\tau_{N}}(\theta)$ for general $\rho$ is rather involved and thus deferred to \hyperref[sec:appendix_metrics]{Appendix (ref)}, but it considerably simplifies for binary $\rho$ as in our application.
The set $\mathcal S_{\tau_N}(\theta)$ replaces $\mathcal S(\theta)$ in the bootstrap population. The precise algorithm proceeds as follows. For each $\theta \in \Theta$:
A confidence interval for $\theta$ collects values of $\theta$ that are not rejected.
\hyperref[thm:validity]{Theorem (ref)} below establishes asymptotic validity of the above procedure. Let $$ \mathcal F := \left\{(\theta,\pi) \; \left| \; \theta \in \overline{\vartheta}, \pi \in \mathcal S(\theta) \cup \mathcal P \right. \right\} $$ where ${\mathcal{P}}$ denote the set of all $\pi$ that satisfy \hyperref[condition 1]{Condition (ref)} in \hyperref[sec:appendix_metrics]{Appendix (ref)}.
The proof of \hyperref[thm:validity]{Theorem (ref)} is in \hyperref[sec:appendix_metrics]{Appendix (ref)}.
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {Empirical Applications}
In this section, we present two separate applications meant to demonstrate how both the deterministic and random versions of our model can be tested and employed for welfare analysis.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{Augmented utility model: testing and welfare analysis on Progresa data}
We apply the deterministic model to the Progresa-Oportunidades data set, a workhorse of the treatment evaluation literature. Progresa was a conditional cash transfer program aimed at poor communities in Mexico. The program was remarkable in that it was rolled out in random order so the causal effect of the cash transfers could be studied. For brevity, we do not describe the program in detail; information on the program is widely available including in the paper we discuss next.
Our application builds on recent work of attanasio2020 (henceforth \citetalias{attanasio2020}) who analyze whether the program led to changes in the market prices for basic staples: rice, kidney beans, and sugar. This is an important question because the welfare effect of these transfers would clearly depend in part on their impact on prices. While the previous literature had documented that average prices were not affected by the program hoddinott2000, \citetalias{attanasio2020} argue that sellers charge nonlinear prices and that these nonlinear price schedules had changed.
Because treatment was randomized across villages but means-tested at the household level, some households faced a changing price schedule but no shock to their own income. In our study, we focus our attention on these households because we can be more confident that their augmented utility functions are unchanged across the observation periods. Our objectives are, firstly, to test the augmented utility model and, secondly, to evaluate the welfare impact of price changes using that model. This data set is well suited for analysis using our deterministic model because its panel structure means that we can study each household separately. Following \citetalias{attanasio2020}, we consider nonlinear prices, which allows us to implement the results in \hyperref[sec:nonlinearGAPP]{Section (ref)}.
The theoretical part of \citetalias{attanasio2020} derives the optimal (nonlinear) pricing schedule under the assumption that there is a heterogenous population of households, with each household maximizing a quasilinear utility function, subject to a subsistence constraint. This constraint stipulates that a household needs a certain minimum number of calories, which can be obtained from either the observed bundle $x$ or the numeraire; given $x$, the minimum amount of the numeraire good needed to meet the calorie threshold is denoted by $\underline{z}(x)$. Thus the household can only choose among those bundles $x$ such that $\psi (x)+\underline{z}(x)\leq M$, where $\psi$ is the price system and $M$ is household wealth. It is worth noting that the augmented utility framework is sufficiently flexible to accommodate this behavior. Indeed, the household could be thought of as maximizing an augmented utility function of the modified-quasilinear form $$U(x,-e)=\widetilde U(x)- \mathbf{K}(e+\underline{z}(x)-M)\, e,$$ where $\mathbf{K}(w)=1$ if $w\leq 0$ and $\mathbf{K}(w)$ is a very large positive number if $w>0$. In this way, any $(x,-e)$ (a bundle and its associated expenditure) that leads to a violation of the subsistence constraint incurs a very large disutility and so will never be chosen.
We work with \citetalias{attanasio2020}'s data and refer to them for a detailed explanation. Compared to their analysis, we restrict ourselves to the narrower definition of village (“locality") because the larger units of analysis (“municipality") may not be contained in either the treatment or the control group. Also, because we are interested in intertemporal within-village price variation, we estimate separate price schedules for the same village in different waves as opposed to one price schedule (estimated across waves) per village. This necessitates being slightly more permissive about data needs, and we estimate prices for all village-good-wave triples that have $20$ or more (as opposed to $75$ or more) observations. We follow \citetalias{attanasio2020} in rejecting data for villages where prices strictly increase with quantity sold and where there is insufficient variation in quantities purchased.
We estimate the price schedule for good $i$ in village $v$ at wave $t$ by applying Ordinary Least Squares to
Here $h$ indexes households and $\psi_{vti}(q_{vtih}) =\mathbb{E}[p_{vti}(x_{vtih}) |x_{vtih}]\,x_{vtih}$, where $p_{vti}(x_{vtih})$ is the unit price corresponding to quantity $x_{vtih}$, $\varepsilon$ is measurement error, and the expected value is taken over the empirical distribution of reported unit prices corresponding to the same quantity purchased of good $i$ in village-wave $(v, t)$. This is exactly Equation (15) in \citetalias{attanasio2020} except for being estimated at a less aggregated level.
We test GAPP on households that:
In our final sample, this leaves us with $2488$ households in $177$ villages.\footnote{For 554 of these households we have two observations, for 840 households we have three, for 934 households we have four, and for 160 households we have five. There are so few with five observations because many households were enrolled into the program in the final wave and thus removed from our sample.}
We emphasize that GAPP is not vacuously satisfied on these data. Recall that GAPP cannot be violated when two price systems $\psi$, $\psi'$ are ranked, in the sense that $\psi(x)\geq\psi'(x)$ for all $x\in \mathbb{R}^L_+$. Of the $20556$ possible combinations of pairs of waves encountered by households in the data, about $4\%$ have this feature, and only $20$ out of $2488$ households exclusively face such price pairs and therefore satisfy GAPP vacuously. Nonetheless, $83\%$ of households pass the GAPP test. Most violations were small in the sense of the rationality index $\vartheta$ (defined in \hyperref[sec:index]{Section (ref)}) being close to $1$: fewer than $1\%$ of households were below $.9$, and fewer than $4\%$ were below $.95$.
We carried out some illustrative welfare analysis, the results of which are displayed in Tables (ref) and (ref). \hyperref[table:RP]{Table (ref)} displays the fractions of GAPP-compliant households that reveal prefer a given wave to another wave. Specifically, each cell in the table corresponds to the fraction of GAPP-rationalizable consumers who reveal prefer (directly or indirectly) the price system in the row wave to the price system in the corresponding column wave.\footnote{Note that the (indirect) revealed preference relation $\succeq^*_p$ uses demand information at {\em all} waves in each binary comparison; see the definition of $\succeq^*_p$ in \hyperref[sec:nonlinearGAPP]{Section (ref)}.} Notice that the data indicates a strong tendency to prefer price systems in later waves. For example, 91.3% of households reveal prefer prices in 03/99 to those in 10/98; the same is true even more strongly when 10/98 is compared against later waves.
To have a sense of the scale of this welfare improvement over time, we calculate, for each household, the lower bound on the compensating variation, with the price system faced by the household at 10/98 as the base.\footnote{The formula for the lower bound when prices are nonlinear is in \hyperref[sec:morecomvar]{Section (ref)}.} These values are then ranked. The results are displayed in Table (ref). Since more than 90% of households reveal prefer (price systems at) subsequent waves to 10/98, the lower bound of the compensating variation must be positive for more than 90% of households. For example, between 03/99 and 10/98, the median compensating variation is 3.27; thus, based on its observed behavior, one could remove 3.27 from this household in 03/99 and still leave it as well off in 03/99 as in 10/98. Note that the values in this table are not small, given that the household median expenditure in 10/98 on the items considered is 27.48.
These results are consistent with \citetalias{attanasio2020}'s finding that the change in the income distribution induced by Progresa caused a change in sellers' intensity of price discrimination. As a result, poorer households faced higher average prices and wealthier households faced lower ones; since Progresa was means-tested, untreated households fall into the latter category. Thus, the general equilibrium effects of the program could be the reason for the welfare improvements observed in untreated households.
\@startsection{subsection}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {\normalfont}{RAUM: Testing and welfare analysis on household expenditure data}
We test the RAUM and conduct welfare analyses on two repeated cross-sectional data sets: the U.K. Family Expenditure Survey (FES) and the Canadian Surveys of Household Spending (SHS). Our aim is to show that the data supports the model and to demonstrate that the estimated welfare bounds are informatively tight.
We first present the analysis for the FES, which is widely used in the nonparametric demand estimation literature (for instance, by blundell2008, \citetalias{kitamura2018}, hoderlein2014, adams2020, and Kawaguchi). In the FES, about 7000 households are interviewed each year and they report their consumption expenditures in different commodity groups. Following blundell2008 we derive the real consumption level for each commodity group by deflating it with a price index for that group (which is taken from the annual Retail Prices index). Again following blundell2008, we restrict attention to households with cars and children, leaving us with roughly 25% of the original data. We implement tests for $3$, $4$, and $5$ composite goods. The coarsest partition of $3$ goods---food, services, and nondurables---is precisely what is examined by blundell2008 (and we use their replication files). As in \citetalias{kitamura2018}, we introduce more commodities by first separating out clothing and then alcoholic beverages from the nondurables.
The data set we have is the sample analog of $\mathscr{D}=\{(p^t,\mathring{\pi}^t)\}_{t=1}^T$ as defined in \hyperref[sec:RatRAUM]{Section (ref)}. It is worth reiterating the point that we made at the end of \hyperref[sec:RatRAUM]{Section (ref)}: even though this data set is {\em not} iso-expenditure, we can directly test the RAUM on this data; this contrasts with testing the RUM on this data, which cannot be done directly and must involve a further procedure to estimate an iso-expenditure data set.
We implement the test in blocks of 6 years, i.e., we set $T=6$. We avoid covering a longer period partly due to the computational demands of calculating $A$ (the matrix of GARP-consistent types; see ((ref))),\footnote{That said, new techniques developed in smeulders2021 have significantly reduced the computational demands of the problem.} but also because a time-invariant distribution of augmented utility functions is only plausible over shorter time horizons, for example because of long term first-order changes to the U.K. income distribution jenkins2016.
\hyperref[table:FES_Test]{Table (ref)} displays our results, with columns correspond to different blocks of 6 years and rows contain the values of the test statistic and the corresponding p-values. The test statistic $J_N$ is defined by ((ref)), with the identity matrix serving as $\Omega$. Notice that for the year block 90-95, the test statistic is zero; this means that the sample distribution $\hat\pi$ satisfies the rationality condition in \hyperref[thm:StGAPPTest]{Theorem (ref)} exactly. That is, there is a distribution $\nu$ on GARP-consistent types such that $\hat\pi=A\nu$. Apart from this case, the sample distribution does not exactly satisfy the rationality condition and so the test statistic is strictly positive; nonetheless it is very clear from the p-values that, overall, our model is not rejected by the FES data.
We also estimated the bounds $[\underline{\mathcal{N}}_{p^t\succeq^*_p p^{t'}}\, ,\, \overline{\mathcal{N}}_{p^{t}\succeq^*_p p^{t'}}]$ (as defined by ((ref)) and ((ref))) on the proportion of households that are revealed better off at prices $p^t$ than at prices $p^{t'}$. For brevity, we present a few representative estimates using data for the years 1975-1980 in \hyperref[table:FES_Bounds]{Table (ref)}. The column `Estimated Bounds' are the bounds obtained by calculating $\mathbb{1}_{p^{t}\succeq^*_p p^{t'}}\, \nu$ from the (not necessarily unique) values of $\nu$ that minimize the test statistic ((ref)). In two cases this estimate is unique while it is not in the other two cases. Applying the procedure set out for calculating confidence intervals in \hyperref[sec:econ_wel]{Section (ref)}, we obtain the intervals displayed (which must necessarily contain the estimated bounds). It is worth noting that the width of these intervals is less than .1 throughout, so they are quite informative.\footnote{Note that, even if the true values of the proportion of the population satisfying $p^{t}\succ^*_p p^{t'}$ and $p^{t'}\succ^*_p p^{t}$ are known, they will typically add up to strictly less than 1 because, for part of the population, there will be no revealed preference relation between $p^{t}$ and $p^{t'}$. For example, type 1 consumers in \hyperref[eg:StTesting]{Example (ref)} have no revealed preference relation between $p^{t_1}$ and $p^{t_2}$.}
For our second empirical application using Canadian data, we use the replication kit of norris2013,norris2015. Like the FES, the SHS is a publicly available, annual data set of household expenditures in a variety of different categories. We study annual expenditure in 5 categories that constitute a large share of the overall expenditure on nondurables: food purchased (at home and in restaurants), clothing and footwear, health and personal care, recreation, and alcohol and tobacco. The SHS data is rich enough to allow us to analyze the data separately for the nine most populous provinces: Alberta, British Columbia, Manitoba, New Brunswick, Newfoundland, Nova Scotia, Ontario, Quebec, and Saskatchewan. The number of households in each province-year range from $291$ (Manitoba, 1997) to $2515$ (Ontario, 1997). We use province-year prices indices (as constructed by norris2015) and deflate them using province-year CPI data from Statistics Canada to get real price indices.
\hyperref[table:SHS_Test]{Table (ref)} displays the test statistics and associated p-value for each province and every 6 year block. Compared to the FES data, there are two notable differences. The first is that many more test statistics are exactly zero; that is, the observed choice frequencies are rationalized by the random augmented utility model. The second is that, for a small proportion of year blocks, there are statistically significant positive test statistics (in particular, the last three columns for British Columbia). Nonetheless, the p-values taken together do not reject the model if multiple testing is taken into account; for example, step-down procedures would terminate at the first step (that is, Bonferroni adjustment). Finally, we can also estimate the proportion of the population with a revealed preference for one year's prices over another. We provide an illustration in \hyperref[table:SHS_Test]{Table (ref)}; notice that the confidence intervals are informative, with a width no greater than 0.15.
\@startsection{section}{2} \z@{.7\linespacing\@plus\linespacing}{.5\linespacing} {Conclusion}
We propose a revealed price preference relation that generates a nonparametric ranking of price vectors; a consistency (no-cycles) condition in this relation characterizes an augmented utility model in which consumers get utility from consumption and disutility from expenditure. This model is a natural generalization of quasilinearity and, furthermore, captures some prominent behavioral models of consumption. The model is also flexible enough to accommodate nonlinear prices, discrete choice and other consumption environments. We develop the theoretical basis for welfare analysis in our model.
We generalize our model to a random utility context which is suitable for welfare analysis using repeated cross-sectional (as opposed to single-agent) data and show how to statistically test this random augmented utility model. A strength of this model is that it can be directly taken to household expenditure data in contrast to the standard random utility model which requires an additional round of estimation to account for the endogeneity of expenditure. We develop novel econometric theory to determine the proportion of consumers who are made better or worse off by a price change. This theory---which derives bounds on linear transforms of partially identified vectors---is a standalone contribution which has broader applications beyond those in this paper.
Finally, we operationalize both the deterministic and random versions of our model in separate applications to single-agent and repeated cross-sectional data. We confirm that our model is supported by data and can be used for meaningful welfare analysis.