EconBase
← Back to paper

Consumer Welfare Under Individual Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,263 characters · 15 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Consumer Welfare Under Individual Heterogeneity

titlepage\thispagestyle{empty} \begin{abstract} We propose a nonparametric method for estimating the distribution of consumer welfare from cross-sectional data with no restrictions on individual preferences. First demonstrating that moments of demand identify the curvature of the expenditure function, we use these moments to approximate money-metric welfare measures. Our approach captures both nonhomotheticity and heterogeneity in preferences in the behavioral responses to price changes. We apply our method to US household scanner data to evaluate the impacts of the price shock between December $2020$ and $2021$ on the cost-of-living index. We document substantial heterogeneity in welfare losses within and across demographic groups. For most groups, a naive measure of consumer welfare would significantly underestimate the welfare loss. By decomposing the behavioral responses into the components arising from nonhomotheticity and heterogeneity in preferences, we find that both factors are essential for accurate welfare measurement, with heterogeneity contributing more substantially. \end{abstract} {\bf Keywords}: Consumer welfare, individual heterogeneity, cost-of-living index, compensating variation, equivalent variation {\bf JEL classification}: C14, C31, D11, D12, D63, H22, I31

Introduction

Estimating consumer welfare under price changes is essential for policy evaluation across a wide range of domains, including tax reforms, trade liberalization, market interventions, and competition policies. The recent pandemic and ensuing large and sustained inflationary episodes have further renewed interest in measuring consumer welfare, highlighting the importance of accurate cost-of-living indices (e.g., \citeauthor*{fajgelbaum2016measuring}, fajgelbaum2016measuring; baqaee2022new, baqaee2022new; atkinetal23, atkinetal23; jaravel2023measuring, jaravel2023measuring).

While these advances emphasize the nonhomotheticity of preferences, they often abstract from unobserved preference heterogeneity across households. Indeed, most datasets are cross-sectional and thus offer very limited behavioral insight under individual heterogeneity. In this paper, we introduce a novel nonparametric approach to estimating consumer welfare that is applicable in cross sections of consumers with data on prices and expenditures while accounting for unrestricted unobserved preference heterogeneity.

Our method uses moments of demand to identify local approximations of moments of welfare changes, thereby providing a flexible alternative to traditional parametric models.\footnote{Standard approaches impose parametric restrictions on preferences deatonmuellbauer, lewbelpendakur. These work well when preferences are homogeneous but can misrepresent welfare impacts when unobserved heterogeneity is unrestricted. Accounting for preference heterogeneity is crucial in empirical applications, as traditional microeconometric models typically explain only a small fraction of the variation in consumer demand.} We apply our framework to measure changes in cost-of-living indices (CLIs) in the United States following the recent pandemic. Since our analysis accurately captures the distributional welfare impacts of price changes, it can offer valuable insights for policies related to poverty thresholds and welfare benefits. Our method enables comprehensive analysis of heterogeneity in experienced inflation rates, which has been widely documented (e.g., kaplan2017, kaplan2017; jaravel2019, jaravel2019; argente2021, argente2021).

The core conceptual insight of our approach is that the Slutsky equation, combined with Shephard's lemma, establishes a relationship between moments of demand and the curvature of the expenditure function. By exploiting this connection, we identify local approximations of money-metric welfare measures that are robust to unrestricted preference heterogeneity. We show that our method provides the best welfare point estimates that can be obtained in cross-sectional data under small price changes.\footnote{In this respect, our result nuances the nonidentification result of hausman2016individual, who show that average welfare is not identified from cross-sectional data under arbitrary price changes.} In fact, the entire distribution of welfare effects is point-identified in this setting. Moreover, our approach remains computationally feasible in settings with many goods, thus providing a more comprehensive assessment of the impacts of price changes across different consumer segments.\footnote{hausman2016individual derive bounds on average welfare effects in a setting with two goods, while chernozhukovetal estimate welfare effects in a multigood setting but rely on panel data.}

To understand how our key idea applies in the two-good case, consider a population of heterogeneous consumers indexed by $\omega$, with uncompensated demand $q^\omega$, compensated demand $h^\omega$, expenditure function $e^\omega$, price $p$, and income $y$. For a small price change $\Delta p := p_1 - p_0$ and baseline utility level $v_0^\omega$, the average compensating variation can be written as

equation*[equation* omitted — 550 chars of source]

where the approximation is obtained from a Taylor expansion and the equality follows from Shephard’s lemma. Although the mechanical effect is identified from average demand, the behavioral effect, which reflects substitution behavior, is not directly identified. However, the Slutsky equation yields:

equation*[equation* omitted — 331 chars of source]

While the average price effect is identified from the price derivative of average demand, the average income effect depends on a nonlinear function of demand. A key insight is that this term can be inferred from heteroscedasticity in demand with respect to income:

equation*[equation* omitted — 256 chars of source]

This relationship, in turn, enables identification of the average substitution effect and, via Shephard’s lemma, of the average welfare impact through the curvature of the expenditure function.

The generality of our methodology enables several important extensions. First, we assess the biases introduced by standard CLIs that impose homothetic or homogeneous preference restrictions. In particular, we show that our cost-of-living estimate admits a decomposition that quantifies the shortcomings of these traditional measures. Second, the theoretical insights underlying our approach naturally extend beyond consumer welfare analysis. We illustrate this flexibility by incorporating income effects into the estimation of elasticities of taxable income (e.g., grubersaez02, grubersaez02; brunsziliak16, brunsziliak16). Accounting for income effects is empirically relevant, as recent research documents their quantitative importance golosov2024wealth. More broadly, our framework applies to a wide range of settings where sufficient statistics are used to evaluate efficiency or welfare costs kleven2020.

In our empirical application, we use detailed household-level scanner data from NielsenIQ spanning $2019$ to $2022$ to quantify how the first COVID-19 price shock translated into welfare losses for US grocery shoppers. The data track purchases of fast-moving consumer goods among a large, nationally representative panel of households. For each household, we observe prices paid and quantities purchased, which allow us to construct household-specific price indices across food categories. In addition to detailed purchase data, the dataset includes rich household demographics, such as household income, size, race, and the education level of each household head. These features enable us to estimate the distributional effects of inflation across and within demographic subgroups.

Using our nonparametric method to estimate the CLI, we find that households would have needed over $8\%$ of their monthly food budget in additional compensation to maintain the same utility in December $2021$ as they had in December $2020$. Importantly, the burden of inflation varies substantially across goods. For example, price increases in dry grocery required compensation exceeding $45\%$ of monthly food expenditures, while the corresponding figure for packaged meat was just $6\%$. These differences underscore that the welfare impact of inflation is highly sensitive to the specific categories in which price increases occur.

Next, we find significant heterogeneity in the welfare losses within and across demographic groups. Across demographic groups, this heterogeneity is driven in part by differences in price increases. However, these differences do not follow a systematic pattern based on observable household characteristics such as race or education. Overall, a typical household’s CLI deviates by approximately $2$ percentage points from the mean. In addition, within groups, households with higher education levels tend to exhibit greater dispersion in CLI. These findings highlight the relevance of unobserved heterogeneity, such as in shopping patterns or store access, in shaping welfare outcomes.

We then compare our CLI to a first-order approximation that holds shares fixed and ignores behavioral responses to price changes such as substitution across goods or income effects. We find that this approximation underestimates the welfare losses by about $4\%$ when we apply it to a composite basket. Notably, this bias is relatively stable across demographic subgroups. However, when examined for specific categories of goods, the bias is far more variable and can be substantial. For instance, the approximation overestimates the welfare loss from price increases in dry grocery and alcohol by approximately $11\%$ and $40\%$, respectively. These large discrepancies highlight that the accuracy of the first-order approximation depends heavily on the category of goods affected and the curvature of demand.

Finally, our decomposition of the behavioral component of the CLI highlights how essential it is to incorporate preference heterogeneity to measure welfare accurately, especially across demographic groups. Models that assume homogeneous preferences, whether or not they incorporate nonhomotheticity, can fail to capture $20\%$ to $40\%$ of the behavioral response to price shocks for some demographic groups. In contrast, the explained share is always within $5\%$ of the full behavioral response once preference heterogeneity is allowed, underscoring that heterogeneity is the primary force behind behavioral adjustments.

\paragraph{Related literature.} Our methodology contributes to a long-standing tradition of estimating consumer welfare nonparametrically using cross-sectional data.\footnote{See bhattacharya24 for a recent and comprehensive review.} A prominent early approach in this literature treats average demand as if generated by a representative consumer and thereby abstracts from preference heterogeneity hausman1981exact, vartia1983efficient. Along these lines, hausman1995nonparametric obtain point estimates via nonparametric regression, while FosterHahn2000 and Blundelletal2003 derive conditions under which such estimates provide a first-order approximation of the true welfare changes. Similarly, \citet*{schlee2007measuring} provides sufficient conditions under which the representative agent framework yields upper bounds on welfare.

Motivated by aggregation results from gorman53, this approach typically assumes that preferences are homogeneous or homothetic. However, interpreting average demand as that of a representative consumer is valid only under restrictive conditions jerison1994optimal, lewbel2001demand, such as when preferences are homothetic or when income effects are negligible. Unfortunately, these restrictions on preferences are often too strong in practice. Notably, the covariance term in lewbel2001demand captures the failure of average demand to satisfy integrability conditions due to preference heterogeneity.

Closer to our approach, banks1996tax improve on first-order approximations by incorporating second-order terms. However, they do not link demand moments to the curvature of the expenditure function, a technique that we exploit to characterize welfare impacts under arbitrary preference heterogeneity. \citet*{hoderleinvanhemsm} provide point identification of quantile demand functions in the two-good case by assuming monotonicity in unobserved scalar heterogeneity. While elegant, their identification strategy requires that individuals retain their rank in the distribution of demand across budget sets.

More fundamentally, hausman2016individual show that average consumer welfare is not point-identified from cross-sectional data under unrestricted heterogeneity and provide worst-case bounds under mild assumptions. Although their results are robust, their approach can lead to wide welfare intervals, particularly when income effects are large. In contrast, our method provides point estimates of local welfare under general heterogeneity, including in settings with multiple goods and unrestricted income effects.

Relatedly, our insight that variation in demand contains a source of identification in cross-sectional data connects to a broader literature. For example, the idea that the variance of demand is informative about average income effects is noted by hildenbrand83 and hoderlein2011many. More generally, hoderleinmammen establish that, in nonseparable models, cross-sectional data identify local average structural derivatives, but not transformations thereof. Their result helps clarify that nonidentification of welfare effects stems from the inability to learn higher-order income effects in cross-sectional data.

A different strand of literature relies on revealed preference methods to bound welfare impacts in random utility models. For instance, \citet*{cosaert18} apply the weak axiom of stochastic revealed preference to repeated cross sections. \citet*{deb22} derive and exploit novel revealed preference conditions over prices in repeated cross sections. \citet*{allen2020counterfactual} use revealed preference conditions for the law of demand that are applicable in panel data. While these methods are robust to functional form assumptions, they typically yield set identification.

Our approach also contributes to the literature on CLIs and inflation heterogeneity. Several recent papers examine how nonhomothetic preferences shape inflation across income levels and consumption baskets fajgelbaum2016measuring, baqaee2022new, jaravel2023measuring, and kaplan2017 construct household-specific inflation indices to reflect differential price exposure across groups. We complement this body of work by allowing for unrestricted heterogeneity in preferences and quantifying the bias introduced by assuming homotheticity.

Finally, our framework contributes to the growing literature on sufficient statistic approaches in public finance. In particular, we extend the tools for estimating elasticities and behavioral responses to settings with income effects and heterogeneous preferences, thereby relaxing the quasi-linear utility assumptions common in earlier work grubersaez02, brunsziliak16. Our methodology also lends itself to the valuation of redistributive policies and government programs, in line with recent unified frameworks as in hendren2020unified, kleven2020, and finkelstein2019.

\paragraph{Outline of the paper.} This paper is organized as follows. Section (ref) presents the notation and conceptual framework. Section (ref) exploits information on the curvature of the expenditure function to approximate moments of consumer welfare. Section (ref) contains extensions of our main result to a quantile-based approach, a decomposition of the behavioral response, and elasticities of taxable income. Section (ref) details the application and illustrates the usefulness of our method for assessing changes in consumer welfare. Finally, Section (ref) concludes. The appendix provides omitted proofs and additional results.

Framework

Our framework allows for unrestricted unobserved heterogeneity in preferences. For ease of exposition, we suppress observed individual characteristics from the notation in the rest of the paper; our results can be thought of as conditional on such characteristics.

\paragraph{Setup and notation.} Let $\Omega$ be the universe of preference types. Each preference type $\omega \in \Omega$ can be viewed as an individual with preferences over bundles of $(k+1)$ goods $\mathbf{q} \in \mathcal{Q} \subseteq \mathbb{R}_{++}^{k+1}$, where $\mathcal{Q}$ is compact and convex. We assume that preferences are represented by smooth and strictly quasi-concave utility function $u^\omega : \mathcal{Q} \to \mathbb{R}$ that is infinitely differentiable everywhere.\footnote{This ensures that the demand functions are infinitely differentiable in prices and income.} Furthermore, we denote prices by $\mathbf{p} \in \mathcal{P}\subset \mathbb{R}_{++}^{k+1}$ and income by $y \in \mathcal{Y} \subset \mathbb{R}_{++}$. We use Euler's notation for differentiation such that $D_{x_1,x_2}^{m,n} f(x_1, x_2) := \frac{\partial^{m+n} f(x_1,x_2)}{\partial x_{1}^m \partial x_{2}^n}$. The derivative of a vector function with respect to another vector is expressed in numerator layout such that $(D_\mathbf{x} \mathbf{f}(\mathbf{x}))_{ij} = D_{x_j} f_i(\mathbf{x})$. A variable in logarithms is denoted with a tilde: e.g., $\widetilde{\mathbf{p}} := \log(\mathbf{p})$ and $\widetilde{y} := \log(y)$. The symbol $\odot$ denotes the element-wise product. Throughout, we use $O$ to denote Landau's big O.\footnote{Formally, $f(x) = O(g(x))$ means there exists constants $C >0$ and $x_0$ such that $|f(x)| \leq C|g(x)|$ for all $x \geq x_0$.}

\paragraph{Consumer demand.} We consider the standard model of utility maximization under a linear budget constraint. Accordingly, an individual demand function $\mathbf{q}^\omega(\mathbf{p}, y) : \mathcal{P} \times \mathcal{Y} \to \mathcal{Q}$ is defined as

equation*[equation* omitted — 168 chars of source]

For every uncompensated (or Marshallian) demand function $\mathbf{q}^\omega$, there exists a compensated (or Hicksian) demand function $\mathbf{h}^\omega(\mathbf{p}, u) : \mathcal{P} \times \mathbb{R} \to \mathcal{Q}$ defined as

equation*[equation* omitted — 181 chars of source]

Observe that both demand functions are linked through the Slutsky equation.\footnote{The Slutsky equation is defined as $D_\mathbf{p} \mathbf{h}^\omega(\mathbf{p}, u^\omega(\mathbf{q}^\omega(\mathbf{p}, y))) = D_\mathbf{p} \mathbf{q}^\omega(\mathbf{p},y) + D_y \mathbf{q}^\omega(\mathbf{p},y) \mathbf{q}^\omega(\mathbf{p},y)^\intercal$.} Note also that, without loss of generality, one can omit the demand and price for the $(k+1)$st good because of Walras's law. Next, the indirect utility function $v^\omega(\mathbf{p},y) : \mathcal{P} \times \mathcal{Y} \to \mathbb{R}$ is defined as

equation*[equation* omitted — 138 chars of source]

and gives the maximum utility level obtained for the budget set defined by $(\mathbf{p},y)$. Likewise, the expenditure function $e^\omega(\mathbf{p}, u) : \mathcal{P} \times \mathbb{R} \to \mathcal{Y}$ is defined as

equation*[equation* omitted — 141 chars of source]

and gives the minimum amount of income needed to achieve utility level $u$ at prices $\mathbf{p}$.

\paragraph{Moments of consumer demand.} To avoid complex tensor notation, it will be convenient to introduce scalar-valued composite demands. For any $\mathbf{t} \in \mathbb{R}^k$, let

equation*[equation* omitted — 258 chars of source]

denote the composite uncompensated and compensated demand, respectively. These composite demands can be interpreted as the projection of demand bundles on the line through the origin defined by $\mathbf{t}$. For any $n \in \mathbb{N}_{++}$, we define the $n$th (noncentral) conditional moment of composite (uncompensated) demand as

equation[equation omitted — 247 chars of source]

where $F(\omega)$ denotes the distribution of preference types. In particular, note that we recover moments of (noncomposite) demand by setting $\mathbf{t} = (0,\dots,0,1,0,\dots,0)^\intercal$.

These moments generalize the concept of the average structural function blundell_powell_2003. Under budget set exogeneity, they are nonparametrically identified from cross-sectional data, as they are conditional expectation functions. Specifically, budget set exogeneity implies that $F(\omega \mid \mathbf{p},y) = F(\omega)$. While the assumption of exogenous budget sets is strong, it is also standard in the literature on nonparametric identification (e.g., see hausman2016individual, hausman2016individual; \citeauthor*{blomquistnewey}, blomquistnewey). To our knowledge, existing theoretical results do not establish identification under unrestricted preference heterogeneity when budget sets are endogenous in cross-sectional settings. However, certain forms of endogeneity can be addressed by means of a control function approach. We implement this approach to account for endogenous expenditure in our empirical application in Section (ref).

In the remainder of the paper, we assume that every moment exists and is finite. Unless we state otherwise, expectations and moments are always conditional on prices and income. Nevertheless, we refer to such conditional moments as “moments" for the sake of brevity. Finally, we assume that conditions for the dominated convergence theorem hold such that derivative and integral operators can be interchanged.\footnote{That is, we assume that there exists a function $g : \Omega \rightarrow \mathbb{R}$ such that, for all $(\mathbf{p},y) \in \mathcal{P} \times \mathcal{Y}$ and $m,n \in \mathbb{N}$, it holds that $\Big\Vert \text{vec}\left(D_{\mathbf{p}, y}^{m,n} \mathbf{q}^\omega(\mathbf{p},y)\right)\Big\Vert \leq g(\omega)$ with $\int g(\omega) dF(\omega) < \infty$.}

\paragraph{Budget shares.} Some of our results are more conveniently expressed in terms of budget shares. Therefore, define the uncompensated and compensated budget shares as

equation*[equation* omitted — 268 chars of source]

and their associated composite shares as

equation*[equation* omitted — 267 chars of source]

For instance, notice that $w_q^\omega(\mathbf{p}, y; \widetilde{\mathbf{p}})$ corresponds to Stone's price index stone1954. Similarly to the levels of demand, the $n$th moment of the composite compensated budget share is denoted with $W_n(\mathbf{p}, y; \mathbf{t}) := \mathbb{E}[w_q^\omega(\mathbf{p}, y; \mathbf{t})^n]$.

Identification

In this section, we show that moments of consumer welfare can be approximated locally by observed moments of demand.\footnote{A more in-depth analysis of the informational content of the moments of demand is provided in maesmalhotra.} Our identification argument is constructive and naturally leads to plug-in estimators that are simple to implement. A summary of our main argument is presented in Figure (ref).

\tikzset{every shadow/.style={fill=none,shadow xshift=0pt,shadow yshift=0pt}}

figure[figure omitted — 1,096 chars of source]

Curvature of the expenditure function

Before proceeding to our first main result, we establish a fundamental connection between price and income effects and moments of uncompensated demand.

lemFor every $n \in \mathbb{N}_{++}$ and $\mathbf{t} \in \mathbb{R}^k$, it holds that \begin{equation*} \begin{split} \mathbb{E}\left[q^\omega(\mathbf{p},y; \mathbf{t})^{n-1} D_y q^\omega(\mathbf{p},y; \mathbf{t})\right] &= \frac{1}{n} D_y M_{n}(\mathbf{p},y; \mathbf{t}), \\ \mathbb{E}\left[q^\omega(\mathbf{p},y; \mathbf{t})^{n-1} D_\mathbf{p} q^\omega(\mathbf{p},y; \mathbf{t})\right] &= \frac{1}{n} D_\mathbf{p} M_{n}(\mathbf{p},y; \mathbf{t}), \end{split} \end{equation*} at all $(\mathbf{p}, y) \in \mathcal{P}\times \mathcal{Y}$.
proofFrom the definition of moments of demand, we have \begin{equation*} \begin{split} D_y M_{n}(\mathbf{p},y; \mathbf{t}) &= D_y \left(\int q^\omega(\mathbf{p},y; \mathbf{t})^n dF(\omega)\right) \\ &= \int D_y q^\omega(\mathbf{p},y; \mathbf{t})^n dF(\omega) \\ &= n \mathbb{E}\left[q^\omega(\mathbf{p},y; \mathbf{t})^{n-1} D_y q^\omega(\mathbf{p},y; \mathbf{t})\right], \end{split} \end{equation*} where the second equality follows from the dominated convergence theorem and the third follows from the chain rule. The proof for the price effects is analogous, mutatis mutandis.

While straightforward, Lemma (ref) serves as the key insight underlying all subsequent results. Indeed, Lemma (ref) establishes that cross-sectional data are informative about the (noncentered) covariance between powers of composite demand and the marginal propensity to consume. The information on these covariances will be crucial when we reconstruct the compensated price responses and conduct welfare analysis. For instance, it implies that the expected income effect is identified from the derivative of the second moment of demand with respect to income:

equation*[equation* omitted — 168 chars of source]

That the variance of demand is informative about average income effects has been observed before by hildenbrand83 and hoderlein2011many. Lemma (ref) generalizes this key insight to higher-order income effects.

Let the curvatures of the expenditure function and log expenditure function in the direction $\mathbf{t}$ be defined as

equation*[equation* omitted — 385 chars of source]

respectively. This curvature captures the sensitivity of the consumer's cost of maintaining her current utility to price changes in the direction $\mathbf{t}$. Its magnitude plays a central role in the approximation of money-metric welfare measures such as the compensating variation and the CLI (see Section (ref)). Using Lemma (ref) and leveraging the Slutsky equation, we can now identify features of the curvature of the expenditure function.

propFor every $n \in \mathbb{N}_{++}$ and $\mathbf{t} \in \mathbb{R}^k$, it holds that \begin{equation*} \begin{split} \mathbb{E} \left[q^\omega(\mathbf{p}, y; \mathbf{t})^{n-1} c^\omega(\mathbf{p}, y; \mathbf{t}) \right] &= \frac{1}{n} D_\mathbf{p} M_n(\mathbf{p}, y; \mathbf{t})\mathbf{t} + \frac{1}{n+1} D_y M_{n+1}(\mathbf{p}, y; \mathbf{t}), \\ \mathbb{E} \left[ w^\omega_q(\mathbf{p}, y; \mathbf{t})^{n-1}\widetilde{c}^\omega(\mathbf{p}, y; \mathbf{t}) \right] &= \frac{1}{n} D_{\widetilde{\mathbf{p}}} W_n(\mathbf{p}, y; \mathbf{t})\mathbf{t} + \frac{1}{n+1} D_{\widetilde{y}} W_{n+1}(\mathbf{p}, y; \mathbf{t}) + W_{n+1}(\mathbf{p}, y; \mathbf{t}), \end{split} \end{equation*} at all $(\mathbf{p}, y) \in \mathcal{P}\times \mathcal{Y}$.
proofFor clarity, we focus on the curvature of the expenditure function; see Appendix (ref) for the proof for the log expenditure function. For any $(\mathbf{p}, y) \in \mathcal{P}\times \mathcal{Y}$, we have \begin{equation*} \begin{split} c^\omega(\mathbf{p}, y; \mathbf{t}) &= \mathbf{t}^\intercal D_\mathbf{p}^2 e^\omega(\mathbf{p}, v^\omega(\mathbf{p}, y)) \mathbf{t} \\ &= \mathbf{t}^\intercal D_\mathbf{p} \mathbf{h}^\omega(\mathbf{p}, v^\omega(\mathbf{p}, y)) \mathbf{t} \\ &= \mathbf{t}^\intercal \left[D_\mathbf{p} \mathbf{q}^\omega(\mathbf{p}, y) + D_y \mathbf{q}^\omega(\mathbf{p},y) \mathbf{q}^\omega(\mathbf{p},y)^\intercal \right]\mathbf{t} \\ &= D_\mathbf{p}(\mathbf{t}^\intercal \mathbf{q}^\omega(\mathbf{p}, y))\mathbf{t} + (\mathbf{t}^\intercal \mathbf{q}^\omega(\mathbf{p}, y))D_y(\mathbf{t}^\intercal \mathbf{q}^\omega(\mathbf{p}, y)) \\ &= D_\mathbf{p} q^\omega(\mathbf{p}, y; \mathbf{t})\mathbf{t} + q^\omega(\mathbf{p}, y; \mathbf{t})D_y q^\omega(\mathbf{p}, y; \mathbf{t}), \end{split} \end{equation*} where the second equality follows from Shephard's lemma and the third from the Slutsky equation. Therefore, \begin{equation*} \begin{split} \mathbb{E} \left[q^\omega(\mathbf{p}, y; \mathbf{t})^{n-1} c^\omega(\mathbf{p}, y; \mathbf{t}) \right] &= \mathbb{E} [q^\omega(\mathbf{p}, y; \mathbf{t})^{n-1}[D_\mathbf{p} q^\omega(\mathbf{p}, y; \mathbf{t})\mathbf{t} + q^\omega(\mathbf{p}, y; \mathbf{t})D_y q^\omega(\mathbf{p}, y; \mathbf{t})]] \\ &= \frac{1}{n} D_\mathbf{p} M_n(\mathbf{p}, y; \mathbf{t})\mathbf{t} + \frac{1}{n+1} D_y M_{n+1}(\mathbf{p}, y; \mathbf{t}), \end{split} \end{equation*} where the last equality follows from Lemma (ref).
remProposition (ref) implies that moments of demand contain information about compensated price responses. Indeed, by Shephard's lemma, we have \begin{equation*} \begin{split} c^\omega(\mathbf{p}, y; \mathbf{t}) = \mathbf{t}^\intercal D_\mathbf{p} \mathbf{h}^\omega(\mathbf{p}, v^\omega(\mathbf{p}, y)) \mathbf{t}. \end{split} \end{equation*} For instance, Proposition (ref) implies that the average substitution matrix $\mathbb{E}[D_\mathbf{p} \mathbf{h}^\omega(\mathbf{p}, v^\omega(\mathbf{p},y)]$ is identified from cross-sectional data since a symmetric matrix is uniquely determined by its quadratic forms. Using arguments analogous to those above, one can show that \begin{equation*} \begin{split} \mathbb{E}[D_\mathbf{p} \mathbf{h}^\omega(\mathbf{p}, v^\omega(\mathbf{p},y)] & =\frac{1}{2}\left[D_\mathbf{p} \mathbf{M}_1(\mathbf{p},y)+ \left(D_\mathbf{p}\mathbf{M}_1(\mathbf{p},y)\right)^\intercal+ D_y \mathbf{M}_2(\mathbf{p},y)\right], \end{split} \end{equation*} where $\mathbf{M}_1(\mathbf{p}, y)$ is the vector of first moments and $\mathbf{M}_2(\mathbf{p}, y)$ the matrix of second moments.\footnote{This result follows, for example, by adding the Slutsky equation to its transpose and taking expectations.} Since every term on the right-hand side is nonparametrically identified from cross-sectional data, the average price response is itself identified.
remIn the often-studied two-good case, one can obtain a first-order approximation of the $n$th moment of compensated demand at a counterfactual price $p'$ through the $n$th and $(n+1)$st moments of demand. Using a first-order Taylor approximation around $p' \approx p$, we have that \begin{equation*} \begin{split} h^\omega(p', v^\omega(p,y))^n &= h^\omega(p, v^\omega(p,y))^n + \left[h^\omega(p, v^\omega(p,y))^{n-1} D_p h^\omega(p, v^\omega(p,y)) \right] (p'-p) + O((p'-p)^2)\\ &= q^\omega(p, y)^n + \left[q^\omega(p, y)^{n-1} D_p h^\omega(p, v^\omega(p,y)) \right] (p'-p) + O((p'-p)^2)\\ \end{split} \end{equation*} Taking expectations on both sides and using Lemma (ref) gives \begin{equation*} \begin{split} \mathbb{E}[h^\omega(p', v^\omega(p,y))^n] = M_n(p, y) + \left(\frac{1}{n} D_p M_n(p,y) + \frac{1}{n+1} D_y M_{n+1}(p,y)\right) (p' - p) + O((p' - p)^2). \end{split} \end{equation*}

Consumer welfare

Since individual preferences are unobservable, consumer welfare is inherently stochastic from the perspective of the analyst. In this section, we examine how the distribution of consumer welfare can be identified and recovered from cross-sectional data.

\paragraph{Measures of consumer welfare.}

We derive results for two widely used money-metric measures of consumer welfare: the compensating variation (CV) and the cost-of-living index (CLI). Consider a (potentially counterfactual) price change from $\mathbf{p}_0$ to $\mathbf{p}_1$ and define the absolute and logarithmic price differences as $\Delta \mathbf{p} := \mathbf{p}_1 - \mathbf{p}_0$ and $\widetilde{\Delta \mathbf{p}} := \widetilde{\mathbf{p}}_1 - \widetilde{\mathbf{p}}_0$, respectively. Further define the indirect utilities as $v^\omega_0 := v^\omega(\mathbf{p}_0, y)$ and $v^\omega_1 := v^\omega(\mathbf{p}_1, y)$, respectively.

The CV represents the amount of income an individual would require after a price change to restore her initial utility level. Formally, it is defined as\footnote{For clarity of exposition, we adopt a sign convention for the CV and CLI that differs from the standard textbook definition (e.g., see \citeauthor*{mascolelletal95}, mascolelletal95).}

equation[equation omitted — 242 chars of source]

Likewise, the CLI represents the proportional change in income an individual would require after a price change to restore her initial utility level. Formally, it is defined as

equation[equation omitted — 289 chars of source]

Note that both $CV^\omega(\mathbf{p}_0,\mathbf{p}_1, y)$ and $CLI^\omega(\mathbf{p}_0,\mathbf{p}_1, y)$ are positive whenever $\mathbf{p}_1 > \mathbf{p}_0$, reflecting a loss in purchasing power due to a price increase. Conversely, both measures are negative when $\mathbf{p}_1 < \mathbf{p}_0$. We assume uniformly bounded income effects to ensure that consumer welfare is bounded (see Appendix (ref)).

\paragraph{Identification of local effects.} Building on our earlier result on the curvature of the expenditure function, we now state the central result. The theorem below demonstrates that moments of consumer welfare can be approximated in terms of moments of demand.\footnote{Since consumer welfare is bounded, standard results on the Hausdorff moment problem ensure that the distributions of the CV and CLI are uniquely determined by their moments.} The identification argument is constructive, providing a basis for straightforward plug-in estimators.

thmThe $(n+1)$st-order approximation of the $n$th moment of the CV and CLI depends only on the $n$th and $(n+1)$st moment of demand: \begin{equation*} \begin{split} \mathbb{E}[CV^\omega(\mathbf{p}_0, \mathbf{p}_1, y)^n] &= \mathfrak{M}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) + \mathfrak{B}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) + O(||\Delta \mathbf{p}||^{n+2}), \\ \mathbb{E}[CLI^\omega(\mathbf{p}_0, \mathbf{p}_1, y)^n] &= \mathfrak{M}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) + \mathfrak{B}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) + O(||\widetilde{\Delta \mathbf{p}}||^{n+2}), \end{split} \end{equation*} where \begin{equation*} \begin{split} \mathfrak{M}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) &= M_n(\mathbf{p}_0, y; \Delta \mathbf{p}), \\ \mathfrak{B}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) &= \frac{1}{2} \left(D_\mathbf{p} M_n(\mathbf{p}_0, y; \Delta \mathbf{p})\Delta \mathbf{p} + \frac{n}{n+1} D_y M_{n+1}(\mathbf{p}_0, y; \Delta \mathbf{p})\right), \\ \mathfrak{M}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &= W_n(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}), \\ \mathfrak{B}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &= \frac{1}{2} \left(D_{\widetilde{\mathbf{p}}} W_n(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})\widetilde{\Delta \mathbf{p}} + \frac{n}{n+1} D_{\widetilde{y}} W_{n+1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) + nW_{n+1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})\right). \end{split} \end{equation*}
proofFor the sake of brevity, we present only the proof for the CV herein; see Appendix (ref) for the proof for the CLI. The second-order expansion of the expenditure function for $\Delta \mathbf{p} \approx \mathbf{0}$ can be written as \begin{equation*} \begin{split} e^\omega(\mathbf{p}_1, v_0^\omega) &= e^\omega(\mathbf{p}_0, v_0^\omega) + D_\mathbf{p} e^\omega(\mathbf{p}_0, v_0^\omega) \Delta \mathbf{p} + \frac{1}{2}(\Delta \mathbf{p})^\intercal D_\mathbf{p}^2 e^\omega(\mathbf{p}_0, v_0^\omega) \Delta \mathbf{p} + O(||\Delta \mathbf{p}||^3) \\ &= y + q^\omega(\mathbf{p}_0, y; \Delta \mathbf{p}) + \frac{1}{2}c^\omega(\mathbf{p}_0, y; \Delta \mathbf{p}) + O(||\Delta \mathbf{p}||^3). \end{split} \end{equation*} Therefore, using Equation (ref), the second-order approximation of the CV becomes \begin{equation*} CV^\omega(\mathbf{p}_0, \mathbf{p}_1, y) = q^\omega(\mathbf{p}_0, y; \Delta \mathbf{p}) + \frac{1}{2}c^\omega(\mathbf{p}_0, y; \Delta \mathbf{p}) + O(||\Delta \mathbf{p}||^3), \end{equation*} such that, for higher powers, we obtain \begin{equation*} CV^\omega(\mathbf{p}_0, \mathbf{p}_1, y)^n = q^\omega(\mathbf{p}_0, y; \Delta \mathbf{p})^n + \frac{n}{2} q^\omega(\mathbf{p}_0, y; \Delta \mathbf{p})^{n-1}c^\omega(\mathbf{p}_0, y; \Delta \mathbf{p}) + O(||\Delta \mathbf{p}||^{n+2}). \end{equation*} Taking expectations on both sides and using Proposition (ref) gives \begin{equation*} \begin{split} \mathbb{E}[CV^\omega(\mathbf{p}_0, \mathbf{p}_1, y)^n] &= \mathfrak{M}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) + \mathfrak{B}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p})+ O(||\Delta \mathbf{p}||^{n+2}), \end{split} \end{equation*} as desired.

In these expressions, the terms $\mathfrak{M}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p})$ and $\mathfrak{M}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})$ represent the mechanical contributions of price changes to the $n$th moments of consumer welfare---specifically, the effects that would arise if consumers did not adjust their behavior. In contrast, the terms $\mathfrak{B}_n^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p})$ and $\mathfrak{B}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})$ capture behavioral adjustments, reflecting the reoptimization of consumer choices in response to the change in prices.

remTheorem (ref) admits further simplification in two empirically relevant cases. First, in the two-good setting, the approximation for moments of the CV simplifies to \begin{equation*} \mathbb{E}[CV^\omega(p_0, p_1, y)^n] \approx (\Delta p)^n\left(M_n(p_0,y)+\frac{\Delta p}{2}\left[D_p M_n(p_0,y) +\frac{n}{n+ 1} D_y M_{n+1}(p_0,y)\right]\right). \end{equation*} This case is of practical importance because when the price of a single good changes, only the demand for that good must be modeled. This substantially reduces the dimensionality of the commodity space \citep*{hausman1981exact}. Second, even in the many-good setting, the second-order approximation of the average CV depends solely on the first two moments of demand: \begin{equation*} \begin{split} \mathbb{E}[CV^\omega(\mathbf{p}_0, \mathbf{p}_1, y)] \approx M_{1}(\mathbf{p}, y; \Delta \mathbf{p}) + \frac{1}{2} \left(D_\mathbf{p} M_{1}(\mathbf{p}, y; \Delta \mathbf{p}) \Delta \mathbf{p} + \frac{1}{2} D_y M_{2}(\mathbf{p}, y; \Delta \mathbf{p})\right). \end{split} \end{equation*} Thus, average welfare effects can be computed directly from estimates of the mean and variance of demand, offering a tractable approach in applied work. Analogous simplifications hold for the CLI.
remA useful way to interpret our approach to estimating the CLI is as a localized application of the {translog unit cost function} christensenjorgensonlau, diewert76. Specifically, the approximation developed in Theorem (ref) corresponds to a second-order expansion of the log expenditure function around observed prices. This is similar to evaluating welfare by means of the translog form \begin{equation*} \widetilde{e}_{TL}^\omega(\mathbf{p}, u) := \alpha_0 + \widetilde{\mathbf{p}}^\intercal \boldsymbol{\alpha}_1 + \frac{1}{2} \widetilde{\mathbf{p}}^\intercal \boldsymbol{\Gamma} \widetilde{\mathbf{p}}, \end{equation*} where $\mathbf{1}^\intercal \boldsymbol{\alpha}_1 = 1$, $\boldsymbol{\Gamma}$ is symmetric and $\boldsymbol{\Gamma} \mathbf{1} = \mathbf{0}$. Our method can therefore be seen as recovering a local translog representation of preferences to approximate the welfare effects of price changes without imposing the global structure of a fully specified demand system.

\paragraph{Nonidentification of global effects.} Our local approximations of consumer welfare represent the most informative estimates obtainable in a cross-sectional setting. The inability to identify higher-order approximations stems from the fact that cross-sectional data lack information about how the variance---and higher moments---of income effects vary across demand bundles. While the issue of nonidentification was previously noted by hausman2016individual, Proposition (ref) sharpens this insight by formally characterizing the limits of identification.

propThe $(n+2)$nd-order approximation of the $n$th moment of the CV or CLI is not identified.
proofFor clarity, we focus on the average CV with two goods. Suppose the true series expansion of the average CV at a given budget set $(p_0, y)$ is $\mathbb{E}[CV^\omega(p_0,p_1,y)]=a_0 +a_1\Delta p+a_2(\Delta p)^2+a_3(\Delta p)^3$. Extending the argument from the proof of Theorem (ref), recovering $a_3$ requires identifying $\mathbb{E}\left[D_p^2 h^\omega(p_0,v^\omega_0) \right]$, i.e., the expected second derivative of compensated demand with respect to price. Differentiating the identity $h^\omega(p_0,v_0^\omega) \equiv q^\omega(p_0,e^\omega(p_0,v_0^\omega))$ twice with respect to price and substituting in the Slutsky equation, we obtain \begin{equation*} \begin{split} D_p^2 h^\omega(p_0,v^\omega_0) &= D_p^2 q^\omega(p_0,y)+ q^\omega(p_0,y) D_{p,y} q^\omega(p_0,y) \\ &\qquad + D_p q^\omega(p_0,y) D_y q^\omega(p_0,y) + q^\omega(p_0,y) \left(D_y q^\omega(p_0,y)\right)^2 . \end{split} \end{equation*} Taking expectations and interchanging differentiation and integration yields \begin{equation*} \begin{split} \mathbb{E}\left[D^2_p h^\omega(p_0, v^\omega_0)\right] &= D_p^2 M_1(p_0,y) + \frac{1}{2} D_{p,y} M_2(p_0,y) + \mathbb{E}\left[q^\omega(p_0,y) \left( D_y q^\omega(p_0,y)\right)^2 \right]. \end{split} \end{equation*} As a consequence of Lemma (ref) in the appendix, the final term on the right-hand side cannot be identified from cross-sectional data: Two observationally equivalent models can yield different values for $\mathbb{E}\left[q^\omega(p_0,y) \left(D_y q^\omega(p_0,y)\right)^2 \right]$. Consequently, the third-order approximation of the average CV is also not identified.

The proof of Proposition (ref) reveals that the source of nonidentification in cross-sectional data stems from terms such as $\mathbb{E}\left[q^\omega(p_0,y) \left(D_y q^\omega(p_0,y)\right)^2 \right]$. To understand the underlying issue, note that by the law of iterated expectations, this term can be rewritten as

equation*[equation* omitted — 226 chars of source]

The right-hand side represents the (noncentered) covariance between the demand bundle and the second moment of the income effect at that demand bundle. Consequently, the failure to identify the third-order approximation of the average CV arises because cross-sectional data do not provide information on how the variance of the income effect varies across demand bundles.

This negative finding is closely linked to fundamental results in the literature on nonseparable models. A direct application of Theorem 2.1 in \citet*{hoderleinmammen} establishes that, in nonseparable models, cross-sectional data identify local average structural derivatives (e.g., $\mathbb{E}\left[ D_y q^\omega(p_0,y) \mid q^\omega(p_0,y)\right]$), but not transformations of these derivatives (e.g., $\mathbb{E}\left[\left( D_y q^\omega(p_0,y)\right)^2 \mid q^\omega(p_0,y)\right]$). As a result, while moments of compensated price derivatives are identified, moments of second-order price derivatives are not. This reasoning extends naturally to higher-order approximations and the many-good case, mutatis mutandis.

Further results and extensions

The results developed in Section (ref) extend beyond the direct analysis of moments of consumer welfare. In this section, we briefly explore extensions and alternative applications, including decomposition of CLIs and estimation of taxable income elasticities.

A quantile-based estimator of consumer welfare

If the objective is to recover the full distribution of consumer welfare, an alternative strategy based on quantile regression can be employed. This approach leverages information in the quantiles of composite demand.\footnote{Dettehoderlein2016 and gunsilius2025nonparametrictestslutskysymmetry leverage quantiles of linear combinations of demand to test for consumer rationality.}

Let $F_q(z \mid \mathbf{p}, y; \mathbf{t}) := \Pr[ q^\omega(\mathbf{p}, y; \mathbf{t}) \leq z \mid \mathbf{p}, y]$ denote the conditional cumulative distribution function (CDF) of composite demand, and let the corresponding conditional quantile function be defined by $K_{q, \tau}(\mathbf{p}, y; \mathbf{t}) := \inf \{z : \tau \leq F_q(z \mid \mathbf{p}, y; \mathbf{t})\}$ for all $\tau \in (0,1)$. Similarly, for budget shares, define the conditional CDF by $F_w(z \mid \mathbf{p}, y; \mathbf{t}) := \Pr[ w_q^\omega(\mathbf{p}, y; \mathbf{t}) \leq z \mid \mathbf{p}, y]$ and the associated conditional quantile function by $K_{w, \tau}(\mathbf{p}, y; \mathbf{t}) := \inf \{z : \tau \leq F_w(z \mid \mathbf{p}, y; \mathbf{t})\}$.

propUnder standard regularity conditions hoderleinmammen, the CDFs of the CV and CLI can be approximated as \begin{equation*} \begin{split} F_{CV}(z \mid \mathbf{p}_0, y; \Delta \mathbf{p}) &\approx \Pr_{\tau}\left[K_{q, \tau}(\mathbf{p}_0, y ; \Delta \mathbf{p}) + S_\tau^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) \leq z \mid \mathbf{p}_0, y; \Delta \mathbf{p}\right], \\ F_{CLI}(z \mid \mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &\approx \Pr_{\tau}\left[K_{w, \tau}(\mathbf{p}_0, y ; \widetilde{\Delta \mathbf{p}}) + S_\tau^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) \leq z \mid \mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}\right], \end{split} \end{equation*} where $\tau \sim U(0,1)$ and \begin{equation*} \begin{split} S_\tau^{CV}(\mathbf{p}_0, y; \Delta \mathbf{p}) &= \frac{1}{2} \left(D_\mathbf{p} K_{q, \tau}(\mathbf{p}_0, y ; \Delta \mathbf{p}) \Delta \mathbf{p} + K_{q, \tau}(\mathbf{p}_0, y ; \Delta \mathbf{p}) D_y K_{q, \tau}(\mathbf{p}_0, y ; \Delta \mathbf{p}) \right), \\ S_\tau^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &= \frac{1}{2} \left(D_{\widetilde{\mathbf{p}}}K_{w, \tau}(\mathbf{p}_0, y ; \widetilde{\Delta \mathbf{p}})\widetilde{\Delta \mathbf{p}} + K_{w, \tau}(\mathbf{p}_0, y ; \widetilde{\Delta \mathbf{p}})D_{\widetilde{y}}K_{w, \tau}(\mathbf{p}_0, y ; \widetilde{\Delta \mathbf{p}}) + K_{w, \tau}(\mathbf{p}_0, y ; \widetilde{\Delta \mathbf{p}})^2\right). \\ \end{split} \end{equation*}
proofSee Appendix (ref).

Proposition (ref) may be viewed as a local extension of the result established by hoderleinvanhemsm, generalized to accommodate multiple goods and unrestricted unobserved heterogeneity. In the two-good case, hoderleinvanhemsm show using conditional quantile demands that exact individual-level welfare analysis is feasible under the restrictive assumption that demand is monotonic in scalar-valued unobserved heterogeneity.\footnote{Under this assumption, the conditional quantiles coincide with individual demand functions, and consumer welfare can be computed by the method of vartia1983efficient.} Proposition (ref) demonstrates that a version of this insight carries over to the more general setting with multiple goods and arbitrary heterogeneity provided that the price changes are sufficiently small. In this local setting, the distribution of consumer welfare can be approximated by means of the average compensated responses of hypothetical consumers located at different quantiles of the composite demand distribution. In particular, we show that all moments of these approximate welfare distributions match those derived in Theorem (ref), up to the same order of approximation. Whether more accurate quantile-based approximations exist remains an open question, which is left for future research.

Decomposition of CLIs

A growing literature raises the point that standard price index formulas suffer from a bias when preferences are nonhomothetic baqaee2022new,jaravel2023measuring,fajgelbaum2016measuring. Theorem (ref) allows us to decompose the bias of standard formulas into the components arising from heterogeneity and nonhomotheticity in preferences.\footnote{Interestingly, in our empirical application in Section (ref), we find that the contribution to the bias from nonhomotheticity is more substantial than that from heterogeneity.} Such a decomposition can be used to assess the relative importance of accounting for the two features.

propThe behavioral term of the aggregate price index can be decomposed as \begin{equation*} \begin{split} \mathfrak{B}_n^{CLI}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) = \sum_{k=1}^4 \mathfrak{D}_{n, k}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}), \end{split} \end{equation*} where \begin{equation*} \begin{split} \mathfrak{D}_{n, 1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &:= \frac{1}{2} \left(D_{\widetilde{\mathbf{p}}} W_1(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})^n\widetilde{\Delta \mathbf{p}} + nW_{1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})^{n+1}\right), \\ \mathfrak{D}_{n, 2}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &:= \frac{1}{2}\left(\frac{n}{n+1} D_{\widetilde{y}} W_{1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})^{n+1}\right), \\ \mathfrak{D}_{n, 3}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &:= \frac{1}{2}\left( D_{\widetilde{\mathbf{p}}} \overline{W}_n(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})\widetilde{\Delta \mathbf{p}} + n\overline{W}_{n+1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})\right),\\ \mathfrak{D}_{n, 4}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}}) &:= \frac{1}{2}\left( \frac{n}{n+1} D_{\widetilde{y}} \overline{W}_{n+1}(\mathbf{p}_0, y; \widetilde{\Delta \mathbf{p}})\right),\\ \end{split} \end{equation*} and $\overline{W}_n(\mathbf{p}_0, y; \mathbf{t}) := {W}_n(\mathbf{p}_0, y; \mathbf{t}) - {W}_1(\mathbf{p}_0, y; \mathbf{t})^n$.
proofSee Appendix (ref).

Table (ref) provides a schematic overview of our decomposition. The term $\mathfrak{D}_{n, 1}$ captures the behavioral component under the assumption of a homothetic representative agent (RA). Adding $\mathfrak{D}_{n, 2}$ relaxes this assumption by allowing the RA to have nonhomothetic preferences. The sum $\mathfrak{D}_{n, 1} + \mathfrak{D}_{n, 2}$ corresponds to the standard RA approximation.

The term $\mathfrak{D}_{n, 3}$ captures the effect of preference heterogeneity in a population with homothetic preferences. That is, $\mathfrak{D}_{n, 1} + \mathfrak{D}_{n, 3}$ yields an approximation for a heterogeneous population, assuming homotheticity. Finally, $\mathfrak{D}_{n, 4}$ captures deviations from homotheticity at the population level. The sum $\mathfrak{D}_{n, 3} + \mathfrak{D}_{n, 4}$ represents the deviation between our full approximation and the RA benchmark.

table[table omitted — 498 chars of source]

Using an RA approach, jaravel2023measuring propose a correction for nonhomothetic preferences based on $D_y \mathbb{E}[CLI^\omega(\mathbf{p}_0, \mathbf{p}_1, y)]$, which captures the sensitivity of the aggregate CLI to income changes. Our approach extends this analysis by incorporating unobserved preference heterogeneity into the same parameter. Formally,

equation*[equation* omitted — 658 chars of source]

Our decomposition allows direct assessment of the relative contributions of nonhomotheticity and heterogeneity in preferences to the overall income sensitivity of the aggregate CLI.

Elasticity of taxable income

A central theme in the public economics literature concerns the measurement of the impact of income tax reforms. If a reform is small, a sufficient-statistic approach based on compensated elasticities delivers a convenient approximation of the welfare cost of taxation feldstein99, grubersaez02, chetty2009.\footnote{Income effects enter the first-order approximations because of the nonlinear nature of the tax schedule. That is, virtual income (see infra) is affected by price or tax changes.} In practice, however, most studies proceed by means of uncompensated elasticities, implicitly assuming that individuals have quasi-linear utilities (e.g., see brunsziliak16 and the references therein). When this assumption is unwarranted, this can bias the efficiency estimates. Our method provides a means of accounting for income effects without increasing data requirements.

Consider a setting with two goods: before-tax income $q^\omega$ and consumption $c^\omega$. An individual derives utility from consumption and disutility from before-tax income, as the latter requires effort to earn. To allow for nonproportional taxation, we assume the tax schedule is piecewise linear. At a linear part of this schedule, the budget constraint is given by $c^\omega = (1-\tau)q^\omega + r$, in which $\tau$ denotes the (local) marginal tax rate and $r$ virtual income.\footnote{Following the literature, we abstract from individuals located at the kinks of the budget set.} Figure (ref) illustrates this relationship, showing the piecewise-linear budget constraint alongside its linear approximation at a specific segment (indicated by the dashed line).

figure[figure omitted — 1,104 chars of source]

feldstein95, feldstein99 shows that the welfare cost of taxation can be summarized by means of the compensated elasticity of taxable income

equation*[equation* omitted — 146 chars of source]

where $h^\omega$ denotes compensated demand for before-tax income. Insights similar to those in Section (ref) allow us to nonparametrically identify and estimate the average of this compensated price elasticity.\footnote{This result is related to the work of \citet*{paluchkneiphildenbrand}, who derive a connection between individual and aggregate income elasticities.}

propThe average compensated elasticity of taxable income can be written as \begin{equation*} \mathbb{E}[\varepsilon^\omega(\tau, v_0^\omega)] = (1-\tau) \Big(D_{1-\tau}\mathbb{E}[\widetilde{q}^\omega(1-\tau,r)] + D_r \mathbb{E}[q^\omega(1-\tau,r)]\Big). \end{equation*}
proofFollowing arguments similar to those above, it holds that \begin{equation*} \begin{split} \mathbb{E}[\varepsilon^\omega(\tau, v_0^\omega)] &= (1-\tau) \mathbb{E}\left[\frac{1}{h^\omega(1-\tau, v_0^\omega)} D_{1-\tau} h^\omega(1-\tau, v_0^\omega) \right] \\ &= (1-\tau) \mathbb{E}\left[\frac{1}{q^\omega(1-\tau, r)}\Big(D_{1-\tau} q^\omega(1-\tau, r) + q^\omega(1-\tau, r) D_r q^\omega(1-\tau,r)\Big)\right] \\ &= (1-\tau) \Big(D_{1-\tau} \mathbb{E}[\widetilde{q}^\omega(1-\tau, r)] + D_r \mathbb{E}[q^\omega(1-\tau,r)]\Big). \end{split} \end{equation*}

Empirical application: Inflation and consumer welfare

This section introduces the data, outlines the estimation, and presents our empirical findings. More precisely, we first apply our results to detailed scanner data on grocery purchases to assess the welfare effects of inflation during the initial COVID-$19$ price shock.\footnote{Our empirical strategy exploits both cross-sectional and intertemporal price variation to identify welfare effects. However, the method is equally applicable in settings with only one source of price variation (e.g., see \citeauthor*{blundellhorowitzparey}, blundellhorowitzparey).} Then, we compare our CLI estimate against a first-order approximation and show that the latter consistently underestimates the true cost of living. Furthermore, we show there is sizable heterogeneity in the cost of living across households, much of which cannot be explained by observed household characteristics. Finally, we quantify the bias introduced by various restrictions on preferences in the CLI. We find that heterogeneity in preferences is crucial for capturing the true cost of living and that nonhomotheticity further eliminates a meaningful and systematic bias.

Sample

Our empirical analysis uses household-level scanner data from NielsenIQ, which track fast-moving consumer goods purchased by a representative panel of households in the United States. This dataset covers a wide range of products, including food items and nonfood categories such as cosmetics, pet care, and office supplies. The panel consists of approximately $60$,$000$ households, with some participating for many years and others entering or exiting the panel over time. Using in-home scanners or a mobile application, panelists record the universal product codes (UPCs) for all the products they purchase. For each transaction, we observe the price paid, quantity purchased, and detailed product characteristics. Additionally, the data include rich household-level demographics, such as household income, household size, and the education level of each head household member.

Our sample spans January $2019$ to December $2022$, allowing us to capture price dynamics before, during, and after the peak of the COVID-$19$ lockdown. We construct monthly, household-specific price indices across food departments as defined by NielsenIQ: dry grocery, frozen foods, dairy, deli, packaged meat, fresh produce, and alcohol. Following the approach of HoderleinMihaleva2008, these indices are linear in prices, with UPC-level prices weighted by household-specific expenditure shares. This method preserves price heterogeneity and captures substitution patterns as the relative prices faced by each household change. Further details about the price aggregation are provided in Appendix (ref).

Our final sample consists of $269$,$593$ household--year--month observations spanning the years $2019$--$2022$. For each observation, we observe prices and expenditures across food departments. We equivalize household income and total expenditures using a modified OECD equivalence scale and remove observations with missing information or extreme values in shares, prices, or equivalized expenditures.\footnote{Equivalized expenditure is computed as $y = y/(1 + 0.5\cdot \mathds{1}(\text{nadults} = 2) + 0.3\cdot \text{nchildren})$, where $y$ is household expenditure, $\mathds{1}(\cdot)$ is the indicator function that equals one if there are two adults in the household (nadults $= 2$), and nchildren is the number of children in the household. The computations of equivalized household income are identical.} The resulting dataset includes only household--year--month observations with nonmissing purchases in each food department category. For expositional simplicity, we may refer to departments as “goods" throughout the paper. Further details on the sample construction are provided in Appendix (ref).

Summary statistics

This section documents prices and expenditures across demographic groups and tracks how prices and expenditure shares have evolved over time across departments. First, Table (ref) reports summary statistics across demographic groups, highlighting several noteworthy consumption patterns.

table[table omitted — 3,353 chars of source]

The results in the table show that log prices are highest among single-person households and decline with household size, consistent with cost savings from bulk purchases. Similarly, average prices increase with household income and education, reflecting preferences for differentiated or higher-quality brands among higher-income and more educated households.

Interestingly, total expenditure exhibits an inverse pattern with income whereby households in higher income groups spend less on average than lower-income households. This seemingly counterintuitive result is driven primarily by differences in household composition. Indeed, higher-income households tend to be smaller and often do not have children, which lowers their overall consumption levels. This is confirmed by the corresponding number of children, which falls to zero for the two highest income brackets.

Table (ref) also shows heterogeneity in prices paid across racial groups. For example, Asian households pay the highest average prices, while Black households pay the lowest, potentially reflecting differences in access to retailers, brand preferences, or store loyalty. Finally, note that the standard deviations in prices and expenditures are typically higher among low-income and larger households, indicating greater variability in their shopping behavior and product selection.

Next, Figure (ref) displays the evolution of average log prices and average shares across goods over time. Panel A shows the evolution of average log prices across goods over time, with temporary dips occurring around December. These seasonal declines are well documented and driven primarily by holiday promotions and year-end inventory clearance. Overall, prices rose rapidly after the onset of the COVID-$19$ pandemic in early $2020$, with cumulative inflation reaching approximately $10$% by the end of $2020$ and an additional $10$% by the end of $2022$.

figure[figure omitted — 1,377 chars of source]

Panel B displays the evolution of average shares across goods over time. With the exception of dry grocery, which represents approximately $40\%$ of household food expenditures, the average share is around $10\%$ for any other good. Furthermore, though the shares appear relatively stable, we observe noticeable shifts during the COVID-$19$ period, especially for dry grocery and alcohol. While the other departments have more modest fluctuations, they still exhibit some adjustments in household spending. Taken together, the two panels show that household consumption adapted in response to the price changes during the pandemic, an important feature to consider when we turn to interpreting welfare impacts.

Estimation and inference

To obtain estimates for average welfare and its spread across the population, we need to model the first three moments of demand.\footnote{We do this using the consumer panel in pooled form, abstracting from its panel structure and treating all observations as a single cross section.} We recover those moments semiparametrically using a generalized additive model (GAM) in prices and expenditure.\footnote{A GAM generalizes generalized linear models by allowing for smooth functions of predictors. See Hastie1986 for an introduction to the theory and Wood2017 for an extensive review and details about its implementation in $R$.} Such an approach is flexible in budget sets while avoiding the curse of dimensionality. In our application, the GAM expresses moments of composite shares as the sum of smooth functions in prices and equivalized expenditure: \[ W_{n}(\mathbf{p}, y; \widetilde{\Delta \mathbf{p}}) = \mathbb{E}[w_q^\omega(\mathbf{p}, y; \Delta \mathbf{p})^n \mid \mathbf{p}, y , \mathbf{x}] =\sum_{j=1}^J f_{jn}(p_{j}) + g_{n}(y) + \mathbf{x}^{\intercal} \boldsymbol{\beta}_{n}, \] where $f_{jn}(\cdot)$ and $g_{n}(\cdot)$ are unknown smooth functions, $\mathbf{x}$ is a set of control variables, and $\boldsymbol{\beta}_{n}$ is a vector of parameters.\footnote{We include age, race, and education as covariates to control for observable heterogeneity.} The functions $f_{jn}(\cdot)$ and $g_{n}(\cdot)$ are represented by means of basis expansions (e.g., cubic splines), and smoothness is controlled through penalization. The estimation then proceeds by minimizing a penalized likelihood:

align[align omitted — 437 chars of source]

where the penalty terms $\lambda_n := (\lambda_{1n}, \lambda_{2n}, \dots, \lambda_{Jn}, \lambda_{yn})^{\intercal}$ control the trade-off between fit and smoothness and $h''(\cdot)$ denotes the second-order derivative of a function $h(\cdot)$. To penalize complexity more effectively, we use the restricted maximum likelihood (REML) approach to estimate the smoothing parameters $\lambda_{n}$.

Following the literature, we use equivalized household income $z$ as an instrument in a control function approach blundell_powell_2003, blundellmatzkin14. Specifically, we regress total equivalized household expenditure on equivalized household income and prices:

equation*[equation* omitted — 141 chars of source]

Then, we add the residual $\widehat{e}$ of this regression as an additional explanatory variable when estimating moments of the composite share. This approach helps eliminate the bias induced by endogeneity in our nonlinear moment equations. We apply our results using consistent plug-in estimators, and all confidence intervals are reported at the $95\%$ level and constructed from a nonparametric bootstrap with $199$ replications.

Welfare analysis

In what follows, we estimate the welfare impacts from the initial COVID-$19$ price shock, defined as the change in log prices faced by households between December $2020$ and December $2021$. Comparing log prices from the same month one year apart allows us to control for natural cyclical price variations and thus to isolate the impacts of the COVID-$19$ pandemic on prices. To allow meaningful interhousehold welfare comparisons, we fix the vector of baseline prices $(\mathbf{p}_{0})$ to the average price in December $2020$ for every household. Finally, we set the change in log prices $(\widetilde{\Delta \mathbf{p}})$ to the difference in average log prices between December $2020$ and December $2021$.\footnote{For Table (ref), we consider a single price change at a time such that the results for a good $j$ are obtained with $\widetilde{\Delta \mathbf{p}} = (0, \dots, 0, \widetilde{\Delta p_j}, 0, \dots, 0)$. For Tables (ref)--(ref), we consider a joint price change such that $\widetilde{\Delta \mathbf{p}} = (\widetilde{\Delta p_1}, \widetilde{\Delta p_2}, \dots, \widetilde{\Delta p_J})$.} When applicable, the price change is demographic specific to capture heterogeneity in prices faced arising from differences in location, store availability, and consumption behavior across demographic groups.

\paragraph{Average welfare effects.}

We begin our analysis by using Theorem (ref) to compute the CLI for each category of goods separately and to assess the bias that would be introduced from using a first-order approximation.\footnote{The first-order approximation is given by $\mathfrak{M}_1^{CLI}$ and represents the welfare impact of the price shock with the share held fixed at its baseline value.} This approach allows us to isolate the welfare impact of each individual price change, thereby clarifying the relative importance of different goods in driving the overall cost of living. The results are reported in Table (ref).

table[table omitted — 1,356 chars of source]

The first column of Table (ref) shows significant heterogeneity in the average compensation needed for households to retain their pre-pandemic utility levels. On the higher end, a typical household would have had to be compensated by as much as $45.44\%$ of its total food expenditures given the price increase in dry grocery to keep its pre-pandemic utility constant. On the lower end, the required compensation is as little as $6.03\%$ for packaged meat. For other goods, the average compensation hovers around $10\%$ of households' total food expenditures. The relatively large CLI for dry grocery likely reflects that it accounts for a disproportionately large share of households' food expenditures, thus amplifying the welfare loss from its price increases.

Next, the second column of Table (ref) presents the bias that would be introduced from using a first-order approximation of the CLI. Given the definition of our first-order approximation, a positive number means that the first-order approximation underestimates the true welfare loss, while a negative number means it overestimates the welfare loss. The results show that a first-order approximations can be highly misleading in estimating the welfare impacts across different categories of goods. In some cases, the first-order approximation dramatically underestimates the true welfare loss---for example, by approximately $44\%$ and $11\%$ for alcohol and dry grocery, respectively. These results highlight that the accuracy of the first-order approximation varies widely across goods, depending on income and substitution effects.

\paragraph{Welfare effects across observed heterogeneity.}

To better assess the welfare loss from the COVID-$19$ price shock, we next evaluate the changes in CLI using a composite share weighted by the change in the goods' log prices. The CLI for this composite share can be interpreted as a measure of a household's overall exposure to the price shock. Since a household's ability to mitigate price increases through substitution or price search may differ systematically across demographics, we consider demographic-specific price shocks that reflect the change in average log price faced by each demographic group.\footnote{The price shock used for the row “All" in Tables (ref)--(ref) is the average change in log price in the whole sample.} Similarly to Table (ref) above, Table (ref) presents the results for the CLI and the bias from the first-order approximation.

table[table omitted — 2,402 chars of source]

The first column shows that the average CLI is $8.47\%$ across the full sample. This implies that, on average, a household would have required approximately $108\%$ of its December $2020$ food expenditures in December $2021$ to maintain the same utility level. Notably, Table (ref) also reveals considerable heterogeneity in CLI values, with point estimates ranging from $-1.59\%$ for Other--low education households to $13.86\%$ for Asian--low education households. The negative CLI for the former subgroup is driven by a sharp decline in the average price of dry grocery\textemdash the department with the largest budget share. Overall, there is no clear pattern in the CLIs across race and education, suggesting that the welfare impacts of the COVID-19 price shock were not systematically linked to these observable characteristics.

The second column shows that the bias from the first-order approximation is $4.28\%$ on average, meaning that the approximation tends to underestimate the welfare losses by approximately $4\%$. As with the CLI, there is substantial heterogeneity in the magnitude of this bias. For example, Asian households with low and mid education levels have a bias of approximately $5\%$ and $-2\%$, respectively, while the bias for White households is at around $4\%$. These disparities reflect significant price and income effects, which cause the true welfare response to diverge from the first-order approximation.

\paragraph{Welfare effects across unobserved heterogeneity.} This section looks into variation in the cost of living within race--education subgroups to uncover heterogeneity driven by unobservable characteristics. To this end, we compute the standard deviation of the CLI from the first two moments of the CLI provided by Theorem (ref). The results for each subgroup are reported in Table (ref).

table[table omitted — 1,651 chars of source]

Table (ref) reveals notable within-group variation in the CLIs across race--education subgroups. The standard deviations range from $1.00$ to $2.89$, generally with greater dispersion for higher education levels. For example, the standard deviations of the CLI are $2.21$ and $2.68$ for Black and Asian households with high education levels while only $1.91$ and $1.08$ for Black and Asian households with low education levels. These results suggest that households within the same subgroups experienced markedly different welfare impacts from the price changes, particularly among the more highly educated.

To gain further insight into the heterogeneity in the cost of living, we exploit Proposition (ref) to recover the CDFs of both the CLI and the first-order approximation. The CDFs are displayed in Figure (ref), where the horizontal axis shows the required percentage increase in total expenditure needed to achieve the utility level from December $2020$ in the face of the joint price shock from December $2020$ to December $2021$. The vertical axis shows the share of households whose welfare loss is at or below each value.

figure[figure omitted — 838 chars of source]

Consistent with our previous results, the two curves in Figure (ref) reveal that the first-order approximation systematically underestimates the welfare loss and that most households are within $2$ percentage point of the average CLI.

\paragraph{Sensitivity to preference restrictions.}

In this section, we take advantage of Proposition (ref) to analyze the sensitivity of the CLI across demographics to various restrictions on preferences. Table (ref) reports the decomposition of the behavioral component of the CLI across race--education subgroups.

table[table omitted — 3,293 chars of source]

Table (ref) shows that every model is within $3\%$ of the true behavioral response when applied to the whole sample ("All"). At the subgroup level, however, the specification with homogeneous and homothetic preferences $(\mathfrak{D}_{1,1})$ does not perform as well. The estimated response is too low at $73.92\%$ for Black households with mid education, and too high at $138.14\%$ for Other households with high education. Allowing for nonhomotheticity while maintaining homogeneity $(\mathfrak{D}_{1,1} + \mathfrak{D}_{1,2})$ improves the estimate for Black households with mid education to $80.12\%$, but in other cases it further inflates the response, reaching $138.29\%$ for Other households with low education. By contrast, introducing heterogeneity while maintaining homotheticity $(\mathfrak{D}_{1,1} + \mathfrak{D}_{1,3})$ yields consistent improvements across all race–education subgroups, with behavioral responses deviating by no more than $5\%$ from the full behavioral response. Taken together, these results highlight heterogeneity as the key first-order feature required to accurately capture cost-of-living adjustments.

Conclusion

In this paper, we introduce a novel method to identify the curvature of the expenditure function through moments of demand. Leveraging this insight can enable more accurate counterfactual exercises in applied welfare analysis. Importantly, we establish that no better approximations can be derived from cross-sectional data. Looking ahead, we highlight several promising directions for future research. First, it would be valuable to investigate the additional identifying power short consumer panels provide for estimating counterfactuals. Initial steps in this direction have been taken by \citet*{crawford19}, \citet*{coopriderhoderleinmeister}, and \citet*{chernozhukovetal}. Second, extending our approach to general equilibrium models could improve the measurement of welfare gains and losses from trade, as explored in baqaeeburstein. Third, further work is needed to deepen our understanding of welfare analysis in models that depart from standard rationality assumptions. Recent contributions by \citet*{apesteguiaballester} and \citet*{aguiarserrano} provide a foundation for such research. Finally, efficient estimation in high-dimensional settings with many goods remains an open challenge. Advances in high-dimensional statistics offer potential solutions to the curse of dimensionality, a problem further exacerbated by the endogeneity of budget sets. In this regard, building on the results of \citet*{chernozhukovhausmannewey} appears to be a particularly promising avenue for future study.

{2pt}

\setcounter{page}{1}

center[center omitted — 126 chars of source]