EconBase
← Back to paper

Individual Welfare Analysis: Random Quasilinear Utility, Independence, and Confidence Bounds

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,294 characters · 19 sections · 40 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Individual Welfare Analysis: Random Quasilinear Utility, Independence, and Confidence Bounds

titlepage\begin{abstract} We introduce a novel framework for individual-level welfare analysis. It builds on a parametric model for continuous demand with a quasilinear utility function, allowing for heterogeneous coefficients and unobserved individual-good-level preference shocks. We obtain bounds on the individual-level consumer welfare loss at any confidence level due to a hypothetical price increase, solving a scalable optimization problem constrained by a novel confidence set under an independence restriction. This confidence set is computationally simple and robust to weak instruments, nonlinearity, and partial identification. The validity of the confidence set is guaranteed by our new results on the joint limiting distribution of the independence test by chatterjee2020. These results together with the confidence set may have applications beyond welfare analysis. Monte Carlo simulations and two empirical applications on gasoline and food demand demonstrate the effectiveness of our method.\\ \\ \noindentKeywords: Welfare analysis, nonlinear models, inferential method, independence.\\ \\ \noindentJEL Codes: C12, C20, C50, D12\\ \end{abstract}

Introduction

Consumer welfare analysis is an important topic in microeconomics. However, commonly adopted methods for welfare analysis are typically at the aggregated or averaged level, unable to recover individual heterogeneity. This paper proposes a novel framework to conduct individual-level welfare inference under a simple demand model. More specifically, we aim to obtain confidence bounds on the welfare loss under a hypothetical price increase for an individual whose consumption level before the price hike is pre-specified.

Our framework builds on a parametric model for continuous demand following brown2007nonparametric and BDW. The model has a quasilinear utility function that allows for multiple goods and unobservables. Structural parameters in the model can depend on observables such as income or can be individual-specific. We derive the welfare loss under a price increase as a function of the parameters. Suppose we have a $(1-\alpha)\in (0,1)$ confidence set for these parameters. We can then infer the maximum and minimum of the welfare loss with confidence $(1-\alpha)$ by solving an optimization problem constrained by the confidence set. We show that under the chosen utility function, the welfare loss function is strictly increasing in each parameter, so the welfare loss bounds are very easy to obtain. Moreover, the function is also concave in the parameters. Therefore, if a researcher wishes to impose additional convex constraints on the parameters, the maximal welfare loss bound can still be quickly computed even for each individual in the sample with a large number of goods.

Individual-level welfare bounds provide a deeper understanding of the data as they allow for the examination of heterogeneity within the data. Unlike a point counterfactual estimate, which is commonly produced through econometric procedures in structural modeling, our bounds are designed to incorporate a specified level of confidence. This means that there is no need for any additional computationally intensive inferential procedures to be conducted at the individual level.

To implement the method, we need first to obtain a confidence set at the desired confidence level for the structural parameters. We propose a new box confidence set under independence. The set is constructed by inverting a recently developed test of independence between two random variables by chatterjee2020. Our confidence set is robust to nonlinearity, weak instruments, and partial identification. It can have other applications in different areas of economics beyond welfare analysis.

To guarantee the desired joint coverage probability of our confidence set for the multidimensional parameters on multiple goods, we prove the following result. For two independent continuous random vectors $\{Z_{k}\}_{k=1}^K$ and $\{W_{k}\}_{k=1}^K$ for some $K>1$, the asymptotic distribution of the vector of Chatterjee's statistics evaluated at $(W_{k},Z_{k})$ for each $k$ is jointly normal with variances all equal to 0.4 and covariances all equal to 0 under a mild condition. Notably, we allow for almost arbitrary correlation among $W_{k}$s and also among $Z_{k}$s across $k$; in particular, we allow some of or all the $Z_{k}$s to be identical. Under such asymptotic independence between these statistics, a multidimensional confidence set can be easily constructed by building multiple one-dimensional confidence intervals, enhancing computation efficiency. The new asymptotic result is general as it does not rely on the demand model we consider. We believe it is of independent interest.

To obtain tight bounds on welfare loss, it is desirable to have a tight confidence set as the constraint. Our proposed confidence set may not always be bounded, similar to the inversion of the Anderson-Rubin statistic. To overcome this issue, we can intersect our confidence set with an alternative confidence set obtained through a different method. In our demand model, we can construct a confidence set for the parameters by 2SLS or OLS with nonlinear transformations. Due to the nonlinear nature of the model, this alternative confidence set may also be large. The intersection of this set and the confidence set by inverting Chatterjee's test can be tighter than either of them, as observed in our empirical applications even with a small sample size.

To control the coverage of the intersection of the two confidence sets, we need to choose the coverage probability for each set appropriately. For this purpose, we derive the joint asymptotic distribution of a vector of Chatterjee's statistics under independence and a large class of asymptotically linear estimators based on which the alternative confidence set is constructed. We show that Chatterjee's statistics are asymptotically independent of the estimators, so joint inference can be again easily carried out.

This paper speaks to two strands of literature, demand estimation with welfare calculation and inference under independence. There is a large literature on demand estimation. See dube2019microeconometric for a review of recent developments. The typical approach constructs a plug-in estimator of welfare change using a point estimator of the parameters of a structural model. However, usually, there is no corresponding confidence interval for the welfare analysis.

For inference under independence, one popular approach is estimation based. That is, first estimate the parameters by exploiting the relationship between the joint and marginal distributions of the unobservable and the exogenous variable that is independent of it, derive the limiting distribution of the estimator, and construct a confidence set based on that. Examples of this approach include manski1983closest, brown2002weighted, BDW, komunjer2010semi, poirier2017efficient and torgovitsky2017minimum. Like any estimation-based inferential approach, the quality of the confidence set typically depends on the quality of the asymptotic approximation of the estimator. Hence, weak instruments (if instrumental variables are needed) or partial identification may complicate the inferential procedure or result in an invalid confidence set.

This paper adopts an alternative approach by inverting a test. Confidence sets obtained by inverting a test, such as the Anderson-Rubin confidence set, are more robust to weak instruments and partial identification. In principle, one can invert any test of independence to obtain a confidence set following the general idea of this paper. Examples of such tests include Hoeffding's $D$ hoeffding1948non, Blum-Kiefer-Rosenblatt's $R$ blum1961distribution, and Bergsma-Dassios-Yanagimoto's $\tau^{*}$ bergsma2014consistent,yanagimoto1970measures. As documented in shi2022power, these tests of independence have better power properties than Chatterjee's test we adopt. We choose Chatterjee’s test for two major reasons. First, it does not require bootstrap/simulation to obtain the critical value. One even does not need to estimate the asymptotic covariance matrix to obtain the critical value since we show that the limiting distribution of a vector of Chatterjee's statistics is jointly normal with known covariance matrix as mentioned before. In contrast, the critical values of all the other aforementioned tests need to be simulated/bootstrapped, adding computation complexity. Specifically, for the scalar case, the limiting distributions of those test statistics under independence are infinite weighted sums of demeaned $\chi_{1}^{2}$ random variables shi2022power; the joint distribution for a vector of such statistics, to the best of our knowledge, is unknown. Second, inverting Chatterjee's test computes much faster than inverting the others. Although all these test statistics including Chatterjee's can be computed in $O(n\log n)$ time in theory, we find in our Monte Carlo simulations that our confidence set by inverting Chatterjee's test computes around 1,000 times faster than inverting the other tests.

The rest of the paper is organized as follows. Section (ref) sets up the demand model and develops our framework of inferring the bounds on the welfare loss with confidence. Section (ref) proposes a new method to construct confidence sets under independence and states the new asymptotic results. Section (ref) presents Monte Carlo simulations. Section (ref) demonstrates two empirical examples of gasoline and food demand. The Appendices provide heuristics on the shape and size of our confidence set, implementation details, some extensions, robustness checks for our empirical applications, and all the proofs.

Individual-Level Welfare Confidence Bounds

We consider the demand model in BDW (henceforth BDW for simplicity). Suppose there are $K>0$ goods of interest. For $k=1,\ldots,K$, let $Y_{k}$ and $P_{k}$ be the consumption and price of good $k$, respectively. Let $Y_{0}$ be the consumption of the num\'eraire good with $P_{0}=1$. Let $W_{k}$ be an unobserved nonnegative random variable that affects the consumer's utility for good $k$. Let the $K\times 1$ vectors $\bm{Y}, \bm{P}$ and $\bm{W}$ collect $Y_{k}$, $P_{k}$ and $W_{k}$ for $k=1,\ldots,K$ respectively. For a consumer with income level $I$, she maximizes her quasilinear utility $U(Y_{0},\bm{Y},\bm{W},\theta)$ subject to the budget constraint:

align[align omitted — 277 chars of source]

where $\bm{\theta}\coloneqq (\theta_{k})\in (0,1)^{K}$. The function $g_{k}$ is a known concave and strictly increasing function for all $k$.

The model (ref) captures good-specific consumer heterogeneity in the sense that it allows for an unobserved preference shock $W_k$ for each good $k$. For concreteness, let $g_{k}(\cdot) = \log(\cdot)$ for each $k$. This choice of $U_{0}$, together with $Y_{0}$ and $\bm{W}'\bm{Y}$, makes $U$ the logarithm of some perturbed Cobb-Douglas utility function. We also allow $U_{0}$ to take other functional forms, or in principle allow $\bm{W}$ to be nonseparable. We provide more details on these extensions in Appendix (ref) and Remark (ref), respectively.

On the other hand, we maintain linearity in $Y_{0}$ even in those extensions. Such linearity makes the utility we consider a special case of a random quasilinear utility function proposed by brown2007nonparametric that develops theoretical results for quasilinear rationalizations of consumer demand data. Under a quasilinear utility function, welfare changes due to a shift in price can be equivalently measured by the change in consumer surplus, compensating variation (CV), or equivalent variation (EV): all the three measures coincide and are equal to the difference in the indirect utility functions at the new and old prices. This property is useful to gain computational efficiency for the empirical method we propose.

One major limitation of quasilinearity utility is the absence of the income effect. Similar to the idea of griffith2018income that allows income to enter the utility function in a flexible way to capture income effects in a discrete choice model, we mitigate this issue of missing income effect by allowing $\bm{\theta}$ to depend on income groups. We consider such a model in our first empirical application in Section (ref). Alternatively, when a panel data set is available, we can allow $\bm{\theta}$ (and more fundamentally $g_{k}$s) to differ across individuals but constant across time. In this case, the individual-wise income effect can exist; we consider this model in the empirical application in Section (ref).

We now derive the welfare loss function. Assuming an interior solution for the optimal $Y_{0}$ and $\bm{Y}$, BDW derives the following first-order condition for the optimal consumption $\bm{Y}^{*}(\bm{P},\bm{W};\bm{\theta})$:

equation[equation omitted — 117 chars of source]

The model exploits separability, so that the first order condition for each good $k$ only involves $(W_k, P_k, \theta_k, Y^{*}_k)$ but not those from other goods. It follows from (ref) that the demand for good $k$ is

align[align omitted — 94 chars of source]

Substituting the budget constraint (ref) and the optimal demand (ref), we then obtain the following indirect utility function from (ref):

align[align omitted — 152 chars of source]

We now focus on a change in consumer welfare due to a change in prices: the difference in the indirect utility function at the old and new prices by quasilinearity. Specifically, let $\bar{\bm{w}}$ be a vector of realized $\bm{W}$ and ($\bm{p}^{0}$, $\bm{p}^{1}$) be two price levels; neither $\bm{p}^{0}$ nor $\bm{p}^{1}$ needs to be the actual observed price. Denote the optimal consumption of goods $1,\ldots,k$ under $\bm{p}^{j}$ and $\bar{\bm{w}}$ by $\bm{y}^{j}$ and the optimal consumption of the num\'eraire by $y_{0}^{j}$; by equation (ref), $\bm{y}^{j}\coloneqq \bm{Y}^{*}(\bm{p}^{j},\bar{\bm{w}};\bm{\theta})$, $j=0,1$. Assume that $I$ is sufficiently large and $\bar{\bm{w}}$ is smaller than both $\bm{p}^{0}$ and $\bm{p}^{1}$ so that the optimal consumption levels are interior solutions. These are automatically satisfied if $y_{0}^{0},\bm{y}^{0}>\bm{0}$ and $\bm{p}^{1}\geq \bm{p}^{0}$; see Remark (ref) for details. Let $\bm{\Delta} := \bm{p}^{1} - \bm{p}^{0}$ be the vector of price changes. By the definition of $\bm{y}^{0}$ and $\bm{y}^{1}$, equation (ref) implies that the welfare loss denoted by $\text{WL}(\bm{\theta})$, when the price changes from $\bm{p}^{0}$ to $\bm{p}^{1}$ while $\bm{W}$ is fixed at $\bar{\bm{w}}$, is as follows:

align[align omitted — 363 chars of source]

Hence, the welfare loss solely depends on $\bm{y}^{0}$, the price change $\bm{\Delta}$, and the unknown parameters $\bm{\theta}$.

So far, all the prices in equations (ref)-(ref) are real terms under the normalization that the price of the num\'eraire is 1. In most applications, however, we can only observe nominal prices in data. Denoting the nominal price for good $k=0,\ldots,K$ by $P_{k}^{n}$, we have $P^{n}_{k}/P_{k}\equiv P^{n}_{0}/1$ for every $k$. Equation (ref) written in terms of the nominal prices then becomes $P_{k}^{n}-P_{0}^{n}W_{k}=P^{n}_{0}\theta_{k}/Y_{k}$. If we only observe $P_{k}^{n}$ and $Y_{k}$ in data, we can in fact only treat $W_{k}^{n}\coloneqq P_{0}^{n}W_{k}$ as the unobservable and $\theta^{n}_{k}\coloneqq P^{n}_{0}\theta_{k}$ as the parameter. Consequently, although the real $\theta_{k}$ is assumed to lie in $(0,1)$ by BDW, we do not restrict the confidence interval to be upper bounded by 1 in the empirical applications: The parameter that data can recover is $\bm{\theta}^{n}$, which is $P^{n}_{0}$ times of the real $\bm{\theta}$ and $P^{n}_{0}$ is unknown.

As a result, if we do not observe the real prices in a data set, the welfare loss function can be at most identified up to scale. To see it, plugging $\bm{\theta}^{n}$ and the observable prices $\bm{P}^{n}$ into the welfare loss function by treating them as the actual parameters $\bm{\theta}$ and the real prices $\bm{P}$, we obtain $\text{WL}^{n}(\bm{\theta}^{n})$ which is proportional to the actual welfare loss:

align[align omitted — 439 chars of source]

where $\Delta_{k}^{n}$ is the nominal price change. Hence, the welfare loss that we can make inferences about based on nominal price data is $P_{0}^{n}$ times the actual welfare loss.

Nonetheless, we can still identify the following standardized welfare loss function based on nominal data:

align*[align* omitted — 280 chars of source]

where $\|\cdot\|$ is the Euclidean norm of a vector. Equality (1) is by equation (ref). The standardized welfare loss $\text{WL}^{st}(\bm{\theta})$ measures the welfare loss given a unit price change. The above equation shows that the standardized welfare loss is invariant to how the prices are measured. In the rest of the paper, we focus on finding confidence bounds on the standardized welfare loss. Since we can identify it with nominal prices, with a bit of abuse of notation, we still denote the nominal price $P^{n}_{k}$ and the parameter $\theta_{k}^{n}$ by $P_{k}$ and $\theta_{k}$ respectively for simplicity.

remarkThe model in general can have corner solutions: Optimal consumption of the num\'eraire good is zero if $W_{k}\geq P_{k}$ for some $k$ or $I\leq \sum_{k}(P_{k}\theta_{k}/(P_{k}-W_{k}))$. In that case, the model is effectively no longer quasilinear and our welfare loss function is invalid. For good of interest $k=1,\ldots,K$, on the other hand, if the derivative of $g_{k}(\cdot)$ at 0, denoted by $g'_{k}(0)$, is well-defined, then zero consumption occurs when $W_{k}+g_{k}'(0)\leq P_{k}$. In this case, since we cannot write down a reduced form that expresses $\bm{W}$ as a function of the observables and parameters, we cannot form moment equations for the parameters. We focus on interior solutions throughout the paper and leave the study of corner solutions to future research. Hence, for simplicity, we pick $g_{k}(\cdot)=\log(\cdot)$ and assume $W_{k}<P_{k}$ for all $k$ with sufficiently large $I$ such that $I> \sum_{k}(P_{k}\theta_{k}/(P_{k}-W_{k}))$, then all the optimal $Y_{0}^{*},Y_{1}^{*},\ldots,Y_{K}^{*}$ are strictly positive.
remarkOur perturbed Cobb-Douglas utility function is related to the perturbed utility model (PUM) in allen2019identification with several key differences. In our notation, PUM assumes a more restrictive $U_{0}(\bm{Y};\bm{\theta})=\sum_{k}Y_{k}\theta_{k}$ whereas our $U_{0}$ can be nonlinear in $Y_{k}$s. On the other hand, our $\bm{W}'\bm{Y}$ is replaced in their work by a more general nonparametric function $D(\bm{Y},\bm{W})$. Their $\theta_{k}$s are functions of regressors; identification of $\theta_{k}$s requires that for every $k$, there exists a continuous $k$-specific regressor only entering $\theta_{k}$. In contrast, our approach does not rely on continuous variation in the regressors: in our baseline model, the $\theta_{k}$s do not depend on any regressor, whereas in Section (ref), they only depend on some discrete variation of a regressor. Another benefit of adopting the linear form for $\bm{W}$ for our purpose is that the welfare loss function does not depend on the realization of $\bm{W}$, simplifying the analysis. To accommodate a model that is nonseparable in $\bm{W}$ is possible. In that case, the welfare loss function in general depends on the realization of $\bm{W}$. One can either set $\bm{W}$ to be a counterfactual value or use its actual realization in a data set if that can be backed out using the observed consumption and price.

Predicting the Individual Welfare Loss Bounds with Confidence

Suppose that we have micro-level i.i.d. data on $\{(P_{k,i}, Y_{k,i}, {\bm{X}_{i}}):i=1,\ldots,n;k=1,\ldots,K\}$, where subscript $i$ refers to observation $i$. Vector $\bm{X}_{i}$ contains other observables in the data. For a given price change $\bm{\Delta}$ {and consumption level $\bm{y}^{0}$}, it remains to deal with the unknown $\bm{\theta}$ {to compute the welfare change}. {In principle, we allow $\bm{\theta}$ to depend on $i$ as long as there are sufficient data to make inference about it; each $\theta_{k}$ can be either a nonparametric function of the other observables $\bm{X}_{i}$, or contain $i$-specific fixed coefficients; in that case, we allow $\bm{y}^{0}$ and $\bm{\Delta}$ to be $i$-specific as well. In our empirical applications, the varying coefficient model in Section (ref) and the model in Section (ref) follow these two setups respectively. For notational simplicity, we treat $\bm{\theta}$, $\bm{\Delta}$ and $\bm{y}^{0}$ as homogeneous across $i$ in this section and in Section (ref).}

For a given $\alpha\in (0,1)$, assume that we have a $(1-\alpha)$ confidence set for $\bm{\theta}$. We can then predict the maximal and minimal standardized welfare loss with confidence $(1-\alpha)$ for an individual with consumption {$\bm{y}^{0}\equiv\{y^{0}_{1},\ldots,y^{0}_{K}\}$} due to a price change $\bm{\Delta} \equiv (\Delta_1,\ldots,\Delta_K)$. Denoting the confidence set by $CS(1-\alpha)$, we solve the following problem:

align[align omitted — 302 chars of source]
thmLet the maximum and the minimum of the objective function under the constraint above be $\overline{\mathrm{WL}}^{st}(1-\alpha)$ and $\underline{\mathrm{WL}}^{st}(1-\alpha)$, respectively. For each $k=1,\ldots,K$, if $\Pr(\bm{\theta}\in CS(1-\alpha))\to (1-\alpha)$ with the sample size $n\to \infty$, then \begin{align*} \lim_{n\to\infty}\Pr\left(\mathrm{WL}^{st}(\bm{\theta})\leq\overline{\mathrm{WL}}^{st}(1-\alpha)\right)\geq& (1-\alpha),\\ \lim_{n\to\infty}\Pr\left(\mathrm{WL}^{st}(\bm{\theta})\geq\mathrm{WL}^{st}(1-\alpha)\right)\geq&(1-\alpha),\\\ \lim_{n\to\infty}\Pr\left(\mathrm{WL}^{st}(1-\alpha)\leq\mathrm{WL}^{st}(\bm{\theta})\leq\overline{\mathrm{WL}}^{st}(1-\alpha)\right)\geq& (1-\alpha). \end{align*}

In the next section, we will propose a method to construct a box, or, a rectangular confidence set $CS(1-\alpha)$. The optimization problem will then be very convenient to solve. As shown in Appendix (ref), the welfare loss $\text{WL}(\tilde{\bm{\theta}})$ as a function of $\tilde{\bm{\theta}}\in\mathbb{R}^{K++}$ is strictly increasing in each $\tilde{\theta}_{k}$ on $(\max\{0,-\Delta_{k}y^{0}_{k}\},\infty)$ for all $\Delta_{k}y_{k}^{0}\neq 0$ and $\Delta_{k}y_{k}^{0}>-\theta_{k}$. Hence, for instance, under a price increase such that all $\Delta_{k}$s are positive, if $CS(1-\alpha)$ is rectangular as the cartesian product of confidence intervals of the $\theta_{k}$s, one can easily obtain the maximum and minimum of the welfare loss by substituting the upper and lower bounds of those intervals, respectively. When there are additional convex constraints besides the convex confidence set, for instance, $\sum_{k=1}^{K}\tilde{\theta}_{k}=1$, one can still solve the maximization problem quickly by off-the-shelf convex optimization algorithms because the objective function is concave in $\bm{\theta}$; see Appendix (ref) for more details.

Individual-level welfare bounds contain more information than a bound at the aggregated level; one can view and study the heterogeneity from them. Moreover, our bound is different from a point estimate; by construction, it incorporates the desired confidence level, so no further inferential procedure is needed.

remarkWe only consider a price increase in this paper, i.e., $\Delta_{k}\geq 0$ for all $k>0$ with the inequality being strict for some $k$. This is a simple sufficient (but not necessary) condition to make sure that if $y_{0}^{0}>0$ and $\bm{y}^{0}>\bm{0}$ satisfies equation (ref), meaning that the optimal consumption including the num\'eraire good is an interior solution at some income level, then $\bm{y}^{1}$ also satisfies equation (ref). Consequently, the welfare loss function we derive is valid. To see this, $\bm{y}^{0}$ satisfying equation (ref) implies that the realization of the unobservable $\bar{\bm{w}}$ satisfies $\bm{p}^{0}>\bar{\bm{w}}$. Hence, if the new prices $\bm{p}^{1}>\bm{p}^{0}$, we have $\bm{p}^{1}>\bar{\bm{w}}$. Meanwhile, the total expenditure for the $K$ goods of interest $\sum_{k}(p_{k}^{1}\theta_{k}/(p_{k}^{1}-w_{k}))\leq \sum_{k}(p_{k}^{0}\theta_{k}/(p_{k}^{0}-w_{k}))$. So, the new optimal consumption including the num\'eraire good is an interior solution under the original budget. Nonetheless, we can analyze the welfare change under a price decrease, too, so long as the optimal consumption at the new price is still an interior solution. To satisfy the interior solution requirements, we first need $p_{k}^{1}-\bar{w}_{k}=\Delta_{k}+p_{k}^{0}-\bar{w}_{k}>0$, which leads to $\Delta_{k}>-\theta_{k}/y_{k}^{0}$; this is exactly the condition that makes the right-hand side of equation (ref) well-defined. So, for a price drop ($\Delta_{k}<0$), if we have a box confidence set $CS(1-\alpha)$ formed by the cartesian product of confidence intervals for each $\theta_{k}$, then we can compute the welfare loss bounds under $\Delta_{k}$ as long as $\Delta_{k}>-\text{LB}_{k}/y^{0}_{k}$ where $\text{LB}_{k}>0$ is the lower bound of the confidence interval for $\theta_{k}$. In the meantime, we assume that the budget $I$ is sufficiently high, greater than the expenditure of the $K$ goods of interest under the new price.

A New Confidence Set for Nonlinear Models

So far, we have assumed the existence of a box confidence set for $\bm{\theta}$. In this section, we introduce a novel approach to obtain it. This approach is based on a new correlation coefficient to measure independence developed by chatterjee2020. We first briefly review this statistic for completeness in Section (ref). We then derive a new result about the joint limiting distribution of this coefficient for multiple pairs of independent random variables in Section (ref); this result enables us to invert individual test statistics to construct a joint box confidence set for $\bm{\theta}$. Section (ref) derives the joint limiting distribution of Chatterjee's correlation coefficients and a general class of estimators. Built on it, we propose an intersection method to sharpen the confidence set.

Review of Chatterjee's Correlation Coefficient

Consider an i.i.d. sample $\{(R^{(1)}_{i},R^{(2)}_{i}):i=1,\ldots,n\}$ for random variables $R^{(1)}$ and $R^{(2)}$. Suppose that the $R^{(1)}_{i}$s and $R^{(2)}_{i}$s have no ties.\footnote{When there are ties in the $R_{i}^{(1)}$s, chatterjee2020 proposes to break the ties uniformly at random. When there are ties in the $R^{(2)}_{i}$s, define $l_{i}$ as the number of $j$ such that $R^{(2)}_{(j)}\geq R^{(2)}_{(i)}$. Then redefine $\xi_{n}(R^{(1)},R^{(2)})\coloneqq 1-n\sum_{i=1}^{n-1}|r_{i+1}-r_{i}|/[2\sum_{i=1}^{n}l_{i}(n-l_{i})]$.} Sort the data by $R^{(1)}$ in ascending order. Let the rearranged data be $(R^{(1)}_{(1)},R^{(2)}_{(1)}),\ldots,(R^{(1)}_{(n)},R^{(2)}_{(n)})$ where $R^{(1)}_{(1)}<\cdots<R^{(1)}_{(n)}$. Let $r_{i}$ be the rank of $R^{(2)}_{(i)}$. chatterjee2020 proposes the following statistic to measure the dependence between $R^{(1)}$ and $R^{(2)}$:

equation[equation omitted — 116 chars of source]

When $R^{(1)}\perp R^{(2)}$, where $\perp$ denotes independence, Chatterjee shows in his Theorem 1.1 that $\xi_{n}(R^{(1)},R^{(2)})\overset{a.s.}{\to}0$. Meanwhile, $R^{(1)}\not\perp R^{(2)}$ if and only if $\text{plim } \xi_{n}(R^{(1)},R^{(2)})>0$. Moreover, under $R^{(1)}\perp R^{(2)}$, his Theorem 2.2 shows that

equation*[equation* omitted — 86 chars of source]

if $R^{(2)}$ is continuous. When $R^{(2)}$ is discrete, the asymptotic variance needs to be estimated and Chatterjee provides a consistent estimator for it which can be computed in time $O(n\log n)$.

Joint Limiting Distribution

For $k=1,\ldots,K$, let $Z_{k}$ be an observed exogenous random variable for which we also have an i.i.d. sample. Depending on the application, variable $Z_{k}$ can be price itself or its instrumental variable. Let the short-hand notation $\xi_{n}^{(k)}$ denotes $\xi_{n}(W_{k},Z_{k})$. Let $\bm{\xi}_{n}\coloneqq (\xi_{n}^{(1)},\ldots,\xi_{n}^{(K)})$. Chatterjee establishes the limiting distribution of $\xi^{(k)}_{n}$ when $W_{k}\perp Z_{k}$. For joint inference of $\bm{\theta}$, we now derive the asymptotic distribution of $\bm{\xi}_{n}$.

assum$\{(W_{1,i},\ldots,W_{K,i},Z_{1,i},\ldots,Z_{K,i}):i=1,\ldots,n\}$ are i.i.d.
assumi) Random vector $(W_{1},\ldots,W_{K},Z_{1},\ldots,Z_{K})$ is continuously distributed. ii) $(W_{1},\ldots,W_{K})\perp (Z_{1},\ldots,Z_{K})$.
remarkWe do not require there to be $K$ different $Z_{k}$s; the current setup is for convenience only. We allow that some of or even all these $Z_{k}$s are equal.

The continuity requirement for the unobservables $W_{k}$ in Assumption (ref)-i) is common in economics. For $Z_{k}$, we only focus on continuous instruments in the applications of this paper. Asymptotic theory for the case of discrete instruments is left for future research. Under this assumption, there are no ties in a data set with probability one. For Assumption (ref)-ii), although the independence assumption is strong, it is indispensable in many nonseparable models, which, in principle, we could handle as mentioned in Remark (ref).

We now introduce an assumption regulating the dependence structure of the $W_{k}$s. For $k=1,\ldots,K$, let $\pi_{k}:\{1,\ldots,n\}\mapsto\{1,\ldots,n\}$ be a random variable such that $\pi_{k}(i)$ indicates the location of the original index $i$ in the permuted index set after arranging $W_{k,1},\ldots,W_{k,n}$ in ascending order. When there are no ties among $W_{k,i}$s, which has probability one under Assumption (ref)-i), $\pi_{k}(i)$ is the number of $W_{k,j}$s satisfying $W_{k,j}\leq W_{k,i}$. The mapping $\pi_{k}$ in this case is one-to-one. When there are ties, we can still uniquely define a one-to-one $\pi_{k}$ such that for any $W_{k,i}=W_{k,j}$ with $i\neq j$, we let $\pi_{k}(i)<\pi_{k}(j)$ if and only if $i<j$.

Notation-wise, to distinguish the post-permutation indices under permutation $(\pi_{k}(1),\ldots,\pi_{k}(n))$ from the original ones, denote the post-permutation indices by $(i)_{k}$. In other words, $W_{k,(1)_{k}}\leq W_{k,(2)_{k}}\leq \cdots\leq W_{k,(n)_{k}}$. For example, suppose the $l$-th smallest element in $W_{k,i}$s (assuming there are no ties) is $W_{k,j}$, then $\pi_{k}(j)=(l)_{k}$.

assumFor any $k\neq k'$, $$\Pr\left(\left|\pi_{k}(i)-\pi_{k}(j)\right|=1\cap \left|\pi_{k'}(i)-\pi_{k'}(j)\right|=1\right)=o\left(\frac{1}{n}\right),\forall 1\leq i<j\leq n.$$

Assumption (ref) says that for any original indices $i$ and $j$, if they are adjacent after arranging the $W_{k,i}$s in ascending order, then they are not likely to be adjacent after arranging the $W_{k',i}$s in ascending order. To be more precise, note that under Assumption (ref),

equation[equation omitted — 116 chars of source]

see Appendix (ref) for a proof. Then, Assumption (ref) is equivalent to $$\Pr\left(\left|\pi_{k'}(i)-\pi_{k'}(j)\right|=1\Big| \left|\pi_{k}(i)-\pi_{k}(j)\right|=1\right)=o\left(1\right),\forall 1\leq i<j\leq n.$$

Assumption (ref) allows for correlated $W_{k}$s, but requires that $W_{k'}$ still has sufficient variation conditional on $W_{k}$. For further illustration, we now provide a sufficient condition for it, an example under which the sufficient condition holds, and a counterexample where Assumption (ref) fails. Let $F_{W_{k}}$ denote the cumulative distribution function of $W_{k}$ for all $k$.

propLet $V_{k}\coloneqq F_{W_{k}}(W_{k})$ and $V_{k'}\coloneqq F_{W_{k}'}(W_{k}')$. If there exists some constant $C>0$ such that the following holds for all $v,v'\in (0,1)$: \begin{equation} \lim_{n\to\infty}\Pr\left(V_{k'}\in \left[v'-C\sqrt{\frac{\log\log{n}}{n}},v'+C\sqrt{\frac{\log\log{n}}{n}}\right]\Bigg|V_{k}=v \right)= 0, \end{equation} then Assumption (ref) holds under Assumptions (ref) and Assumption (ref)-i).

Condition (ref) is mild. It says that conditional on any value of $V_{k}$, there is no atom in the distribution of $V_{k'}$. Note that Proposition (ref) does not rule out nonrectangle support for $(W_{k},W_{k'})$: For $v'$ such that $F^{-1}_{W_{k'}}(v')$ is no longer in the support of $W_{k'}|W_{k}=F_{W_{k}}^{-1}(v)$, condition (ref) trivially holds.

We now provide an example and a counterexample.

exmCondition (ref) holds if $(W_{k},W_{k'})$ are jointly normal for any correlation coefficient $\rho_{k,k'}\in(-1,1)$. This is because conditional on any value of $W_{k}$, the random variable $W_{k'}$ is still continuously distributed on the entire real line as a normal distribution with variance equal to $(1-\rho_{k,k'}^{2})$ times its original variance.
exmSuppose $W_{k}\sim \mathcal{N}(0,1)$. Let $D\sim Ber(0.5)$ be a dummy variable and $\widetilde{W}\sim \mathcal{N}(0,1)$. Assume $W_{k},D$ and $\widetilde{W}$ are mutually independent. Now suppose $W_{k'}=W_{k}D+\widetilde{W}(1-D)$. Then one can verify that for all $i\neq j$, \begin{align*} &\Pr\left(\left|\pi_{k}(i)-\pi_{k}(j)\right|=1\cap \left|\pi_{k'}(i)-\pi_{k'}(j)\right|=1\right) \\ =&\frac{1}{2}\Pr\left(\left|\pi_{k}(i)-\pi_{k}(j)\right|=1|D=1\right)+\frac{1}{2}\Pr\left(\left|\pi_{k}(i)-\pi_{k}(j)\right|=1\cap \left|\pi_{k'}(i)-\pi_{k'}(j)\right|=1|D=0\right)\\ =&\frac{1}{2}\cdot\frac{2}{n}+\frac{1}{2}\cdot\left(\frac{2}{n}\right)^{2} \neq o\left(\frac{1}{n}\right), \end{align*} where the first equality is by the law of iterated expectation, by the distribution of $D$, and by $W_{k'}=W_{k}$ when $D=1$. The second equality is by mutual independence among $D,W_{k},\widetilde{W}$, by $W_{k'}=\widetilde{W}$ under $D=0$, and by equation (ref). One can verify that condition (ref) does not hold, too.

These examples show that we can allow for arbitrarily high correlation among the unobservables, so long as there is no atom in the conditional distribution.

We are now ready to state our main result.

thmUnder Assumptions (ref) to (ref), \begin{equation} \sqrt{\frac{n}{0.4}}\bm{\xi}_{n}\overset{d}{\to}\mathcal{N}(\bm{0},I_{K}), \end{equation} where $I_{K}$ is the $K\times K$ identity matrix.

Theorem (ref) says that under our assumptions, the $\sqrt{n}\xi^{(k)}_{n}$s are asymptotically jointly normal and mutually independent.

Three remarks are in order. First, Assumption (ref) is not needed to obtain joint normality; it is a sufficient condition to obtain zero limiting covariances. So, for instance, if one believes that there is only one unobservable for all different goods, i.e., $W_{1}=\cdots = W_{K}$, then joint normality still holds but the asymptotic covariances may need to be estimated. Second, Assumption (ref) is silent about the dependence structure of the instruments $Z_{k}$s. So Theorem (ref) holds even if there is only one instrument; see also Remark (ref). Finally, Theorem (ref) is independent of our demand model; it holds for any independent random pairs satisfying Assumptions (ref) to (ref).

Based on Theorem (ref), we propose the following procedure to construct a $(1-\alpha)$ joint confidence set for $\bm{\theta}$. Suppose there is a known parameter space $\Theta_{k}$ for each $\theta_{k}$. Let $z_{(1-\alpha)^{1/K}}$ denote the $(1-\alpha)^{1/K}$-th quantile of the standard normal distribution. For each $k$, construct

equation[equation omitted — 205 chars of source]

Then construct the joint confidence set by

equation[equation omitted — 133 chars of source]

By Theorem (ref), the coverage probability for each $CS_{k;\alpha}^{\xi}$ is asymptotically $(1-\alpha)^{1/K}$ and the joint coverage probability is approaching $(1-\alpha)$. We formally state this result in the following corollary.

corUnder Assumptions (ref)-(ref), for $CS(1-\alpha)$ constructed by equation (ref), \[ \Pr\left(\bm{\theta}\in CS(1-\alpha) \right)\to 1-\alpha. \]

Note that Corollary (ref) does not require any relevance condition. The validity of the confidence set holds regardless of the strength of the instrument. Even if the model is only partially identified or the instrument is weak, $CS(1-\alpha)$ always yields the correct asymptotic coverage probability for $\bm{\theta}$. Moreover, one does not need to distinguish the point and partially identified cases. On the other hand, the strength of the instrument does affect the size of the confidence set. For instance, if $Z_{k}$ is independent of both $P_{k}$ and $W_{k}$, then one can see that $CS^{\xi}_{k;\alpha}$ can be the entire real line. Appendix (ref) provides more details on the shape and size of the confidence set.

In practice, one can do a grid search over $\Theta_{k}$ for each $k$ and take the minimum and maximum grid nodes that fall into the set (ref) to get a confidence interval. The joint box confidence set can be constructed as the cartesian product of these intervals. Such a construction enables fast and scalable computation even when $K$ is large because we only need to construct $K$ individual confidence intervals by one-dimensional grid search.

Intersecting Confidence Sets

One possible drawback of Chatterjee's test of independence is its low power shi2022power. As a consequence, the sets $CS^{\xi}_{k;\alpha}$s could be wide. In this section, we propose an intersection approach to mitigate the low power problem and tighten the confidence set.

For the demand model we consider, we can construct a confidence set for $\bm{\theta}$ based on the estimation approach instead of by inverting Chatterjee's $\xi$; that is, first estimate $\bm{\theta}$ and then build a confidence set based on the estimator. Specifically, rearranging equation (ref), we have $1/Y_{k}=\beta _{k}P_{k}-\beta_{k}W_{k}$ where $\beta_{k}=1/\theta_{k}$ for each $k$. Treating $1/Y_{k}$ as the dependent variable, we can estimate $\beta_{k}$ by OLS or seemingly unrelated regression (SUR) if we assume the prices are exogenous, or by 2SLS or multi-equation GMM if the prices are endogenous but we have instruments for them. One can then, for instance, obtain a sup-t confidence band for $\bm{\theta}$ with any desired coverage probability by the delta method and the procedure described in montiel2019simultaneous.

Now we can obtain a potentially tighter confidence set by intersecting the one introduced in Section (ref) with an estimation-based confidence set. To control the coverage probability, we derive the joint limiting distribution of $(\bm{\xi}_{n},\hat{\bm{\theta}})$ where estimator $\hat{\bm{\theta}}$ comes from a class of estimators specified in Assumption (ref) below. Let $\widetilde{W}_{k,i}\coloneqq W_{k,i}-\mathbb{E}(W_{k,i})$ and $\widetilde{\bm{W}}_{i}\coloneqq(\widetilde{W}_{1,i},\ldots,\widetilde{W}_{K,i})'$. Let $\bm{Z}_{i}\coloneqq(Z_{1,i},\ldots,Z_{K,i})'$. We assume the estimator is asymptotically linear with the following influence function:

assumThere exist functions $\omega_{i}:\mathbb{R}^{K}\mapsto\mathbb{R}^{K\times K}$ such that \begin{equation} \sqrt{n}\left(\hat{\bm{\theta}}-\bm{\theta}\right)=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\omega_{i}(\bm{Z}_{i})\widetilde{\bm{W}}_{i}+o_{p}(1), \end{equation} and \begin{equation} Var\left(\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\omega_{i}(\bm{Z}_{i})\widetilde{\bm{W}}_{i}\right)\overset{p}{\to}\Sigma, \end{equation} where the $K\times K$ matrix $\Sigma$ is positive definite and consistently estimable.

One can verify that a large class of estimators including all the aforementioned estimators satisfy Assumption (ref) under certain regularity conditions. The following Theorem (ref) says that $\bm{\xi}_{n}$ and such an estimator $\hat{\bm{\theta}}$ are asymptotically independent under regularity conditions. Let $\omega_{(k,k'),i}$ be the $(k,k')$-th entry in matrix $\omega_{i}$.

thmUnder Assumptions (ref)-(ref), if $\max_{k,k'}\mathbb{E}\left(\omega_{(k,k'),i}(\bm{Z}_{i})\widetilde{W}_{k',i}\right)^{8}<\infty$, then \begin{equation} \sqrt{n}\begin{pmatrix}\frac{\bm{\xi}_{n}}{\sqrt{0.4}}\\ \hat{\bm{\theta}}-\bm{\theta}\end{pmatrix}\overset{d}{\to}\mathcal{N}\left(0,\begin{pmatrix}I_{K}&\bm{0}\\\bm{0}&\Sigma \end{pmatrix}\right). \end{equation}
remarkAsymptotic normality and independence between $\bm{\xi}_{n}$ and $\hat{\bm{\theta}}$ hold even if Assumption (ref) is dropped. That is, without Assumption (ref), equation (ref) still holds by only replacing $I_{K}$ on the right-hand side with a possibly more complicated $K\times K$ matrix.
remarkThe existence of the eighth moment is required by the central limit theorem in chatterjee2008new that we adopt; it is assumed only because of $\hat{\bm{\theta}}$.

Theorem (ref) provides a simple way to construct a $(1-\alpha)$ confidence set for $\bm{\theta}$. Suppose we have a $\sqrt{1-\alpha}$ box confidence set for $\bm{\theta}$ based on Assumption (ref), formed as the cartesian product of individual intervals $\left[\hat{\theta}_{k}-c_{k;\alpha}/\sqrt{n},\hat{\theta}_{k}+c_{k;\alpha}/\sqrt{n}\right]$ where the positive constants $c_{k;\alpha}$s are known or can be consistently estimated. One way to obtain such a confidence set is by the sup-t method by montiel2019simultaneous, which has exact $\sqrt{1-\alpha}$ asymptotic coverage. Then, one can construct an intersected confidence set $CS(1-\alpha)$ as follows: For each $k$, construct

equation[equation omitted — 310 chars of source]

Then construct the joint confidence set by

equation[equation omitted — 140 chars of source]
corUnder Assumptions (ref)-(ref), for $CS(1-\alpha)$ constructed by equation (ref), \[ \Pr\left(\bm{\theta}\in CS(1-\alpha) \right)\to 1-\alpha. \]

Since the confidence sets by inverting Chatterjee's $\xi$ and by the estimation-based approach utilize information differently, it is possible that when using alone, neither is a subset of the other. Hence, the intersection approach may improve the power by yielding a smaller confidence set. In our Monte Carlo simulations and empirical applications, we do observe improvement in this intersection approach compared to using one method alone.

remarkIn principle, the estimation-based confidence set can be either an ellipsoid (e.g. Wald-type) or a box. Neither dominates the other per se since they both can have exact $\sqrt{1-\alpha}$ asymptotic coverage probability. A box estimation-based confidence set, as we propose in this section, has the following advantages in our specific problem. First, it provides a compact parameter space for each $\theta_{k}$ so that a one-dimensional grid search is feasible to get $CS^{inter}_{k;\alpha}$ defined in equation (ref). Second, for the welfare loss optimization problem, although the maximum can still be easily solved for under an ellipsoid confidence set as a convex constraint, the minimum is difficult to obtain. In contrast, when the confidence set is a box, one can solve the constrained minimization problem simply by plugging in the lower bounds of the individual confidence intervals utilizing coordinate-wise strict monotonicity of the welfare loss function.

Monte Carlo Simulations

Comparison of Chatterjee's $\xi$ and Other Tests of Independence

In this section, in a model with three goods, we examine the performance of our confidence set by i) checking the rejection frequencies of the joint test, ii) examining the individual confidence intervals' length for each parameter and computation time and iii) computing the welfare loss bounds under various constraints with coverage frequencies. In each task, we compare the performance of our method with the other tests of independence mentioned in the introduction.

We generate the simulated samples as follows. Let $\Sigma=

pmatrix[pmatrix omitted — 44 chars of source]

$. Independently draw $\bm{P}_{raw}$ and $\bm{W}_{raw}$ from $\mathcal{N}(0,\Sigma)$. Then for $k=1,2,3$, construct $P_{k}$ by $\Phi(P_{raw,k})+1$ and $W_{k}$ by $\Phi(W_{raw,k})$, where $\Phi$ is the cumulative distribution function of the standard normal distribution. This ensures $P_{k}>W_{k}>0$ with probability 1 and thus the optimal demand is an interior solution. We then generate $Y_{k}$ by the first order condition $Y_{k}=\theta_{k}/(P_{k}-W_{k})$, where $(\theta_{1},\theta_{2},\theta_{3})=(0.2,0.3,0.5)$.

Rejection Frequencies

We first compare the rejection frequencies for Chatterjee's $\xi$, Hoeffding's $D$ hoeffding1948non, Blum-Kiefer-Rosenblatt's $R$ blum1961distribution and Bergsma-Dassios-Yanagimoto's $\tau^{*}$ bergsma2014consistent,yanagimoto1970measures. We set $\alpha=0.1$. We consider two null hypotheses, one equal to the true value $(0.2,0.3,0.5)$, whereas the other equal to $(0.2+0.1/3, 0.3+0.1/3, 0.5-0.2/3)$.

For Chatterjee's $\xi$, we reject if there is at least one $k$ such that $\sqrt{n}\xi_{n}(P_{k}-\theta_{k}^{0}/Y_{k},P_{k})/\sqrt{0.4}> z_{(1-\alpha)^{1/3}}$ where $\bm{\theta}^{0}\coloneqq(\theta_{k}^{0})$ is the null hypothesis and the critical value is implied by Theorem (ref). The value of $\xi_{n}$ under each $\bm{\theta}^{0}$ is computed using the R package XICOR chatterjee2020xicor.

For the other tests, since their joint limiting distributions are unavailable to the best of our knowledge, we do Bonferroni correction by using the critical values at $\alpha/3$ for each $\theta_{k}$. Computationwise, we use the R package independence even2020independence. The algorithms the package adopts achieve $O(n\log n)$ computation time for all these tests.

Table (ref) reports the results. The rejection frequencies are computed based on 500 simulation replications. When the null is equal to the true parameter value, the rejection frequencies of Chatterjee's $\xi$ is close to the nominal level $0.1$ for all three sample size. The other methods are less stable. When the null is not equal to the true value, the rejection frequency of Chatterjee's $\xi$ is lower than the others when $n=200$; this is expected as Chatterjee's $\xi$ has lower power than the other three as documented in shi2022power. However, this difference becomes much smaller at $n=1,000$. At $n=5,000$, all four tests reject the false hypothesis in all 500 replications.

table[table omitted — 839 chars of source]

Confidence Intervals

For Chatterjee's $\xi$, we construct a $(1-\alpha)$ joint confidence set for $\bm{\theta}$ by equation (ref), treating $P_{k}$ as $Z_{k}$; specifically, for each $k$, we search over a grid of 1,000 nodes in $(0,1)$, denoted by $\Theta$, to find points that are in the set defined in (ref), and then take the minimum and the maximum. For the other tests, we invert the individual tests with the $\alpha/3$ critical value by a similar grid search.

Table (ref) shows the results averaged over 500 repetitions. The lower and upper bounds (LB and UB) of our confidence intervals shrink toward the true parameters as the sample size increases. Compared to the confidence intervals obtained by inverting the other tests, our intervals have longer lengths, but the differences in length are small especially when the sample size is large.

In terms of computation time, on the other hand, our method is about 1,000 times faster than the other three. This suggests that when the sample size (and the number of goods $K$, since computation time increases linearly in $K$ for all these methods) is large, our method will be much more efficient in computation while maintaining reasonably large power.

table[table omitted — 1,878 chars of source]

Welfare Loss Bounds

For the welfare loss, we consider a hypothetical price change $\bm{\Delta}=(0.5,0.8,0.2)$ with consumption fixed at $(0.2,0.6,0.8)$. We maximize the standardized welfare loss function (ref) under one of the four constraints:

itemize• Confidence intervals and $\sum_{k}\theta_{k}=1$; • Confidence intervals only; • $\theta_{k}\in [10^{-6},1]$ and $\sum_{k}\theta_{k}=1$; • $\theta_{k}\in [10^{-6},1]$ only.

For constraints ii) and iv), we maximize equation (ref) simply by evaluating it at the upper bounds of the $\theta_{k}$s. For constraints i) and iii), we use the CVXR package for R CVXR to conduct optimization. See Appendices (ref) for details. Note that the maximum loss under iii) and iv) are constant across different sample sizes. The true value of the standardized welfare loss is 0.525.

table[table omitted — 1,068 chars of source]

Table (ref) shows the upper bounds on the standardized welfare loss, as well as the coverage frequencies of these bounds, averaged over 500 simulation repetitions\footnote{For $D$, $R$, and $\tau^{*}$, there are up to three simulation repetitions in which constraints i) yields an empty set because there are no values in the confidence intervals satisfying $\sum_{k}\theta_{k}=1$. The upper bounds and the coverage frequencies for these three methods are thus averaged over the rest of the repetitions.}. We have the following four findings. i) All the upper bounds shrink towards the true welfare loss as the sample size increases. ii) Information from data does improve the bounds compared to bounds without using data, as the bounds under constraint i) (or ii)) are smaller than those under constraint iii) (or iv)). iii) The standardized welfare loss bounds under $\xi$ are larger than those under the other three statistics, but the differences are small and shrink toward zero as the sample size increases. iv) The coverage frequencies of these upper bounds are always close or equal to one, greater than $(1-\alpha)$. This is expected because Theorem (ref) shows that the coverage probability for the standardized welfare loss bounds by construction can be greater than the coverage probability for the parameters.

Scalability When There Are Many Goods

In this section, we examine the performance and computation time of our confidence set and welfare loss bounds when there are many goods. We set the number of goods $K$ equal to $10,50$ or $100$. Parameters $\theta_{k}$s are $K$ equally spaced numbers between $0.1$ and $0.9$. For each good $k$, $P_{k}$ and $W_{k}$ are independently drawn from $\text{Unif}[1,2]$ and $\text{Unif}[0,1]$, respectively. Consumption $Y_{k}$ is constructed by the first-order condition.

We start by examining the rejection frequency of the joint test. We consider two null hypotheses. The first hypothesis $\bm{\theta}^{0}_{1}$ is equal to the true value. For the second hypothesis $\bm{\theta}^{0}_{2}$, $\theta_{2,k}^{0}=\theta_{k}+0.03$ if $k\leq K/2$ and $\theta_{2,k}^{0}=\theta_{k}-0.03$ if $k> K/2$. Computation is done similarly as Section (ref) with critical value equal to $z_{(1-\alpha)^{1/K}}$, where $\alpha=0.1$.

table[table omitted — 659 chars of source]

Table (ref) shows the rejection frequencies in 500 simulation repetitions. At the true value, the rejection frequencies are close to the nominal level $0.1$. Under the false hypothesis, the rejection frequencies increase with the sample size and are close to or equal to 1 when $n$ reaches 1,000.

Next, we compute the $(1-\alpha)$ confidence sets by (ref), treating $Z_{k}=P_{k}$ and using grid search as before. For the bounds on the standardized welfare loss, we only consider constraint ii) in Section (ref). We draw the hypothetical price change $\bm{\Delta}$ and consumption level $\bm{y}^{0}$ from $\text{Unif}[0.1,1.1]$ and $\text{Unif}[0.2,1.2]$ once, respectively; for each $K$, they stay the same across all simulation replications and all sample sizes. We compute the upper and lower bounds on the standardized welfare loss by substituting the upper and lower bounds of the confidence intervals of the $\theta_{k}$s, respectively.

Table (ref) presents the results averaged over 500 simulation replications. Numbers in columns “Length” are the average length of the $K$ confidence intervals. The confidence intervals shrink as the sample size increases for all $K$. Columns “Welfare” demonstrate the upper and lower bounds on the standardized welfare loss, the true values of which are shown as the numbers in parentheses. The bounds are tight even when the number of goods reaches 100 while the sample size is only 200. Columns “Coverage” are frequencies that these bounds cover the true standardized welfare loss, which are always 1 across all setups.

sidewaystable[htbp]{2pt} \caption{Confidence Sets, Computation Time and Standardized Welfare Loss with Many Goods; $\alpha=0.1$} \begin{tabular}{cccccccccccccccc} \toprule \toprule & \multicolumn{5}{c}{$K=10$} & \multicolumn{5}{c}{$K=50$} & \multicolumn{5}{c}{$K=100$} \\ \midrule & \multicolumn{2}{c}{CI} & \multicolumn{3}{c}{Welfare (1.34)} & \multicolumn{2}{c}{CI} & \multicolumn{3}{c}{Welfare (2.81)} & \multicolumn{2}{c}{CI} & \multicolumn{3}{c}{Welfare (4.15)} \\ \midrule $n$ & Length & Time (s) & Lower& Upper&Coverage& Length & Time (s) & Lower& Upper&Coverage & Length & Time (s) & Lower& Upper&Coverage \\ \midrule 200 & 0.43 &1.01 & 1.19 & 1.54 & 1& 0.50 & 5.05 & 2.48 & 3.21 &1& 0.52 & 10.35 & 3.62 &4.86&1\\ 1000 & 0.21 &2.77 & 1.27 & 1.43 &1& 0.25 & 13.81& 2.64 & 3.01 &1& 0.26 & 27.62& 3.88 & 4.48 &1\\ 5000 & 0.10 & 13.51& 1.31 & 1.38 &1& 0.11 & 59.17& 2.73 & 2.89 &1& 0.12 & 134.31& 4.03 &4.29&1 \\ \bottomrule \end{tabular}

Table (ref) also demonstrates the computation time (in seconds) for confidence intervals\footnote{The average computation time for the standardized welfare loss bound is always below $10^{-4}$ seconds and thus suppressed. Note that the objective function (ref) does not depend on data, so sample size does not affect computation.} under various combinations of $n$ and $K$. By simple calculation, one can see that the computation time is approximately linear in $K$ for each $n$. This is expected because we calculate the confidence set for each good independently. As an implication, parallel computation is feasible to keep the computation time constant across $K$, and thus the problem is scalable.

The Intersection Approach and Rejection Probability

In this simulation exercise, we examine the performance of the intersection approach. Different from the previous sections, the data-generating process is now designed to match some key moments in the food consumption data set in Section (ref). Specifically, we utilize the sample extract from the Stanford Basket Dataset, a household-level scanner panel data set, created by pump for the empirical results presented in Table 4 of their paper. There are two goods in the sample: Ice cream and other foods.

We first construct a data set with valid observations. For each of the 494 households in the sample, there are 26 observations: One observation corresponds to a period of 4 weeks over a two-year span. In Section (ref), we will perform household-level welfare analysis using the intersection approach for households with at least 19 periods of strictly positive consumption of ice cream and other foods. When $\alpha=0.1$, the intersection approach yields nonempty confidence intervals for 38 households. In this section, for simplicity, we pool all these 38 households together and obtain a data set of 763 observations.

We then design the following data-generating process to match the sample mean and variance of the observed prices and consumption in the pooled data set. For each of the two consumption goods, we assume its price follows a truncated normal distribution. The lower and upper bounds are set to equal the observed minimum and maximum of the price data. For the pre-truncation mean and variance, we calibrate them by grid search so that the resulting post-truncation mean and variance match those of the observed price data.

For the unobservables and consumption, we first draw $(W_{raw,1},W_{raw,2})$ from $\mathcal{N}(0,\Sigma_{W})$ where $\Sigma_{W}=

pmatrix[pmatrix omitted — 41 chars of source]

$ and $\rho_{W}\in\{0.1,0.7\}$. Then to guarantee $0<W_{k}<P_{k}$ with probability 1, we construct $W_{1}=0.553\Phi(W_{raw,1})^{2}$ and $W_{2}=1.1596\Phi(W_{raw,2})^{2}$. Finally, we construct $Y_{k}=\theta_{k}/(P_{k}-W_{k})$ with $(\theta_{1},\theta_{2})=(1.1,0.6)$.

Table (ref) presents the sample mean and variance of the observed and simulated prices and consumption averaged across 1,000 simulation replications; in each simulation replication, the sample size is 5,000. We can see that our data-generating process matches these moments well.

table[table omitted — 554 chars of source]

We now turn to our main simulation results. We first examine the joint rejection frequencies of three different approaches: Chatterjee's $\xi$, SUR and the delta method, and the intersection approach. For Chatterjee' $\xi$, since there are two goods, we calculate the rejection frequencies similar to Sections (ref) with critical value $ z_{\sqrt{1-\alpha}}$. For SUR, we calculate the rejection frequencies based on the sup-t test. The asymptotic rejection probability is also exactly $\alpha$ under the null. For the intersection approach, we reject if and only if at least one of the tests based on Chatterjee and SUR suggests rejection; the size of each test is adjusted to $1-\sqrt{1-\alpha}$ so that the joint rejection probability under the null is controlled at $\alpha$ based on Theorem (ref).

table[table omitted — 941 chars of source]

Table (ref) presents the results based on 500 simulation replications. We consider two hypotheses, the true value $(1.1,0.6)$ and the false value $(1.3,0.8)$. The results show that all three methods, under the true parameter value, have rejection frequencies close to $\alpha=0.1$. This verifies the asymptotic independence results in our Theorems (ref) and (ref). Under the false hypothesis, these methods all have large power, and the power increases as the sample size increases. In particular, SUR and the intersection approach improve when the unobservables are more correlated.

Finally, we compute the standardized welfare loss bounds based on the confidence intervals obtained by the three methods. Again, we treat $Z_{k}=P_{k}$ for all $k$. When using Chatterjee's $\xi$ alone, the confidence sets are constructed following (ref) and (ref) by grid search in $\Theta_{k}=[10^{-6},2]$ with 2,000 grid points. For the intersection approach, grid search with the same number of grid nodes is conducted in the sup-t confidence interval following (ref) and (ref). The consumption level is fixed at the sample median of the observed consumption. The price increase is 20% of the sample median of the observed prices.

Table (ref) presents the standardized welfare loss bounds and the length between the lower and upper bounds averaged over 500 simulation replications. We can see that the intersection approach always achieves the shortest length. Due to the efficiency gain from SUR, the improvement compared to using $\xi$ alone is larger when the unobservables are more correlated.

The table also shows the frequencies of the true standardized welfare loss covered by the confidence bounds. It is always close or equal to one, greater than $(1-\alpha)$, similar to the previous sections.

table[table omitted — 1,151 chars of source]

Empirical Applications

Demand for Gasoline

We study the standardized welfare loss of a hypothetical gasoline price increase. We take the data set constructed by blundell2017nonparametric using the household-level 2001 National Household Travel Survey (NHTS). The sample contains 3,640 observations with annual gasoline demand ($\tilde{Y}$), price of gasoline ($P$) and the distance between one of the major oil platforms in the Gulf of Mexico and the state capital\footnote{In the data set, the distance takes on 34 values ranging from 0.361 to 3.391. We treat it as a continuous variable. In Appendix (ref), we further smooth it by adding a small noise term to it. The results are almost the same.} ($Z$; see blundell2017nonparametric for a more detailed discussion). We consider the daily demand $Y$ by dividing $\tilde{Y}$ by 365.

Under our utility function (ref) and the first order condition (ref) with $K=1$, we can construct confidence intervals for $\theta$ using our method, 2SLS and the delta method, and the intersection approach. When using Chatterjee's $\xi$ alone, our confidence interval follows (ref) by grid search with 5,000 grid nodes over $\Theta=[10^{-6},6]$. For the intersection approach (ref), grid search is done over the $\sqrt{1-\alpha}$ confidence interval obtained by 2SLS and the delta method. Table (ref) presents the results for $\alpha=0.05$. For the standardized welfare loss bounds, we let the price of gasoline increase by twice of the standard deviation (0.076). The current demand is set at the $u$-th sample quantile ($u\in (0,1)$) of $Y$, $q_{Y}(u)$, for $u=0.9,0.5$ and $0.1$.

table[table omitted — 677 chars of source]

From Table (ref), the intersection approach improves the lower bound of the confidence interval a lot compared to the lower bound under 2SLS alone, whereas the upper bound only increases a bit. See Appendix (ref) for more discussion. Similarly, the standardized welfare loss bounds are tightest under the intersection approach (entries in boldface). In particular, the difference between the upper and lower bounds shrinks by about 40%, 44%, and 47% when intersecting the two confidence sets compared to using 2SLS alone for the three demand levels, respectively. {Table (ref) shows that a price increase leads to more severe damage to the welfare of individuals who consume more gas, which is reasonable.}

A Varying Coefficient Model

One limitation of the quasilinear utility function is the missing income effect. To address this issue, we now allow $\theta$ to depend on income, $I$. Let $med(I)$ be the median of the income distribution. Random variable $\theta^{I}=\theta^{H}$ if $I>med(I)$ and $\theta^{I}=\theta^{L}$ if $I\leq med(I)$. Assume that the instrument is independent of the unobservable conditional on income. In this section, we compute the welfare loss bounds in the two subsamples defined by whether the individual's income is greater than the sample median. Similar to the case of a homogeneous $\theta$, the intersection approach largely improves the lower bound of the welfare loss in both subsamples compared to using 2SLS alone. To save space, we only report the results obtained by this approach. Again, $\alpha=0.05$.

table[table omitted — 596 chars of source]

Table (ref) suggests heterogeneity in $\theta$ and consequently in the standardized welfare loss. From the results, individuals with higher incomes tend to have a larger $\theta$. This is reasonable because higher-income individuals may rely more on driving than public transportation, so the marginal utility of consuming one more unit of gas, captured by $\theta$, is higher. Since they value gas more than those with lower income, the welfare of higher income individuals suffers more when facing an increase in gas price, coherent with the estimated standardized welfare loss bounds.

Food Demand

In this application, we consider food demand using the Stanford Basket Dataset. As mentioned in Section (ref), it is a household-level scanner panel data set, and we use the sample extract created by pump. Two goods, ice cream and other foods, are in the sample.

We conduct household-level welfare analysis taking advantage of the panel data structure. Recall that there are 26 observations for each of the 494 households in the sample. We focus on 41 households that have at minimum 19 periods of strictly positive consumption of ice cream and other foods; each household can have a different $\bm{\theta}$. We use the last observation for the counterfactual welfare analysis and the remaining observations for the construction of confidence sets. The i.i.d. assumption required in our theory may be strong here; we regard it as a convenient approximation.

figure[figure omitted — 193 chars of source]

We assume the prices are exogenous. To construct a $(1-\alpha)$ confidence set where $\alpha=0.05$, we first construct a $\sqrt{1-\alpha}$ sup-t confidence band for each household by SUR and the delta method. We then do a grid search with 5,000 grid nodes within the resulting confidence intervals as equations (ref) and (ref). Two out of 41 households have an empty confidence set.

For the welfare analysis for the 39 households with nonempty confidence sets, we compute the standardized welfare loss bounds in the same way as before. Figure (ref) shows the lower and upper bounds on the standardized welfare loss for the 39 households when both prices increase by 10%. The maximal standardized welfare loss of the households is sorted in ascending order. Household heterogeneity is visible and the lower and upper bounds are relatively tight for most households.\footnote{In Appendix (ref), we test one assumption in our model (ref) that the demand for a good is only affected by its own price. This hypothesis is rejected for only 3 of the 39 households. The figure excluding these 3 households is presented in that appendix; it has a very similar shape as Figure (ref). We thank an anonymous referee for suggesting this test.}

Conclusion

In this paper, we propose a novel framework for individual-level welfare analysis. At any desired confidence level, our method can compute the bounds on the standardized welfare loss under a price increase for every individual in a sample. The inferential method is computationally simple and scalable by solving a simple scalable optimization problem constrained by the confidence set for the parameters in the model.

We also propose a new method to construct confidence sets under independence based on the new asymptotic results of Chatterjee's test of independence developed in this paper; the new method is easy to compute and robust to nonlinearity, weak instruments, and partial identification; these results may be of independent interest. To sharpen the confidence set, we propose to intersect our confidence set with an alternative one when available. As in our second empirical example, the intersection approach sometimes yields an empty set. In that case, we may try a very small $\alpha$, for instance, 0.5%. If the intersection becomes nonempty at the new $\alpha$, we can report the new confidence interval as auxiliary information. In our empirical example, all the households have nonempty confidence sets under the interaction approach when we set $\alpha=1\%$. If, on the other hand, the intersection is still empty under the new $\alpha$, then it is a warning that the model may be misspecified. Misspecification is beyond the scope of this paper, and we leave it to future work.

Another direction of future research would be to apply our framework to demand models such as dubois2014prices and allcott2019food to conduct individual-level welfare analysis.