Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
19,428 characters · 4 sections · 4 citation commands
Nonparametric Counterfactuals in Random Utility Models
\address{Cowles Foundation for Research in Economics, Yale University, New Haven, CT 06520.} \email{[email removed]} \address{Department of Economics, Cornell University, Ithaca, NY 14853.} \email{[email removed]}
Consider the random utility model of demand analyzed in \citeasnoun{mcfadden-richter}, \citeasnoun{McFadden05}, and \citeasnoun{KS}: Repeated cross-sections of demand are observed on a finite sequence of budgets; the maintained assumption is that these cross-sections are of a population of individually rational (in the sense of maximizing utility) individuals; however, one does not substantively restrict utility functions nor their distribution, that is, one allows for unrestricted and possibly infinite dimensional unobserved heterogeneity.
\citeasnoun{McFadden05} characterizes the empirical content of this model in population. We build on his results to provide tight bounds on the distribution of counterfactual demand, i.e. of stochastic demand on as yet unobserved budgets, as well as explicit bounds on the expectation and c.d.f. of linear functions of demand vectors, e.g. demand for a specific good. Many of these bounds turn out to be the values of linear programs, hence are easy to compute even in moderately high dimensional applications.\footnote{Some of these results were reported in section 9.2 of \citeasnoun{KS13} and implemented at the time. We make them available not least because other work already built on them Adams16,Manski14. Code is available from the authors. See also \citeasnoun{Adams16}, \citeasnoun{Hubner}, and \citeasnoun{smeulders} for recent results on computational implementation.} We next describe the setup and recall an important characterization of stochastic rationalizability, then provide the bounds, and close by mentioning some extensions.
We use notation from \citeasnoun{KS}. There are $J$ observed budgets $\{\mathcal{B}_j\}_{j=1}^J, J \in \bf N$, each characterized by price vectors $p_j \in \mathbf{R}^K_+$, where expenditure is normalized to $1$:
Suppose that we know a stochastic demand system
for $j=1,\dots,J$, where the random variable $y(p_j)$ is demand on budget $\mathcal{B}_j$. This collection of distributions is rationalizable by a random utility model if there exists a distribution $P_u$ over locally nonsatiated (for simplicity) utility functions $u:\mathcal{\mathbf{R}}_+^K \mapsto \mathbf{R}$ s.t.
Our motivation is demand estimation from repeated cross-section with unobserved heterogeneity, but the model has also been used to describe choices made by an individual with random utility. We next recall a succinct description of its empirical content.
Let $\mathcal{X} \equiv \{x_1,...,x_I\}$ be the coarsest partition of $\cup _{j=1}^J\mathcal{B}_j$ such that for any $i \in \{1,...,I\}$ and $j \in \{1,...,J\}$, $x_i$ is either completely on, completely strictly above, or completely strictly below budget plane $\mathcal{B}_j$. Equivalently, any $y_1,y_2 \in \cup _{j=1}^J\mathcal{B}_j$ are in the same element of the partition iff $\text{sg}(p_j'y_1-1)=\text{sg}(p_j'y_2-1)$ for all $j=1,...,J$. Elements of $\mathcal{X}$ will be called patches. Each budget can be uniquely expressed as union of patches; the number of patches that jointly comprise budget $\mathcal{B}_j$ will be called $I_j$. For future reference, we emphasize that any patch is the intersection of finitely many open or closed half spaces and therefore its closure (though not necessarily the patch itself) is a finite polytope.
An important insight of the aforecited papers is that stochastic rationalizability constrains the aggregate choice probabilities of patches, but not at all the distribution of demand on any patch. Intuitively, this is because all choices that are on the same patch generate the same revealed preference information. Formally, let the vector representation of $(\mathcal{B}_1,\dots,\mathcal{B}_J)$ be the $\bigl(\sum_{j=1}^J I_j\bigr)$-vector $(x_{1|1},\dots,x_{I_1|1},x_{1|2},\dots,x_{I_J|J})$, where $(x_{1|j},\dots,x_{I_j|j})$ lists all patches comprising $\mathcal{B}_j$ in arbitrary but henceforth fixed order.\footnote{Note that elements of $\mathcal{X}$ that appear as components of distinct budgets make corresponding repeat appearances, under different labels, in the vector representation.} Let the vector representation of $(P_1,\dots,P_J)$ be the $\bigl(\sum_{j=1}^{J}I_{j}\bigr)$-vector $\pi \equiv (\pi_{1|1},\dots,\pi_{I_1|1},\pi_{1|2},\dots,\pi_{I_J|J})$, where $\pi_{i|j} \equiv P_j(x_{i|j})$. Next, note that any rationalizable nonstochastic demand system can be thought of as a degenerate stochastic demand system with binary vector representation. A stochastic demand system is rationalizable iff it is a mixture of such rationalizable nonstochastic demand systems because the latter can be thought of as representing choice types in the population. But due to the discretization of the choice universe into patches, there are only finitely many such types. Collect their vector representations in the $H < \infty$ columns of the rational demand matrix $A$: see \citeasnoun{KS} (in particular Definition 3.5 and discussions in Sections 3.2 - 3.4). Then we have:\footnote{The statement follows \citeasnoun{KS}, who also prove it, provide algorithms for computing $A$, and point out that $\nu \in \Delta^{H-1}$ can be conveniently weakened to $\nu \geq 0$. However, the discretization step is clearly anticipated in \citeasnoun{McFadden05}, and the result was otherwise proved in \citeasnoun{mcfadden-richter}. See also \citeasnoun{Stoye18}.}
We next take a rationalizable stochastic demand system $(P_1,\dots,P_J)$ as given and ask what discipline it places on $$y(p_0) := \operatornamewithlimits{argmax}_{y\in \mathbf{R}_+^K:p_0'y=1} u(y), \quad u \sim P_u, $$ the stochastic demand at some counterfactual budget $\mathcal{B}_0$ corresponding to counterfactual price $p_0$.\footnote{For the very special case of $K=2$, \citeasnoun{HS15} provide closed-form bounds. \citeasnoun{BBC08} and many others provide bounds under slightly stronger, e.g. aggregation, assumptions.} As with nonstochastic demand, this discipline will typically take the form of bounds, although these are now on a distribution. They are tightly related to testing rationalizability because a distribution $P_0$ of demands on $\mathcal{B}_0$ is inside the bounds iff $(P_0,...,P_J)$ are jointly rationalizable; thus, Theorem (ref) implies an exact characterization of bounds on $P_0$ implied by knowledge of $(P_1,...,P_J)$. We will now formally state this characterization.
Recall that the matrix $A$ in Theorem (ref) is obtained for the set of observed budgets $(\mathcal B_1,...,\mathcal B_J)$. We can apply the same algorithm to the augmented set of budgets $(\mathcal B_0, \mathcal B_1,...,\mathcal B_J)$ to obtain patches on it and its vector representations: for completeness we write them
The patches for the original system $(\mathcal B_1,...,\mathcal B_J)$ remain unchanged in the augmented system if they do not intersect with $\mathcal B_0$. Therefore if $\mathcal B_{j'} \cap \mathcal B_0=\emptyset$ holds for some $j' \in \{1,\dots,J\}$, then $I_{j'}^* = I_j$ and $$ (x^*_{1|j'},\dots,x^*_{I^*_{j'}|j'}) = (x_{1|j'},\dots,x_{I_{j'}|j'}) $$ for such $j'$. Moreover we can apply the algorithm discussed in Section (ref) to the augmented system $(\mathcal B_0,\dots,\mathcal B_J)$ to obtain its rational demand matrix $A^* \in {\bf R}^{\left(\sum_{j=1}^{J+1}I_j^*\right) \times {H^*}}$, where $H^* \geq H$. Note that each row of $A^*$ corresponds to a patch in the new vector representation (ref). Once $A^*$ is obtained, we can define a probability vector $\nu^* \in \Delta^{H^* -1}$, now defined over the columns of $A^*$, and the choice probability vector $\pi^*$ for the patches $\{ x^*_{1|0},...,x^*_{I^*_0|0},x^*_{1|1},...,x^*_{I^*_J|J} \}$. Note that the elements of $\pi^*$ corresponding to $\cup_{j=1}^J \mathcal B_j$ are observed, while the rest remain unobserved: the latter are counterfactual conditional probabilities. To make this point clear, we write
where $A_0^*$ collects rows of $A^*$ that correspond to patches that do not belong to $\cup_{j=1}^J \mathcal B_j$, $A_1^*$ collects all other patches, and similarly for $\pi^*$. It continues to be the case that $\pi^*$ is rationalizable iff $A^*\nu^*=\pi^*$ for some $\nu^* \in \Delta^{H^*-1}$. However, rather than taking $\pi^*$ to be observed and testing rationalizability, we take $\pi_1^*$ to be observed and $\pi_0^*$ to a vector of counterfactual probabilities to be accordingly constrained by the observed $\pi_1^*$. Formally:\footnote{For a setting like ours except that the universal choice set is finite and “budgets" are subsets of it, \citeasnoun{Manski07} anticipates Theorem (ref). Our contribution lies in the connection to nonparametric demand, in laying the groundwork for Theorem (ref) and its corollaries, and in the accompanying computational as well as statistical machinery.}
We next explain how this result translates into extremely tractable, best possible bounds on many parameters of interest. Specifically, we have:
The proof reveals not only that the bounds are sharp, but also that all intermediate values of $\mathbb{E}g(y(p_0))$ are necessarily attainable. Whether the bounds themselves are attainable depends on whether patches are open or closed in the relevant directions and can only be decided on a case-by-case basis.
Computing these bounds requires to solve the linear programs in (ref) and, as an input, the optimization problems in (ref)-(ref). The latter are tractable in relevant cases: If $g$ is continuous, the constraint sets can be taken to be the closures of patches, hence finite polytopes. If $g$ is furthermore linear, then computing the bounds requires only linear programming, though possibly with many constraints.
We further elaborate this result by more explicitly bounding the expected value and c.d.f. of $z'y(p_0)$, where $z \in {\bf R} ^K$ is a user-specified vector. For example, $z=(1,0,\dots,0)'$ extracts demand for good $1$ and $z=(p_0^{[1]},p_0^{[2]},0,\dots,0)'$ (i.e., the first two components of $p_0$ followed by zeroes) extracts joint expenditure on the first two goods. Theorem (ref) then specializes as follows.
We note that computation of these bounds only requires linear programming. Next, we bound probabilities of arbitrary events and hence also c.d.f.'s.
The bounds on the c.d.f. require essentially only linear programming, with a minimal additional check in the finitely many cases where $\underline{m}_{i|0}(z)=t$.\footnote{This case occurs when $\{z'y=t\}$ is a lower supporting hyperplane of patch $x_{i|0}$ but does not intersect it.} They are pointwise but not uniform in $t$; in particular, their upper and lower envelopes do not necessarily describe feasible counterfactual distributions.\footnote{Indeed, they may not even be c.d.f.'s for lack of right-continuity, though they can always be approximated by c.d.f.'s.} Therefore, while the upper and lower envelopes induce bounds on a multitude of more complicated parameters Stoye10, those bounds are not in general tight.
We conclude by mentioning some connections and extensions.
First, we considered the case of one counterfactual budget for expositional clarity. Bounds on the joint c.d.f. of demand on two counterfactual budgets, or on some linear combination of expected values, are straightforward extensions of the above results. In general, they can be considerably tighter than the Cartesian product of budget-by-budget bounds and also need not include the budget-by-budget minimum and maximum bound. For the case of a single (possibly fictitious, e.g. representative) nonstochastic utility maximizer, see \citeasnoun{Adams16} for a much more extensive analysis in this spirit.
Next, these results naturally extend to finite discrete choice settings, i.e. if choices from distinct subsets $\mathcal{C}_1,\dots,\mathcal{C}_J$ of a finite choice universe $\mathcal{X}$ (the duplication of notation is intended) were observed and choices from another such subset $\mathcal{C}_0$ are to be predicted. In this case, the finitely many elements of $\mathcal{X}$ directly play the role of patches, the conditional distributions on patches are trivial, and Theorem 2 characterizes those p.m.f.'s of counterfactual random choice that are consistent with observed choice distributions. That said, the analysis in \citeasnoun{Manski07} anticipates Theorem 2 in this setting. Of course, the result also applies to the further specialization where all observed choice sets are binary, as in the “linear polytope" literature in mathematical psychology fishburn92.
Finally, we developed population-level bounds but ignored estimation and inference. To handle this, note that if one takes $(p_0,\dots,p_J)$ and therefore $A^*$ to be known, then all the above bounds maximize or minimize $\gamma A_0^* \nu^*$ for some known vector $\gamma$; it is only $\pi_1^*$ that must be estimated. This is essentially the estimation and inference problem analyzed in Section 4.2 of \citeasnoun{DKQS16}.