Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,634 characters · 20 sections · 14 citation commands
Nonparametric Analysis of Random Utility Models
\address{Cowles Foundation for Research in Economics, Yale University, New Haven, CT 06520.} \email{[email removed]} \address{Departments of Economics, Cornell University, Ithaca, NY 14853, and University of Bonn, Germany; Hausdorff Center for Mathematics, Bonn, Germany.} \email{[email removed]}
This paper develops new tools for the nonparametric analysis of Random Utility Models (RUM). We test the null hypothesis that a repeated cross-section of demand data might have been generated by a population of rational consumers, without restricting either unobserved heterogeneity or the number of goods. Equivalently, we empirically test McFadden and Richter's (1991) Axiom of Revealed Stochastic Preference. To do so, we develop a new statistical test that promises to be useful well beyond the motivating application.
We start from first principles and end with an empirical application. Core contributions made along the way are as follows.
First, the testing problem appears formidable: A structural parameterization of the null hypothesis would involve an essentially unrestricted distribution over all nonsatiated utility functions. However, the problem can, without loss of information, be rewritten as one in which the universal choice set is finite. Intuitively, this is because a RUM only restricts the population proportions with which preferences between different budgets are directly revealed. The corresponding sample information can be preserved in an appropriate discretization of consumption space.
More specifically, observable choice proportions must be in the convex hull of a finite (but long) list of vectors. Intuitively, these vectors characterize rationalizable nonstochastic choice types, and observable choice proportions are a mixture over them that corresponds to the population distribution of types. This builds on \citeasnoun{mcfadden-2005} but with an innovation that is crucial for testing: While the set just described is a finite polytope, the null hypothesis can be written as a cone. Furthermore, computing the list of vectors is hard, but we provide algorithms to do so efficiently.
Next, the statistical problem is to test whether an estimated vector of choice proportions is inside a nonstochastic, finite polyhedral cone. This is reminiscient of multiple linear inequality testing and shares with it the difficulty that inference must take account of many nuisance parameters. However, in our setting, inequalities are characterized only implicitly through the vertices of their intersection cone. It is not computationally possible to make this characterization explicit. We provide a novel test and prove that it controls size uniformly over a reasonable class of d.g.p.'s without either computing facets of the cone or resorting to globally conservative approximation. This is a contribution of independent interest that has already seen other applications DKQS16,Hubner,LQS15,LQS18. Also, while our approach can become computationally costly in high dimensions, it avoids a statistical curse of dimensionality (i.e., rates of approximation do not deteriorate), and our empirical exercise shows that it is practically applicable to at least five-dimensional commodity spaces.
Finally, we leverage recent results on control functions (\citeasnoun{imbens2009}; see also \citeasnoun{blundell2003}) to deal with endogeneity for unobserved heterogeneity of unrestricted dimension. These contributions are illustrated on the U.K. Family Expenditure Survey, one of the work horse data sets of the literature. In that data, estimated demand distributions are not stochastically rationalizable, but the rejection is not statistically significant.
The remainder of this paper is organized as follows. Section (ref) discusses the related literature. Section (ref) lays out the model, develops a geometric characterization of its empirical content, and presents algorithms that allow one to compute this characterization in practice. All of this happens at population level, i.e. all identifiable quantities are known. Section (ref) explains our test and its implementation under the assumption that one has an estimator of demand distributions and an approximation of its sampling distribution. Section (ref) explains how to get the estimator, and a bootstrap approximation to its distribution, by both smoothing over expenditure and adjusting for endogenous expenditure. Section (ref) contains a Monte Carlo investigation of the test's finite sample performance, and Section (ref) contains our empirical application. Section (ref) concludes. Supplemental materials collect all proofs (Appendix A), pseudocode for some algorithms (Appendix B), and some algebraic elaborations (Appendix C).
Our framework for testing Random Utility Models is built from scratch in the sense that it only presupposes classic results on nonstochastic revealed preference, notably the characterization of individual level rationalizability through the Weak Samuelson38, Strong Houthakker50, or Generalized Afriat67 Axiom of Revealed Preference (WARP, SARP, and GARP henceforth). At the population level, stochastic rationalizability was analyzed in classic work by \citeasnoun{mcfadden-richter} updated by \citeasnoun{mcfadden-2005}. This work was an important inspiration for ours, and we will further clarify the relation later, but they did not consider statistical testing nor attempt to make the test operational, and could not have done so with computational constraints even of 2005.
An influential related research project is embodied in a sequence of papers by Blundell, Browning, and Crawford (2003, 2007, 2008; BBC henceforth), where the 2003 paper focuses on testing rationality and bounding welfare and later papers focus on bounding counterfactual demand. BBC assume the same observables as we do and apply their method to the same data, but they analyze a nonstochastic demand system generated by nonparametric estimation of Engel curves. This could be loosely characterized as revealed preference analysis of a representative consumer and in practice of average demand. \citeasnoun{Lewbel01} gives conditions on a RUM that ensure integrability of average demand, so BBC effectively add those assumptions to ours. Also, the nonparametric estimation step in practice constrains the dimension of commodity space, which equals three in their empirical applications.\footnote{BBC's implementation exploits only WARP and therefore a necessary but not sufficient condition for rationalizability. This is remedied in \citeasnoun{BBCCDV}.}
\citeasnoun{Manski07} analyzes stochastic choice from subsets of an abstract, finite choice universe. He states the testing and extrapolation problems in the abstract, solves them explicitly in simple examples, and outlines an approach to non-asymptotic inference. (He also considers models with more structure.) While we start from a continuous problem and build a (uniform) asymptotic theory, the settings become similar after our initial discretization step. However, methods in \citeasnoun{Manski07} will only be practical for choice universes with a handful of elements, an order of magnitude less than in Section (ref) below. In a related paper, \citeasnoun{Manski14} uses our computational toolkit for choice extrapolation.
Our setting much simplifies if there are only two goods, an interesting but obviously very specific case. \citeasnoun{BKM14} bound counterfactual demand in this setting through bounding quantile demands. They justify this through an invertibility assumption. \citeasnoun{HS15} show that with two goods, this assumption has no observational implications.\footnote{A similar point is made, and exploited, by \citeasnoun{HN16}.} Hence, \citeasnoun{BKM14} use the same assumptions as we do; however, the restriction to two goods is fundamental. \citeasnoun{BKM17} conceptually extend this approach to many goods, in which case invertibility is a restriction. A nonparametric estimation step again limits the dimensionality of commodity space. They apply the method to similar data and the same goods as BBC, meaning that their nonparametric estimation problem is two-dimensional.
\citeasnoun{HN16} nonparametrically bound average welfare under assumptions resembling ours, though their approach additionally imposes smoothness restrictions to facilitate nonparametric estimation and interpolation. Their main identification results apply to an arbitrary number of goods, but the approach is based on nonparametric smoothing, hence the curse of dimensionality needs to be addressed. The empirical application is to two goods.
With more than two goods, pairwise testing of a stochastic analog of WARP amounts to testing a necessary but not sufficient condition for stochastic rationalizability. This is explored by \citeasnoun{HS14} in a setting that is otherwise ours and also on the same data. \citeasnoun{Kawaguchi} tests a logically intermediate condition, again on the same data. A different test of necessary conditions was proposed by \citeasnoun{Hoderlein11}, who shows that certain features of rationalizable individual demand, like adding up and standard properties of the Slutsky matrix, are inherited by average demand under weak conditions. The resulting test is passed by the same data that we use. \citeasnoun{DHN16} propose a similar test using quantiles.
Section (ref) of this paper is (implicitly) about testing multiple inequalities, the subject of a large literature in economics and statistics. See, in particular, \citeasnoun{ghm} and \citeasnoun{wolak-1991} and also \citeasnoun{chernoff}, \citeasnoun{kudo63}, \citeasnoun{perlman}, \citeasnoun{shapiro}, and \citeasnoun{takemura-kuriki} as well as \citeasnoun{andrews-hac}, \citeasnoun{BCS15}, and \citeasnoun{guggenberger-hahn-kim}. For the also related setting of inference on parameters defined by moment inequalities, see furthermore \citeasnoun{andrews-soares}, \citeasnoun{bugni-2010}, \citeasnoun{Canay10}, \citeasnoun{cht}, \citeasnoun{imbens-manski}, \citeasnoun{romano-shaikh}, \citeasnoun{rosen-2008}, and \citeasnoun{stoye-2009}. The major difference to these literatures is that moment inequalities, if linear (which most of the papers do not assume), define a polyhedron through its faces, while the restrictions generated by our model correspond to its vertices. One cannot in practice switch between these representations in high dimensions, so that we have to develop a new approach. This problem also occurs in a related problem in psychology, namely testing if binary choice probabilities are in the so-called Linear Order Polytope. Here, the problem of computing explicit moment inequalities is researched but unresolved (e.g., see \citeasnoun{DR15} and references therein), and we believe that our test is of interest for that literature. Finally, \citeasnoun{HS14} only compare two budgets at a time, and \citeasnoun{Kawaguchi} tests necessary conditions that are directly expressed as moment inequalities. Therefore, inference in both papers is much closer to the aforecited literature.
We now show how to verify rationalizability of a known set of cross-sectional demand distributions on $J$ budgets. The main results are a tractable geometric characterization of stochastic rationalizability and algorithms for its practical implementation.
Throughout this paper, we assume the existence of $J < \infty$ fixed budgets $\mathcal{B}_j$ characterized by price vectors $p_j \in \mathbf{R}^K_+$ and expenditure levels $W_j >0$. Normalizing $W_j=1$ for now, we can write these budgets as
We also start by assuming that the corresponding cross-sectional distributions of demand are known. Thus, assume that demand in budget $\mathcal{B}_j$ is described by the random variable $y(p_j)$, then we know
for $j=1,...,J$.\footnote{To keep the presentation simple, here and henceforth we are informal about probability spaces and measurability. See \citeasnoun{mcfadden-2005} for a formally rigorous setup.} We will henceforth call $(P_1,...,P_J)$ a stochastic demand system.
The question is if this system is rationalizable by a RUM. To define the latter, let
denote a utility function over consumption vectors $y \in \mathbf{R}_+^K$. Consider for a moment an individual consumer endowed with some fixed $u$, then her choice from a budget characterized by normalized price vector $p$ would be
with arbitrary tie-breaking if the solution is not unique. For simplicity, we restrict utility functions by monotonicity (\textquotedblleft more is better\textquotedblright) so that choice is on budget planes, but this is not conceptually necessary.
The RUM postulates that $$ u \sim P_u, $$ i.e. $u$ is not constant but is distributed according to a constant (in $j$) probability law $P_u$. In our motivating application, $P_u$ describes the distribution of preferences in a population of consumers, but other interpretations are conceivable. For each $j$, the random variable $y(p_j)$ then is the distribution of $y$ defined in (ref) that is induced by $p_j$ and $P_u$. Formally:
This model is completely parameterized by $P_u$, but it only partially identifies $P_u$ because many distinct $P_u$ will induce the same stochastic demand system. We do not place substantive restrictions on $P_u$, thus we allow for minimally constrained, infinite dimensional unobserved heterogeneity across consumers.
Definition (ref) reflects some simplifications that we will drop later. First, $W_j$ and $p_j$ are nonrandom, which is the framework of \citeasnoun{mcfadden-richter} and others but may not be realistic in applications. In the econometric analysis in Section (ref) as well as in our empirical analysis in Section (ref), we treat $W_j$ as a random variable that may furthermore covary with $u$. Also, we initially assume that $P_u$ is the same across price regimes. Once $W_j$ (hence $p_j$, after income normalization) is a random variable, this is essentially the same as imposing $W_j {\bot\negthickspace\negthickspace\bot} u$, an assumption we maintain in Section (ref) but drop in Section (ref) and in our empirical application. However, for all of these extensions, our strategy will be to effectively reduce them to (ref), so testing this model is at the heart of our contribution.
The model embodied in (ref) is extremely general; again, a parameterization would involve an essentially unrestricted distribution over utility functions. However, we next develop a simple geometric characterization of the model's empirical content and hence of stochastic rationalizability.
To get an intuition, consider the simplest example in which (ref) can be tested:
Consider Figure (ref), whose labels will become clear. (The restriction to $\mathbf{R}^2$ is only for the figure.) It is well known that in this example, individual choice behavior is rationalizable unless choice from each budget is below the other budget, in which case a consumer would revealed prefer each budget to the other one. Does this restrict repeated cross-section choice probabilities? Yes: Supposing for simplicity that there is no probability mass on the intersection of budget planes, it is easy to see (e.g. by applying Fr\'{e}chet-Hoeffding bounds) that the cross-sectional probabilities of the two line segments labeled $(\pi_{1|1},\pi_{1|2})$ must not sum to more than $1$. This condition is also sufficient Matzkin06.
Things rapidly get complicated as budgets are added, but the basic insight scales. The only relevant information for testing (ref) is what fractions of consumers revealed prefer budget $j$ to $k$ for different $(j,k)$. This information is contained in the cross-sectional choice probabilities of the line segments highlighted in Figure (ref) (plus, for noncontinuous demand, the intersection). The picture will be much more involved in interesting applications -- see Figure (ref) for an example -- but the idea remains the same. This insight allows one to replace the universal choice set $\mathbf{R}^K_+$ with a finite set and stochastic demand systems with lists of corresponding choice probabilities. But then there are only finitely many rationalizable nonstochastic cross-budget choice patterns. A rationalizable stochastic demand system must be a mixture over them and therefore lie inside a certain finite polytope.
Formalizing this insight requires some notation.
Elements of $\mathcal{X}$ will be called patches. Patches that are part of more than one budget plane will be called intersection patches. Each budget can be uniquely expressed as union of patches; the number of patches that jointly comprise budget $\mathcal{B}_j$ will be called $I_j$. Note that $\sum_{j=1}^{J}I_{j} \geq I$, strictly so (because of multiple counting of intersection patches) if any two budget planes intersect.
The partition $\mathcal{X}$ is the finite universal choice set alluded to earlier. The basic idea is that all choices from a given budget that are on the same patch induce the same directly revealed preferences, so are equivalent for the purpose of our test. Conversely, stochastic rationalizability does not at all constrain the distribution of demand on any patch. Therefore, rationalizability of $(P_1,\dots,P_J)$ can be decided by only considering the cross-sectional probabilities of patches on the respective budgets. We formalize this as follows.
Thus, the vector representation of a stochastic demand system lists the probability masses that it assigns to patches.
Example (ref) continued. This example has a total of $5$ patches, namely the $4$ line segments identified in the figure and the intersection. The vector representations of $(\mathcal{B}_1,\mathcal{B}_2)$ and $(P_1,P_2)$ have $6$ components because the intersection patch is counted twice. If one disregards intersection patches (as we will do later), the vector representation of $(P_1,P_2)$ is $(\pi_{1|1},\pi_{2|1},\pi_{1|2},\pi_{2|2})$; see Figure (ref).
Next, a stochastic demand system is rationalizable iff it is a mixture of rationalizable nonstochastic demand systems. To intuit this, one may literally think of the latter as characterizing rational individuals. It follows that the vector representation of a rationalizable stochastic demand system must be the corresponding mixture of vector representations of rationalizable nonstochastic demand systems. Thus, define:
We then have:
Theorem (ref) reduces the problem of testing (ref) to testing a null hypothesis about finite (though possible rather long) vector of probabilities. Furthermore, this hypothesis can be expressed as finite cone, a simple but novel observation that will be crucial for testing.\footnote{The idea of patches, as well as equivalence of (i) and (ii) in Theorem (ref), were anticipated by \citeasnoun{mcfadden-2005}. While the explanation of patches is arguably unclear and (i)$\Leftrightarrow $(ii) is not explicitly pointed out, the idea is unquestionably there. The observation that (ii)$\Leftrightarrow $(iii) (more importantly: the idea of using this for testing) is new.}
We conclude this subsection with a few remarks.
Simplification if demand is continuous. Intersection patches are of lower dimension than budget planes. Thus, if the distribution of demand is continuous, their probabilities must be zero, and they can be eliminated from $\mathcal{X}$. This may considerably simplify $A$. Also, each remaining patch belongs to exactly one budget plane, so that $\sum_{j=1}^{J}I_{j}=I$. We impose this simplification henceforth and in our empirical application, but none of our results depend on it.
GARP vs SARP. Rationalizability of nonstochastic demand systems can be defined using either GARP or SARP. SARP will define a smaller matrix $A$, but nothing else changes. However, columns that are consistent with GARP but not SARP must select at least three intersection patches, so that GARP and SARP define the same $A$ if $\mathcal{X}$ was simplified to reflect continuous demand.
Generality. At its heart, Theorem (ref) only uses that choice from finitely many budgets reveals finitely many distinct revealed preference relations. Thus, it applies to any setting with finitely many budgets, irrespective of budgets' shapes. For example, the result was applied to kinked budget sets in \citeasnoun{Manski14} and could be used to characterize rationalizable choice proportions over binary menus, i.e. the Linear Order Polytope. The result furthermore applies to the “random utility" extension of any other revealed preference characterization that allows for discretization of choice space; see \citeasnoun{DKQS16} for an example.
We next illustrate with a few examples. For simplicity, we presume continuous demand and therefore disregard intersection patches.
Example (ref) continued. Dropping the intersection patch, we have $I=4$ patches. Index vector representations as in Figure (ref), then the only excluded behavior is $(1,0,1,0)'$, thus
The column cone of $A$ can be explicitly written as $\{(\nu_1,\nu_2+\nu_3,\nu_2,\nu_1+\nu_3)':\nu_1,\nu_2,\nu_3 \geq 0\}$. As expected, the only restriction on $\pi$ beyond adding-up constraints is that $\pi_{1|1}+\pi_{1|2}\leq 1$.
It should be clear now (and we formally show below) that the size of $A$, hence the cost of computing it, may escalate rapidly as examples get more complicated. We next elaborate how to compute $A$ from a vector of prices $(p_{1},...,p_{J})$. For ease of exposition, we drop intersection patches (thus SARP=GARP) and add remarks on generalization along the way. We split the problem into two subproblems, namely checking whether a binary “candidate" vector $a$ is in fact a column of $A$ and finding all such vectors.
Consider any binary $I$-vector $a$ with one entry of $1$ on each subvector corresponding to one budget. This vector corresponds to a nonstochastic demand system. It is a column of $A$ if this demand system respects SARP, in which case we call $a$ rationalizable.
To check such rationalizability, we initially extract a direct revealed preference relation over budgets. Specifically, if (an element of) $x_{i|j}$ is chosen from budget $\mathcal{B}_j$, then all budgets that are above $x_{i|j}$ are direct revealed preferred to $\mathcal{B}_j$. This information can be extracted extremely quickly.\footnote{In practice, we compute a $(I \times J)$-matrix $X$ where, for example, the $i$-th row of $X$ is $(0,-1,1,1,1)$ if $x_{i|1}$ is on budget $\mathcal{B}_1$, below budget $\mathcal{B}_2$, and above the remaining budgets. This allows to vectorize construction of direct revealed preference relations, including strict vs. weak revealed preference, though we do not use the distinction.}
We next exploit a well-known representation: Preference relations over $J$ budgets can be identified with directed graphs on $J$ labeled nodes by equating a directed link from node $i$ to node $j$ with revealed preference for $\mathcal{B}_i$ over $\mathcal{B}_j$. A preference relation fulfills SARP iff this graph is acyclic. This can be tested in quadratic time (in $J$) through a depth-first or breadth-first search. Alternatively, the Floyd-Warshall algorithm Floyd62 is theoretically slower but also computes rapidly in our application. Importantly, increasing $K$ does not directly increase the size of graphs checked in this step, though it allows for more intricate patterns of overlap between budgets and, therefore, potentially for richer revealed preference relations.
If intersection patches are retained, then one must distinguish between weak and strict revealed preference, and the above procedure tests SARP as opposed to GARP. To test GARP, one could use Floyd-Warshall or a recent algorithm that achieves quadatic time TSS15.
A total of $\prod_{j=1}^J I_J$ vectors $a$ could in principle be checked for rationalizability. Doing this by brute force rapidly becomes infeasible, including in our empirical application. However, these vectors can be usefully identified with the leaves (i.e. the terminal nodes) of a tree constructed as follows: (i) The root of the tree has no label and has $I_1$ children labelled $(x_{1|1},\dots,x_{I_1|1})$. (ii) Each child in turn has $I_2$ children labelled $(x_{1|2},\dots,x_{I_2|2})$, and so on for a total of $J$ generations beyond the root. Then there is a one-to-one mapping from leaves of the tree to conceivable vectors $a$, namely by identifying every leaf with the nonstochastic demand system that selects its ancestors. Furthermore, each non-terminal node of the tree can be identified with an incomplete $a$-vector that specifies choice only on the first $j<J$ budgets. The methods from Section (ref) can be used to check rationalizability of such incomplete vectors as well.
Our suggested algorithm for computing $A$ is a depth-first search of this tree. Importantly, rationalizability of the implied (possibly incomplete) vector $a$ is checked at each node that is visited. If this check fails, the node and all its descendants are abandoned. A column of $A$ is discovered whenever a terminal node has been reached without detecting a choice cycle. Pseudocode for the tree search algorithm is displayed in Appendix B.
The cost of computing $A$ will escalate rapidly under any approach, but some meaningful comparison is possible. To do so, we consider three sequences, all indexed by $J$, whose first terms are displayed in Table (ref). First, any two distinct nonstochastic demand systems induce distinct direct revealed preference relations; hence, $H$ is bounded above by $\bar H_J$, the number of distinct directed acyclic graphs on $J$ labeled nodes. This sequence -- and hence the worst-case cost of enumerating the columns of $A$, not to mention computing them -- is well understood Robinson73, increases exponentially in $J$, and is displayed in the first row of Table (ref).
Next, a worst-case bound on the number of terminal nodes of the aforementioned tree, hence on vectors that a brute force algorithm would check, is $2^{J(J-1)}$. This is simply because the number of conceivable candidate vectors $a$ equals $\prod_{j=1}^J I_j$, and every $I_j$ is bounded above by $2^{J-1}$. The corresponding sequence is displayed in the last row of Table (ref).
Finally, some tedious combinatorial book-keeping (see Appendix C) reveals that the depth-first search algorithm visits at most $\sum_{j=2}^J \bar{H}_{j-1}2^{j(J+2-j)-2}$ nodes. This sequence is displayed in the middle row of the table, and the gain over a brute force approach is clear.\footnote{The comparison favors brute force because some nodes visited by a tree search are non-terminal, in which case rationalizability is easier to check. For example, this holds for $16$ of the $64$ nodes a tree search visits in Example (ref).}
It is easy to show that the ratio of any two sequences in the table grows exponentially. Also, all of the bounds are in principle attainable (though restricting $K$ may improve them) and are indeed attained in Examples (ref) and (ref). In our empirical application, the bounds are far from binding (see Tables (ref) and (ref) for relevant values of $H$), but brute force was not always feasible, and the tree search improved on it by orders of magnitude in some cases where it was.
A modest amount of problem-specific adjustment may lead to further improvement. The key to this is contained in the following result.
If the geometry of budgets allows it -- this is particularly likely if budgets move outward over time and even guaranteed if some budget planes are parallel -- Theorem (ref) can be used to construct columns of $A$ recursively from columns of $A$-matrices that correspond to a smaller $J$. The gain can be tremendous because, at least with regard to worst-case cost, one effectively moves one or more columns to the left in Table (ref). A caveat is that application of Theorem (ref) may require to manually reorder budgets so that it applies. Also, while the internal ordering of $(\mathcal{B}_1,\dots,\mathcal{B}_M)$ and $(\mathcal{B}_{M+1},\dots,\mathcal{B}_{J-1})$ does not matter, the theorem may apply to distinct partitions of the same set of budgets. In that case, any choice of partition will accelerate computations, but we have no general advice on which is best. We tried the refinement in our empirical application, and it considerably improved computation time for some of the largest matrices. However, the tree search proved so fast that, in order to keep it transparent, our replication code omits this step.
This section lays out our statistical testing procedure in the idealized situation where, for finite $J$, repeated cross-sectional observations of demand over $J$ periods are available to the econometrician. Formally, for each $1 \leq j \leq J$, suppose one observes $N_j$ random draws of $y$ distributed according to $P_j$ defined in (ref). Define $N = \sum_{j=1}^J N_j$ for later use. Clearly, $P_j$ can be estimated consistently as $N_j \uparrow \infty$ for each $j$, $1 \leq j \leq J$. The question is whether the estimated distributions may, up to sampling uncertainty, have arisen from a RUM. We define a test statistic and critical value and show that the resulting test is uniformly asymptotically valid over an interesting range of d.g.p.'s.
By Theorem (ref), we wish to test:
\
(H$_A$): There exist $\nu \geq 0$ such that $A\nu =\pi$.
\
This hypothesis is equivalent to
\
(H$_{B}$): \quad $\min_{\eta \in \mathcal C}[\pi -\eta ]^{\prime }\Omega \lbrack \pi -\eta ]=0$,
\
where $\Omega $ is a positive definite matrix (restricted to be diagonal in our inference procedure) and $\mathcal{C}:=\{A\nu |\nu \geq 0\}$ is a convex cone in $\mathbf{R}^{I}$. The solution $ \eta _{0}$ of (H$_{B}$) is the projection of $\pi \in \mathbf{R} _{+}^{I}$ onto $\mathcal{C}$ under the weighted norm $\Vert x\Vert _{\Omega }=\sqrt{ x^{\prime }\Omega x}$. The corresponding value of the objective function is the squared length of the projection residual vector. The projection $\eta _{0}$ is unique, but the corresponding $\nu $ is not. Stochastic rationality holds if and only if the length of the residual vector is zero.
A natural sample counterpart of the objective function in (H$_{B}$) would be $\min_{\eta \in \mathcal C}[\hat{\pi}-\eta ]^{\prime }\Omega \lbrack \hat{\pi}-\eta ]$, where $\hat{\pi}$ estimates $\pi $, for example by sample choice frequencies. The usual scaling yields
Once again, $\nu $ is not unique at the optimum, but $\eta =A\nu $ is. Call its optimal value $\hat{\eta}$. Then $\hat{\eta}=\hat{\pi}$, and $\mathcal{J}_N=0$, if the estimated choice probabilities $\hat{\pi}$ are stochastically rationalizable; obviously, our null hypothesis will be accepted in this case.
We next explain how to get a valid critical value for $\mathcal{J}_N$ under the assumption that $\hat{\pi}$ estimates the probabilities of patches by corresponding sample frequencies and that one has $R$ bootstrap replications $\hat{\pi}^{\ast (r)},r=1,...,R$. Thus, $\hat{\pi}^{\ast (r)}-\hat{\pi}$ is a natural bootstrap analog of $\hat{\pi}-\pi$. We will make enough assumption to ensure that its distribution consistently estimates the distribution of $\hat{\pi}-\pi_0$, where $\pi_0$ is the true value of $\pi$. The main difficulty is that one cannot use $\hat{\pi}$ as bootstrap analog of $\pi_0$.
Our bootstrap procedure relies on a tuning parameter $\tau_N$ chosen s.t. $\tau _{N}\downarrow 0$ and $\sqrt{N} \tau _{N}\uparrow \infty $.\footnote{In this section's simplified setting and if $\hat{\pi}$ collects sample frequencies, a reasonable choice would be
where $\underline{N}=\min_{j}N_{j}$ and $N_{j}$ is the number of observations on Budget $\mathcal{B}_{j}$: see (ref). This choice corresponds to the “BIC choice” in \citeasnoun{andrews-soares}. We will later propose a different $\tau_N$ based on how $\pi$ is in fact estimated.} Also, we restrict $\Omega$ to be diagonal and positive definite and let $\mathbf{1}_{H}$ be a $H$-vector of ones\footnote{In principle, $\mathbf{1}_{H}$ could be any strictly positive $H$-vector, though a data based choice of such a vector is beyond the scope of the paper.}. The restriction on $\Omega$ is important: Together with a geometric feature of the column vectors of the matrix $A$, it ensures that constraints which are fulfilled but with small slack become binding through the Cone Tightening algorithm we are about to describe. A non-diagonal weighting matrix can disrupt this property. For further details on this point and its proof, the reader is referred to Appendix A. Our procedure is as follows:
\
The object $\hat \eta_{\tau_N}$ is the true value of $\pi$ in the bootstrap population, i.e. it is the bootstrap analog of $\pi_0$. It differs from $\hat{\pi}$ through a “double recentering.” To disentangle the two recenterings, suppose first that $\tau_N=0$. Then inspection of step (i) of the algorithm shows that $\hat{\pi}$ would be projected onto the cone $\mathcal{C}$. This is a relatively standard recentering “onto the null” that resembles recentering of the $J$-statistic in overidentified GMM. However, with $\tau_N>0$, there is a second recentering because the cone $\mathcal{C}$ itself has been tightened. We next discuss why this recentering is needed.
Our testing problem is related to the large literature on inequality testing but adds an important twist. Writing $\{a_{1},a_{2},...,a_{H}\}$ for the column vectors of $A$, one has
i.e. the set $\mathcal{C}$ is a finitely generated cone. The following result, known as the {Weyl-Minkowski Theorem}, provides an alternative representation that is useful for theoretical developments of our statistical testing procedure.\footnote{See \citeasnoun{Gruber}, \citeasnoun{Grunbaum03}, and \citeasnoun{Ziegler}, especially Theorem 1.3, for these results and other materials concerning convex polytopes used in this paper.}
The “only if” part of the theorem (which is Weyl's Theorem) shows that our rationality hypothesis $ \pi \in \mathcal{C}, \mathcal{C} = \{A\nu|\nu \geq 0\} $ in terms of a $\mathcal{V}$-representation can be re-formulated in an $\mathcal{H}$-representation using an appropriate matrix $B$, at least in theory. If such $B$ were available, our testing problem would resemble tests of
based on a quadratic form of the empirical discrepancy between $B\theta$ and $\eta$ minimized over $\eta \in {\bf R}_+^q$. This type of problem has been studied extensively; see references in Section (ref). Its analysis is intricate because the limiting distribution of such a statistic depends discontinuously on the true value of $B\theta$. One common way to get a critical value is to consider the globally least favorable case, which is $\theta =0$. A less conservative strategy widely followed in the econometric literature on moment inequalities is Generalized Moment Selection (GMS; see \citeasnoun{andrews-soares}, \citeasnoun{bugni-2010}, \citeasnoun{Canay10}). If we had the $\mathcal{H}$-representation of $\mathcal{C}$, we might conceivably use the same technique. However, the duality between the two representations is purely theoretical: In practice, $B$ cannot be computed from $A$ in high-dimensional cases like our empirical application.
We therefore propose a tightening of the cone $\mathcal{C}$ that is computationally feasible and will have a similar effect as GMS. The idea is to tighten the constraint on $\nu$ in (ref). In particular, define $\mathcal{C}_{\tau _{N}}:=\{A\nu |\nu \geq \tau _{N}\mathbf{1}_{H}/H\}$ and define $\hat{\eta}_{\tau _{N}}$ as
Our proof establishes that constraints in the $\mathcal{H}$-representation that are almost binding at the original problem's solution (i.e., their slack is difficult to be distinguished from zero at the sample size) will be binding with zero slack after tightening. Suppose that $\sqrt{N}(\hat{\pi}-\pi )\rightarrow _{d}N(0,S)$ and let $\hat{S}$ consistently estimate $S$. Let $\tilde{\eta}_{\tau _{N}}:=\hat{\eta}_{\tau _{N}}+\frac{1}{\sqrt{N}}N(0,\hat{S})$ or a bootstrap random variable and use the distribution of
to approximate the distribution of $\mathcal{J}_N$. This has the same theoretical justification as the inequality selection procedure. Unlike the latter, however, it avoids the use of an $\mathcal{H}$-representation, thus offering a computationally feasible testing procedure.
To further illustrate the duality between $\mathcal{H}$- and $\mathcal{V}$-representations, we revisit the first two examples. It is not possible to compute $B$-matrices in our empirical application.
Example (ref) continued. With two intersecting budget planes, the cone $\mathcal{C}$ is represented by
The first two rows of $B$ are nonnegativity constraints (the other two such constraints are redundant), the next two rows are an equality constraint forcing the sum of probabilities to be constant across budgets, and only the last constraint is a substantive economic constraint. If the estimator $\hat{\pi}$ fulfills the first four constraints by construction, then the testing problem simplifies to a test of $(1,0,0,-1)\pi \leq 0$, the same condition identified earlier.
Example (ref) continued. Eliminating nonnegativity and adding-up constraints for brevity, numerical evaluation reveals
The first three rows are constraints on pairs of budgets that mirror the last row of (ref). The next two constraints are not implied by these, nor by additional constraints in \citeasnoun{Kawaguchi}, but they imply the latter.
We now provide a detailed justification. First, we formalize the notion that choice probabilities are estimated by sample frequencies. Thus, for each budget set $\mathcal{B}_j$, denote the choices of $N_{j}$ individuals, indexed by $ n=1,...,N_{j}$, by
Assume that one observes $J$ random samples $\{\{d_{i|j,n}\}_{i=1}^{I_{j}} \}_{n=1}^{N_{j}}$, $j=1,2,...,J$. For later use, define
An obvious way to estimate the vector $\pi $ is to use choice frequencies
The next lemma, among other things, shows that our tightening of the $\mathcal{V}$-representation of $\mathcal{C}$ is equivalent to a tightening its $\mathcal{H}$-representation but leaving $B$ unchanged. For a matrix $B$, let $\mathrm{col}(B) $ denote its column space.
Lemma (ref) is {\it not} just a re-statement of the {Minkowski-Weyl theorem} for polyhedra, which would simply say $\mathcal{C}_\tau = \{A \nu|\nu \geq (\tau/H) \mathbf{1}_{H}\}$ is alternatively represented as an intersection of closed halfspaces. The lemma instead shows that the inequalities in the $\mathcal H$-representation becomes tighter by $\tau \phi$ after tightening the $\mathcal V$-representation by $\tau_N {\bf 1}_H/H$, with the same matrix of coefficients $B$ appearing both for $\mathcal{C}$ and $\mathcal{C}_\tau$. Note that for notational convenience, we rearrange rows of $B$ so that the genuine inequalities come first and pairs of inequalities that represent equality constraints come last.\footnote{In the matrix displayed in (ref), the third and fourth row would then come last.} This is w.l.o.g.; in particular, the researcher does not need to know which rows of $B$ these are. Then as we show in the proof, the elements in $\phi$ corresponding to the equality constraints are automatically zero when we tighten the space for {\it all} the elements of $\nu$ in the $\mathcal V$-representation. This is a useful feature that makes our methodology work in the presence of equality constraints.
The following assumptions are used for our asymptotic theory.
The econometrician also observes the normalized price vector $p_j$, which is fixed in this section, for each $1 \leq j \leq J$. Let ${\mathcal{P}}$ denote the set of all $\pi$'s that satisfy Condition (ref) in Appendix A for some (common) value of $ (c_1,c_2)$.
While it is obvious that our tightening contracts the cone, the result depends on a more delicate feature, namely that we (potentially) turn non-binding inequalities from the $\mathcal{H}$-representation into binding ones but not vice versa. This feature is not universal to cones as they get contracted. Our proof establishes that it generally obtains if $\Omega $ is the identity matrix and all corners of the cone are acute. In this paper's application, we can further exploit the cone's geometry to extend the result to any diagonal $\Omega $.\footnote{It is possible to replace $\Omega$ with its consistent estimator and retain uniform asymptotic validity, if we further impose a restriction on the class of distributions over which we define the size of our test. Note, however, that our $\mathcal P$ in our Theorem (ref) (and its variants in Theorems (ref) and (ref)) allows for some elements of the vector $\pi$ being zeros. This makes the use of the reciprocals of estimated variances for the diagonals of the weighting matrix potentially problematic, as it invalidates the asymptotic uniform validity since the required triangular CLT does not hold under parameter sequences where the elements of $\pi$ converge to zeros. The use of fixed $\Omega$, which we recommend in implementing our procedure, makes contributions from these terms asymptotically negligible, thereby circumventing this problem.} Our method immediately applies to other testing problems featuring $\mathcal{V}$ -representations if analogous features can be verified.
The methodology outlined in Section (ref) requires that (i) the observations available to the econometrician are drawn on a finite number of budgets and (ii) the budgets are given exogenously, that is, unobserved heterogeneity and budgets are assumed to be independent. These conditions are naturally satisfied in some applications. The empirical setting in Section (ref), however, calls for modifications because Condition (i) is certainly violated in it and imposing Condition (ii) would be very restrictive. These are typical issues for a survey data set. This section addresses them.
Let $P_u$ denote the marginal probability law of $u$, which we assume does not depend on $j$. We do not, however, assume that the laws of other random elements, such as income, are time homogeneous. Let $w = \log(W)$ denote log total expenditure, and suppose the researcher chooses a value $\underline{w}_j$ for $w$ for each period $j$. Note that our algorithm and asymptotic theory remain valid if multiple values of $w$ are chosen for each period. Let $w_{n(j)}$ be the log total expenditure of consumer $n(j)$, $1\leq n(j)\leq N_{j}$ observed in period $j$.
The econometrician also observes the unnormalized price vector $\tilde p_j$, which is fixed, for each $1 \leq j \leq J$.
We first assume that the total expenditure is exogenous, in the sense that $ w {\bot\negthickspace\negthickspace\bot} u $ holds under every $P^{(j)}, 1 \leq j \leq J$. This exogeneity assumption will be relaxed shortly. Let $\pi_{i|j}(w):=\Pr\{d_{i|j,n(j)}=1|w_{n(j)}= \underline{w}_j \}$ and writing $\pi _{j}:=(\pi _{1|j},...,\pi _{I_{j}|j})^{\prime }$ and $\pi :=(\pi _{1}^{\prime },...,\pi _{J}^{\prime })^{\prime }=(\pi _{1|1},\pi _{2|1},...,\pi _{I_{J}|J})^{\prime }$, the stochastic rationality condition is given by $ \pi \in \mathcal{C} $ as before. Note that this $\pi$ can be estimated by standard nonparametric procedures. For concreteness, we use a series estimator, as defined and analyzed in Appendix A. The smoothed version of $ \mathcal{J}_N$ (also denoted $\mathcal J_N$ for simplicity) is obtained using the series estimator for $\hat{\pi}$ in (ref). In Appendix A we also present an algorithm for obtaining the bootstrapped version $\tilde {\mathcal J}_N$ of the smoothed statistic.
In what follows, $F_{j}$ signifies the joint distribution of $ (d_{i|j,n(j)},w_{n(j)})$. Let $\mathcal{F}$ be the set of all $ (F_{1},...,F_{J}) $ that satisfy Condition (ref) in Appendix A for some $ (c_1,c_2,\delta,\zeta(\cdot ))$.
Next, we relax the assumption that consumer's utility functions are realized independently from $W$. For each fixed value $\underline{w}_j$ and the unnormalized price vector $\tilde p_j$ in period $j$, $1 \leq j \leq J$, define the endogeneity corrected conditional probability{\footnote{This is the conditional choice probability if $p$ is (counterfactually) assumed to be exogenous. We call it “endogeneity corrected" instead of “counterfactual" to avoid confusion with rationality constrained, counterfactual prediction.}
where $D_j(w,u) := D(\tilde{p}_j/e^w,u)$. Then Theorem (ref) still applies to $$ \pi_{\rm{EC}} := [\pi(p_1,x_{1|1}),...,\pi(p_{1},x_{I_1|1}),\pi(p_2,x_{1|2}),...,\pi(p_{2},x_{I_2|2}),...,\pi(p_J,x_{1|J}),...,\pi(p_{J},x_{I_J|J})]'. $$ Suppose there exists a control variable $\varepsilon$ such that $ {w {\bot\negthickspace\negthickspace\bot} u | \varepsilon} $ holds under every $P^{(j)}, 1 \leq j \leq J$. See (ref) in Section A for an example. We propose to use a fully nonparametric, control function-based two-step estimator, denoted by $\widehat{\pi_{\mathrm{EC}}}$, to define our endogeneity-corrected test statistic ${\mathcal J}_{\mathrm{EC}_N}$; see Appendix A for details. For this, the bootstrap procedure needs to be adjusted appropriately to obtain the bootstrapped statistic $\tilde {\mathcal J}_{\mathrm{EC}_N}$: once again, the reader is referred to Appendix A. This is the method we use for the empirical results reported in Section (ref). Let $z_{n(j)}$ be the $n(j)$-th observation of the instrumental variable $z$ in period $j$.
The econometrician also observes the unnormalized price vector $\tilde p_j$, which is fixed, for each $1 \leq j \leq J$.
In what follows, $F_{j}$ signifies the joint distribution of $ (d_{i|j,n(j)},w_{n(j)},z_{n(j)})$. Let $\mathcal{F}_{\rm EC}$ be the set of all $ (F_{1},...,F_{J}) $ that satisfy Condition (ref) in Appendix A for some $(c_1,c_2,\delta_1,\delta,\zeta_r(\cdot),\zeta_s(\cdot),\zeta_1(\cdot))$. Then we have:
We next analyze the performance of Cone Tightening in a small Monte Carlo study. To keep examples transparent and to focus on the core novelty, we model the idealized setting of Section (ref), i.e. sampling distributions are multinomial over patches. In addition, we focus on Example (ref), for which an $\mathcal{H}$-representation in the sense of Weyl-Minkowski duality is available; see displays (ref) and (ref) for the relevant matrices.\footnote{This is also true of Example (ref), but that example is too simple because the test reduces to a one-sided test about the sum of two probabilities, and the issues that motivate Cone Tightening or GMS go away. We verified that all testing methods successfully recover this and achieve excellent size control, including if tuning parameters are set to $0$.} This allows us to alternatively test rationalizability through a moment inequalities test that ensures uniform validity through GMS.\footnote{The implementation uses a “Modified Method of Moments" criterion function, i.e. $S_1$ in the terminology of \citeasnoun{andrews-soares}, and the hard thresholding GMS function, i.e. studentized intercepts above $-\kappa_N$ were set to $0$ and all others to $-\infty$. The tuning parameter is set to $\kappa_N=\sqrt{ln(N_j)}$.}
Data were generated from a total of 31 d.g.p.'s described below and for sample sizes of $N_j \in \{100,200,500,1000\}$; recall that these are per budget, i.e. each simulated data set is based on 3 such samples. The d.g.p.'s are parameterized by the $\pi$-vectors reported in Table (ref). They are related as follows: $\pi_0$ is in the interior of $\mathcal{C}$; $\pi_2$, $\pi_4$, and $\pi_6$ are outside it; and $\pi_2$, $\pi_3$, and $\pi_5$ are on its boundary. Furthermore, $\pi_1=(\pi_0+\pi_2)/2$, $\pi_3=(\pi_0+\pi_4)/2$, and $\pi_5=(\pi_0+\pi_6)/2$. Thus, the line segment connecting $\pi_0$ and $\pi_2$ intersects the boundary of $\mathcal{C}$ precisely at $\pi_1$ and similarly for the next two pairs of vectors. We compute “power curves" along those $3$ line segments at $11$ equally spaced points, i.e. changing mixture weights in increments of $.1$. This is replicated $500$ times at a bootstrap size of $R=499$. Nominal size of the test is $\alpha=.05$ throughout. Ideally, it should be exactly attained at the vectors $\{\pi_1,\pi_3,\pi_5\}$.
Results are displayed in Table (ref). Noting that the vectors are not too different, we would argue that the simulations indicate reasonable power. Adjustments that ensure uniform validity of tests do tend to cause conservatism for both GMS and cone tightening, but size control markedly improves with sample size.\footnote{We attribute some very slight nonmonotonicities in the “power curves" to simulation noise.} While Cone Tightening appears less conservative than GMS in these simulations, we caution that the tuning parameters and the distance metrics underlying the test statistics are not directly comparable.
The differential performance across the three families of d.g.p.'s is expected because the d.g.p.'s were designed to pose different challenges. For both $\pi_1$ and $\pi_3$, one constraint is binding and three more are close enough to binding that, at the relevant sample sizes, they cannot be ignored. This is more the case for $\pi_3$ compared to $\pi_1$. It means that GMS or Cone Tightening will be necessary, but also that they are expected to be conservative. The vector $\pi_5$ has three constraints binding, with two more somewhat close. This is a worst case for naive (not using Cone Tightening or GMS) inference, which will rarely pick up all binding constraints. Indeed, we verified that inference with $\tau_N=0$ or $\kappa_N=0$ leads to overrejection. Finally, $\pi_2$ and $\pi_4$ fulfill the necessary conditions identified by \citeasnoun{Kawaguchi}, so that his test will have no asymptotic power at a parameter value in the first two panels of Table (ref).
We apply our methods to data from the U.K. Family Expenditure Survey, the same data used by BBC. Our testing of a RUM can, therefore, be compared with their revealed preference analysis of a representative consumer. To facilitate this comparison, we use the same selection from these data, namely the time periods from 1975 through 1999 and households with a car and at least one child. The number of data points used varies from 715 (in 1997) to 1509 (in 1975), for a total of 26341. For each year, we extract the budget corresponding to that year's median expenditure and, following Section (ref), estimate the distribution of demand on that budget with polynomials of order $3$. Like BBC, we assume that all consumers in one year face the same prices, and we use the same price data. While budgets have a tendency to move outward over time, there is substantial overlap of budgets at median expenditure. To account for endogenous expenditure, we again follow Section (ref), using total household income as instrument. This is also the same instrument used in BBC (2008).
We present results for blocks of eight consecutive periods and the same three composite goods (food, nondurable consumption goods, and services) considered in BBC.\footnote{As a reminder, Figure (ref) illustrates the application. The budget is the 1993 one as embedded in the 1986-1993 block of periods, i.e. the figure corresponds to a row of Table (ref).} For all blocks of seven consecutive years, we analyze the same basket but also increase the dimensionality of commodity space to 4 or even 5. This is done by first splitting nondurables into clothing and other nondurables and then further into clothing, alcoholic beverages, and other nondurables. Thus, the separability assumptions that we (and others) implicitly invoke are successively relaxed. We are able to go further than much of the existing literature in this regard because, while computational expense increases with $K$, our approach is not subject to a statistical curse of dimensionality.\footnote{Tables (ref) and (ref) were computed in a few days on Cornell's ECCO cluster (32 nodes). An individual cell of a table can be computed in reasonable time on any desktop computer. Computation of a matrix $A$ took up to one hour and computation of one $\mathcal{J}_N$ about five seconds on a laptop.}
Regarding the test's statistical power, increasing the dimensionality of commodity space can in principle cut both ways. The number of rationality constraints increases, and this helps if some of the new constraints are violated but adds noise otherwise. Also, the maintained assumptions become weaker: In principle, a rejection of stochastic rationalizability at 3 but not 4 goods might just indicate a failure of separability.
Tables (ref) and (ref) summarize our empirical findings. They display test statistics, p-values, and the numbers $I$ of patches and $H$ of rationalizable demand vectors; thus, matrices $A$ are of size $(I \times H)$. All entries that show $\mathcal{J}_N=0$ and a corresponding p-value of $1$ were verified to be true zeros, i.e. $\hat{\pi}_{EC}$ is rationalizable. All in all, it turns out that estimated choice probabilities are typically not stochastically rationalizable, but also that this rejection is not statistically significant.\footnote{In additional analyses not presented here, we replicated these tables using polynomials of degree 2, as well as setting $\tau_N=0$. The qualitative finding of many positive but insignificant test statistics remains. In isolation, this finding may raise questions about the test's power. However, the test exhibits reasonable power in our Monte Carlo exercise and also rejects rationalizability in an empirical application elsewhere Hubner. \newline We also checked whether small but positive test statistics are caused by adding-up constraints, i.e. by the fact that all components of $\pi$ that correspond to one budget must jointly be on some unit simplex. The estimator $\hat{\pi}$ can slightly violate this. Adding-up failures occur but are at least one order of magnitude smaller than the distance from a typical $\hat{\pi}$ to the corresponding projection $\hat{\eta}$.}
We identified a mechanism that may explain this phenomenon. Consider the 84-91 entry in Table (ref), where $\mathcal{J}_N$ is especially low. It turns out that one patch on budget $\mathcal{B}_5$ is below $\mathcal{B}_8$ and two patches on $\mathcal{B}_8$ are below $\mathcal{B}_5$. By the reasoning of Example (ref), probabilities of these patches must add to less than $1$. The estimated sum equals $1.006$, leading to a tiny and statistically insignificant violation. This phenomenon occurs frequently and seems to cause the many positive but insignificant values of $\mathcal{J}_N$. The frequency of its occurrence, in turn, has a simple cause that may also appear in other data: If two budgets are slight rotations of each other and demand distributions change continuously in response, then population probabilities of patches like the above will sum to just less than $1$. If these probabilities are estimated independently across budgets, the estimates will frequently add to slightly more than $1$. With $7$ or $8$ mutually intersecting budgets, there are many opportunities for such reversals, and positive but insignificant test statistics may become ubiquitous.
The phenomenon of estimated choice frequencies typically not being rationalizable means that there is need for a statistical testing theory and also a theory of rationality constrained estimation. The former is this paper's main contribution. We leave the latter for future research.
This paper presented asymptotic theory and computational tools for nonparametric testing of Random Utility Models. Again, the null to be tested was that data was generated by a RUM, interpreted as describing a heterogeneous population, where the only restrictions imposed on individuals' behavior were \textquotedblleft more is better\textquotedblright\ and SARP. In particular, we allowed for unrestricted, unobserved heterogeneity and stopped far short of assumptions that would recover invertibility of demand. We showed that testing the model is nonetheless possible. The method is easily adapted to choice problems that are discrete to begin with, and one can easily impose more, or fewer, restrictions at the individual level.
Possibilities for extensions and refinements abound, and some of these have already been explored. We close by mentioning further salient issues.
(1) We provide algorithms (and code) that work for reasonably sized problem, but it would be extremely useful to make further improvements in this dimension.
(2) The extension to infinitely many budgets is of obvious interest. Theoretically, it can be handled by considering an appropriate discretization argument mcfadden-2005. For the proposed projection-based econometric methodology, such an extension requires evaluating choice probabilities locally over points in the space of $p$ via nonparametric smoothing, then use the choice probability estimators in the calculation of the $\mathcal{J}_N$-statistic. The asymptotic theory then needs to be modified. Another approach that can mitigate the computational constraint is to consider a partition of the space of $p$ such that $\mathbf{R}_{+}^{K}=\mathcal{P}_{1}\cup \mathcal{P}_{2}\cdots \cup \mathcal{P}_{M}$. Suppose we calculate the $\mathcal{J}_N$-statistic for each of these partitions. Given the resulting $M$-statistics, say $\mathcal{J}_N^{1},\cdots ,\mathcal{J}_N^{M}$, we can consider $\mathcal{J}_N^{\mathrm{max}}:=\max_{1\leq m\leq M}\mathcal{J}_N^{m}$ or a weighted average of them. These extensions and their formal statistical analysis are of practical interest.
(3) While we allow for endogenous expenditure and therefore for the distribution of $u$ to vary with observed expenditure, we do assume that samples for all budgets are drawn from the same underlying population. This assumption can obviously not be dropped completely. However, it will frequently be of interest to impose it only conditionally on observable covariates, which must then be controlled for. This may be especially relevant for cases where different budgets correspond to independent markets, but also to adjust for slow demographic change as does, strictly speaking, occur in our data. It requires incorporating nonparametric smoothing in estimating choice probabilities as in Section (ref), then averaging the corresponding $\mathcal{J}_N$-statistics over the covariates. This extension will be pursued.
(4) Natural next steps after rationality testing are extrapolation to (bounds on) counterfactual demand distributions and welfare analysis, i.e. along the lines of BBC (2008) or, closer to our own setting, \citeasnoun{Adams16} and \citeasnoun{DKQS16}. This extension is being pursued. Indeed, the tools from Section (ref) have already been used for choice extrapolation (using algorithms from an earlier version of this paper) in \citeasnoun{Manski14}.
(5) The econometric techniques proposed here can be potentially useful in much broader contexts. Indeed, they have already been used to nonparametrically test game theoretic model with strategic complementarities LQS15,LQS18, a novel model of “price preference” DKQS16, and the collective household model Hubner. Even more generally, existing proposals for testing in moment inequality models andrews-guggenberger-2009,andrews-soares,BCS15,romano-shaikh work with explicit inequality constraints, i.e. (in the linear case) $\mathcal{H}$-representations. In settings in which theoretical restrictions inform a $\mathcal{V}$-representation of a cone or, more generally, a polyhedron, the $\mathcal{H}$-representation will typically not be available in practice. We expect that our method can be used in many such cases.
\nocite{BBC03} \nocite{BBC07} \nocite{BBC08}
\pagenumbering{arabic}
\setcounter{section}{0}