Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,201 characters · 18 sections · 43 citation commands
A Frequentist Approach to Revealed Preference Analysis
{\bf Keywords}: revealed preference, finite data, preferences, hypothesis testing
{\bf JEL classification}: C12, C91, D0, D11, D12
Analysts often seek to verify whether a decision maker (DM) is consistent with a given model, using finite choice data, to ensure that policy predictions about behavior and welfare are reliable. A common approach is revealed preference (RP) analysis, which provides a set of exhaustive inequalities that are satisfied if and only if the DM's choices are consistent with a given class of preferences.\footnote{Following the terminology of varian1983non, we refer to a set of inequalities as exhaustive if they are necessary and sufficient for the data set to pass the RP test.} RP tests typically assume that data consist of a finite number of non-stochastic observations. A subtle issue arises when these observations pass the RP test. Namely, the test cannot distinguish between a DM who is truly consistent with the model and one who is not, but for whom there were insufficient data to detect a deviation.
The frequentist approach to addressing this problem in the context of hypothesis testing focuses on statistical power, the probability of rejecting a decision maker’s choices under the assumption that the DM does not adhere to the model. This approach requires constructing parsimonious alternative hypotheses. However, as noted by adams2015models, “The difficulty is that there are many alternatives to rational choice models and no obvious benchmark.” In other words, there is no straightforward way to assess the power of RP tests.\footnote{For example, bronars1987 proposes a power test based on the assumption that choices are uniformly random on the budget frontier.} This paper shows how results from statistical learning theory can be used to circumvent this limitation and enable power analysis.
To illustrate the problem, suppose an analyst observes a consumer making grocery purchases on ten different occasions, each time facing different prices. The consumer's choices satisfy all the revealed preference conditions—there are no apparent “mistakes”. Should the analyst conclude that the consumer is rational? Not necessarily. Ten observations may simply be too few to detect irrationality, even if it is present. With only ten price--consumption pairs, many non-rational choice patterns can slip through undetected, much as a student who answers randomly on a ten-question multiple-choice exam might still pass by chance.
This raises a natural question: how many observations $n$ are needed to detect with high probability a consumer who deviates from rationality by an amount $\varepsilon$? Standard revealed preference analysis cannot answer this question: it tells us whether observed choices could have come from a rational consumer, not whether the available data are sufficient to reliably detect deviations.
Our paper fills this gap using a simple idea. Suppose we observe a consumer's choices at $n$ randomly selected price points. Because demand functions are smooth (satisfying a Lipschitz condition), any function that fits these observations must be close to the consumer's actual choice function everywhere—not just at the observed points—once $n$ is sufficiently large. Consequently, if the true choice function is far from rational, so will any function consistent with the data, and the revealed preference test will reject rationality. If the consumer is rational, the test will never reject. By quantifying the number of observations needed as a function of the detection threshold $\varepsilon$, we show that explicit power guarantees are possible.
Formally, we consider an economic model $g \in \mathcal{G}$, where $\mathcal{G}$ is a class of models satisfying a Lipschitz condition, which maps an exogenous variable (prices) into an endogenous variable (demand). The analyst observes exogenous prices $p_t$ and demand $x_t$ related by $x_t = g(p_t)$. Consider some subclass $\mathcal{F}\subset \mathcal{G}$ and define the ball of size $\varepsilon>0$ around $\mathcal{F}$ as \[\mathcal{F}(\varepsilon):=\Big\{g\in \mathcal{G}: \inf_{g'\in \mathcal{F}}d(g,g')\leq \varepsilon\Big\}.\]
Suppose the analyst can sample $\{p_t\}_{t=1}^n$ independently at random. This generates a dataset of the form
Note that although $g$ is non-stochastic, the data are stochastic because prices are sampled at random. This aligns with the frequentist tradition in which the state of the world is deterministic and all stochasticity arises from the sampling procedure.
For given $\delta$, $\varepsilon > 0$, we aim to find a computable function $n(\delta,\varepsilon)$ such that, whenever a dataset has more than $n(\delta,\varepsilon)$ observations, one can construct tests of size zero and power greater than $\delta$ to distinguish between the null hyposthesis $\mathbf{H}_0$ and the alternative $\mathbf{H}_1$:
We show that if the class $\mathcal{G}$ is Lipschitz and the analyst has access to a revealed preference characterization of $\mathcal{F}$, such tests and function $n(\delta,\varepsilon)$ can be constructed for any choice of $\varepsilon$. Furthermore, we show that when the dataset size $n$ is fixed, one can compute a lower bound on the power $\delta(n,\varepsilon)$ and a lower bound on the closeness of preferences from the model $\varepsilon(n,\delta)$.
We then generalize our main results by studying cases where the analyst does not have access to RP characterizations, but rather some functional restriction $\mathcal{R}:\mathcal{G} \to \mathbb{R}$ of $\mathcal{F}$ such that: \[\mathcal{R}(g)=0 \iff g \in \mathcal{F}.\]
Under some regularity conditions on $\mathcal{R}$, we can conduct the same analysis as above by constructing \[\mathcal{F}(\varepsilon)=\qty{g^\prime \in \mathcal{G}: \mathcal{R}(g)\leq \varepsilon},\] and then forming the hypothesis test in the manner described in (ref). In that case, the size of the test is not zero, but some computable function of $\mathcal{G}$, $\mathcal{R}$, and $\mathcal{F}$.
Our method applies to preference classes where revealed preference characterizations are unknown or computationally infeasible. For Lipschitz demands, we find that the computational issues of several RP tests stem mostly from their knife-edge nature which have size zero. We demonstrate that this approach can be used to construct confidence intervals for smooth functionals of demand, such as changes in cost-of-living indices. Finally, we show how to integrate additional assumptions on the underlying preference classes and adaptive sampling to improve power and confidence intervals.
We emphasize that the results in this paper are primarily existence results that establish what can be learned from finite choice data, rather than providing readily applicable procedures. Our theorems characterize the sample sizes sufficient to achieve desired power levels and the precision attainable for welfare inference, but operationalizing these bounds requires knowledge of quantities—such as the Lipschitz constant of demand and the sampling distribution of prices—that may be difficult to determine in practice. We view our contribution as providing the theoretical foundation for power analysis in revealed preference settings; developing practical implementations to estimate or bound these quantities from data is an important direction for future work.
Section (ref) describes our setup and notation. Section (ref) shows how to construct our tests with fixed power, sample size, and distance from the model (Theorems (ref)--(ref)), followed by power simulations for RP tests. Section (ref) extends our testing results to functional characterizations (Theorem (ref)) and shows how to conduct inference on smooth estimands such as welfare changes (Theorem (ref)). The section also provides functional characterizations of common classes of preferences, such as substitutes and complements, and of choice under uncertainty (Propositions (ref)--(ref)). Section (ref) studies extensions of our approach. Section (ref) concludes.
The RP approach tests the consistency of a finite dataset with utility maximization. The approach originated with \citet*{samuelson1938} and was refined by \citet*{afriat1967construction} and varian82. In particular, afriat1967construction established that the Generalized Axiom of Revealed Preference (GARP) is a necessary and sufficient condition for any finite dataset to be consistent with utility maximization. The appeal of GARP stems from the ease with which the rationality of a dataset can be verified through simple algebraic inequalities.
The idea of constructing an alternative against which to assess the power of revealed-preference (RP) tests dates back to becker1962irrational, who considered consumers choosing randomly on their budget lines. Building on this, bronars1987 examined how often RP tests would detect violations when choices were generated uniformly at random. More recently, andreoni2013power proposed a more realistic benchmark by drawing from the observed distribution of choices, thereby calibrating the alternative hypothesis to match empirical patterns.
Because a parsimonious alternative to fully rational choice is hard to specify, several papers assess predictive power without positing an explicit behavioural model. andreoni2013power define the Afriat Power Index for datasets with no revealed-preference violations. Reporting the smallest budget perturbation needed to induce a violation—small adjustments indicate a sensitive test, large adjustments a weak one. A related approach, based on selten1991properties and applied by beattycrawford11, considers a method which uses the set of all theory-consistent behaviors with respect to all possible behaviors.
Since those seminal contributions, RP theory has explored various extensions, including functional-form restrictions and intertemporal models. Significant contributions include \citet*{kubler2014asset} for expected utility with objective (known) probabilities and \citet*{echeniquesaito15} for subjective probabilities. A similar problem for “translation-invariant” preferences is considered in \citet*{chambersetal16}. \citet*{polisson2020rp} give a general method to construct RP inequalities that applies to several classes of preferences over risk and uncertainty.\footnote{For recent extensive reviews, see CrawfordEmpirical2014, EcheniqueRPreview2019, and demuynck2019samuelson.} We take the set of Lipschitz demands as the alternative, thus covering these subclasses and rationality itself.
The functional approach for testing decision-makers slutsky1915, antonelli71, DeatonMuell1980 relies on the entire demand function and assumes access to infinite data. We build on this literature by extending these tests to finite data. In this context, our work is related to \citet*{aguiar2018classifying}, who adapt the measure of bounded rationality from \citet*{aguiar2017slutsky} to finite data. While their focus is on measuring and classifying bounded rationality, our approach provides tests for general restrictions under preference monotonicity, making it applicable to a broader range of datasets, including pairwise comparisons.
Papers on extrapolating fundamentals from finite data use regularity conditions similar to ours. chambers2021recovering imposes a structural restriction on the space of preferences that is closely related to Lipschitz continuity, but it allows sample complexity to depend on the underlying (unknown) preference, which we rule out by construction. Related Lipschitz-type upper bounds for utilities are also used in chambers2023recovering to deliver finite-sample guarantee results. Finally, chambers2025decision shows that economic models that approximately satisfy axioms can be rationalized by utilities close to those that exactly rationalize the model, providing microfoundations for using our approach to test axiomatic decision models.
The RP literature focuses on exhaustive restrictions in finite datasets that characterize a model.\footnote{The only revealed preference test that is not a characterization which we are familiar with is for probabilistic sophistication \citep*{Epstein2000probsoph}.} In some instances, those tests can pose a significant computational challenge. Indeed, \citet*{echenique2014testing} and \citet*{Cherchyeinteger2015} show that testing for weak separability is NP-Hard. Accordingly, these tests rely on non-polynomial-time algorithms, such as mixed-integer programming \citep*{Cherchyeinteger2015,Hjertstrandmixint2020}. We show that the computational issues can be mitigated by permitting tests with nonzero size.
The deterministic RP approach has recently been extended to stochastic environments by KS2018 and AK2021. In both cases, RP analysis is embedded within a random utility framework, although the latter treats measurement error as the source of randomness. By contrast, in our framework, all randomness arises from the sampling procedure rather than from the decision-maker. Consequently, our approach is closer in spirit to standard RP tests conducted on individual-level data, while still treating individual observations as random.\footnote{In a different direction, allen2024 propose a statistical consumer model who chooses a distribution of demands given prices. They show that standard RP conditions apply to average demands under a weak mean‑expenditure condition, so our method should extend to their framework as well.}
We consider a DM whose choices are observed a finite number of times. Let $Y\subset \mathbb{R}^\ell_{\geq0}$ be a compact and convex consumption space, where $\ell$ denote the number of goods. A DM faces prices $p\in \mathcal{P}$ and income $I\in \mathcal{I}$, where $\mathcal{P}$ is a compact subset of $\mathbb{R}^\ell_{>0}$ and $\mathcal{I} = [a,b] \subset \mathbb{R}_{>0}$. We denote the budget corresponding to a tuple $(p,I)$ by
and denote the set of all budget sets by $\mathcal{D}$. The DM has a choice function $x(p,I)$ with $x(p,I)\in B(p,I)$. Let $\mathcal{G}$ denote the class of all choice functions $x:\mathcal{P}\times\mathcal{I}\to Y$ that satisfy the following Lipschitz property: there exists a constant $L>0$, known to the analyst, such that for all $(p,I),(p',I')\in\mathcal{P}\times\mathcal{I}$, \[ \|x(p,I)-x(p',I')\|\le L\, d\!\left((p,I)^{\top},(p',I')^{\top} \right), \] where $\|\cdot\|$ is the Euclidean norm and $d$ is the Euclidean distance.
In several applications, we restrict attention to choice functions arising from utility maximisation. A preference relation $\succeq\,\subseteq Y\times Y$ is rational if it is complete and transitive. We further assume that preferences are continuous, strictly monotone and strictly convex.\footnote{ Strict monotonicity means that for any $x,y\in Y$, if $x_i\ge y_i$ for all $i=1,\ldots,\ell$ and $x\neq y$, then $x\succ y$. Strict convexity means that for any $x\neq y$ in $Y$ and any $\alpha\in(0,1)$, $x\succeq y$ implies $\alpha x+(1-\alpha)y\succ y$. } Given a rational preference $\succeq$, let $u_\succeq:Y\rightarrow \mathbb{R}$ be a (continuous) utility representation and define the associated Marshallian demand by \[ x_\succeq(p,I)\;:=\;\operatorname*{arg\,max}_{y\in B(p,I)} u_\succeq(y). \] Under strict convexity this maximiser is unique, so $x_\succeq$ is single-valued. Accordingly, we say that a choice function $x\in\mathcal{G}$ is an admissible rational demand if there exists a rational preference $\succeq$ such that $x = x_\succeq$. In what follows, we use the term choice function for an arbitrary element of $\mathcal{G}$, and the term demand function for an admissible rational demand.
A dataset consists of budget-choice pairs $D_n = \{(B_k, x_k)\}_{k=1}^n$. A choice function $x$ rationalizes a dataset $D_n$ if it exactly reproduces the observed choices $x_k = x(B_k)$ for all $k=1,\ldots,n$.
Let $g\in \mathcal{G}$ denote the true choice function, which we may refer to as the underlying model. The analyst draws $n$ price--income pairs independently at random from some distribution $\mu \in \Delta(\mathcal{P} \times \mathcal{I})$, generating a dataset
Let $\mu^n$ denote the product measure on $(\mathcal{P}\times \mathcal{I})^n$.\footnote{Formally, if $\mu \in \Delta(\mathcal{P}\times\mathcal{I})$ is the sampling distribution of price--income pairs, then $\mu^n$ is the product measure on $(\mathcal{P}\times \mathcal{I})^n$, defined for any measurable set $A \subseteq (\mathcal{P}\times \mathcal{I})^n$ as \[ \mu^n(A) \;=\; \int \mathbf{1}_A\big((p_1,I_1),\ldots,(p_n,I_n)\big) \,\mathrm{d}\mu(p_1,I_1)\cdots \mathrm{d}\mu(p_n,I_n). \]} Since $g\in\mathcal{G}$ is fixed, $\mu^n$ induces a distribution $\mu^n_g$ on datasets of size $n$ by mapping each sampled sequence $\big((p_k,I_k)\big)_{k=1}^n$ to its corresponding budget-choice sequence $\big(B_k,g(B_k)\big)_{k=1}^n$. By treating the dataset as random draws from a fixed model, we can study the size and power of tests over repeated sampling. The timing of this conceptual framework is illustrated in Figure (ref).
Let $\mathbf{H}_0$, $\mathbf{H}_1 \subset \mathcal{G}$ denote the null and alternative hypotheses. We have $g \in \mathbf{H}_0$ if the null hypothesis holds and $g \in \mathbf{H}_1$ if the alternative hypothesis holds. A test is a measurable function $\mathcal{T}: \mathcal{D}\to \{0,1\}$ that takes as input a dataset and returns whether the null hypothesis should be rejected, where by convention $\mathcal{T}(D)=1$ means the null is rejected and $\mathcal{T}(D)=0$ means the null is not rejected.
The distinction between size and power lies in the domain over which probabilities are evaluated: size is defined uniformly over all models in the null, taking the supremum over $\mathbf{H}_0$, while power is evaluated at a particular alternative $g \in \mathbf{H}_1$. For brevity, we will henceforth omit the explicit dependence of $\mu^n_g$ on $g$ whenever there is no risk of confusion, and simply write $\mu^n$.
In words, an RP characterization never rejects a dataset consistent with $\mathcal{F}$, and has maximal power among all such size-zero tests.
This section develops finite–sample power bounds that connect the revealed preference framework to the statistical theory of PAC learning. We proceed in three steps. First, we determine the required sample size to achieve a specified power and accuracy level, which may inform experimental design. Second, for a given sample size and accuracy, we derive the attainable power level, helping to gauge the effectiveness of RP tests. Finally, we characterize the minimal separation from $\mathcal{G}$ that is detectable given a sample size and desired power.
This subsection introduces the notion of PAC learnability and states the main result we use as a building block for our procedure. First, define the generalization error between two choice functions $x$, $x':\mathcal{P}\times \mathcal{I} \to Y$ as:
where $\norm{\cdot}$ denotes the Euclidean norm. The next definition introduces the concept of PAC learnability.
In words, $\mathcal{G}$ is PAC learnable by $\mathcal{H}$ if there exists an algorithm that can approximate every demand function in $\mathcal{G}$ by some hypothesis $h \in \mathcal{H}$ within error $\varepsilon$, with probability at least $1-\delta$. Here, a hypothesis $h$ corresponds to a candidate demand function approximating the true choice function. While in the general learning-theory literature the hypothesis class $\mathcal H$ may differ from the target class $\mathcal G$, we restrict our attention to the case $\mathcal H = \mathcal G$. We can now state a key result for our derivations.
This lemma guarantees that, for any distribution $\mu$, the true demand function can be approximated within any pre-specified accuracy $\varepsilon$, with confidence at least $1-\delta$, from a sample size that grows only polynomially in $1/\varepsilon$ and $1/\delta$. Our proof follows from a simple extension of beigman2006learning, who prove that Lipschitz demand functions are learnable under the above sample guarantees. However, their proof relies on a simple packing-number argument that does not use the fact that the hypothesis class is assumed to consist of demand functions. They also provide a constructive algorithm that achieves this bound, thereby delivering both sample efficiency and polynomial-time computation.\footnote{\citet*{kubler2020identification} show that when the sampling distribution $\mu$ is uniform over $\mathcal{P}$, the assumption of a known Lipschitz constant can be relaxed, but they do not provide an explicit procedure to compute the relevant constants.}
For each $(\varepsilon,\delta)$, let $n(\varepsilon,\delta)$ denote the minimal sample size such that there exists some learning algorithm (possibly depending on $(\varepsilon,\delta)$) that $(\varepsilon,\delta)$--learns $\mathcal G$ in the sense of Definition (ref). In particular, for any specific algorithm $L$ satisfying Definition (ref) with sample size $m_L(\varepsilon,\delta)$, we have $n(\varepsilon,\delta)\le m_L(\varepsilon,\delta)$. Moreover, $n(\varepsilon,\delta)$ is weakly decreasing in $\varepsilon$ for each fixed $\delta$.
Our first result presents the “forward” formulation that characterises the sample size required to ensure that an RP test has power at least $1-\delta$ against all alternatives that are $\varepsilon$–separated from the null class.
In words, this result guarantees that the test rejects any demand strictly more than $\varepsilon$ away from $\mathcal{F}$ with probability at least $1-\delta$ if the analyst collects any sample size of $n\geq n(\varepsilon,\delta)$ observations.\footnote{The distance $\mathsf{erf}_\mu(\cdot,\cdot)$ is computed using the same distribution $\mu$ that governs the sampling of price–income pairs. Thus, both the measure of approximation error and the test are based on the same underlying data distribution.} Here $n(\varepsilon,\delta)$ is the minimal sample complexity, and it suffices to take $n=m_L(\varepsilon,\delta)$. In particular, it is worth noting that Theorem (ref) allows to study power of rational preferences against an alternative Lipschitz demand. Indeed, one can simply take $\mathcal{F}$ to be the set of all choice functions that arise from utility maximisation, and the RP characterization for this class reduces to the standard Afriat inequalities. The geometric structure of the argument in Theorem (ref) is conveyed in Figure (ref).
There, the set of $\mathcal{G}$-rationalising demands is represented by the neighbourhood $\mathcal{N}_{\varepsilon}(g)$. When $n \ge n(\varepsilon,\delta)$, with probability at least $1-\delta$, every $\mathcal{G}$-rationaliser of the observed dataset lies in $\mathcal{N}_{\varepsilon}(g)$. Under the alternative hypothesis, the true model satisfies $g\notin \mathcal{F}(\varepsilon)$, equivalently $\inf_{f\in\mathcal{F}}\mathsf{erf}_\mu(g,f)>\varepsilon$, so $\mathcal{N}_{\varepsilon}(g)\cap\mathcal{F}=\emptyset$. Consequently, with probability at least $1-\delta$, no element of $\mathcal{F}$ can rationalise the observed dataset. Since an RP characterisation rejects precisely when no element of $\mathcal{F}$ rationalises the data, it follows that the test rejects under the alternative with probability at least $1-\delta$.
While Theorem (ref) provided the forward bound, the next result states the corresponding inverse bound: given a fixed sample size $n$ and tolerance $\varepsilon$, it determines the confidence level $1-\delta(\varepsilon,n)$ that can be guaranteed. The difference is only in which parameters are treated as primitive, so the proofs mirror each other. To avoid trivialities, for fixed $\varepsilon>0$ and $n\in\mathbb{N}$ we assume that there exists some $\delta\in(0,1)$ such that $n(\varepsilon,\delta)\le n$.
Algebraically, Theorem (ref) amounts to inverting the sample--complexity map: instead of choosing $n$ large enough to guarantee a target confidence level $1-\delta$, we fix $(\varepsilon,n)$ and identify the best confidence level compatible with $n$, namely $1-\delta(\varepsilon,n)$, where $\delta(\varepsilon,n)$ is the smallest $\delta\in(0,1)$ such that $n(\varepsilon,\delta)\le n$. Geometrically, the idea is the same as in Figure (ref): with probability at least $1-\delta(\varepsilon,n)$, all $\mathcal{G}$-rationalisers of the dataset lie in the $\varepsilon$--ball around the true model, and if this ball is disjoint from $\mathcal{F}$ the test rejects.
Having stated the forward bound (fix $(\varepsilon,\delta)$, solve for $n$) and its inverse (fix $(\varepsilon,n)$, solve for $\delta(\varepsilon,n)$), it is natural to ask the complementary question: given a sample size $n$ and confidence level $1-\delta$, what separation $\varepsilon$ from $\mathcal{F}$ can the test reliably detect? The next result answers this by introducing the minimal detectable separation scale $\varepsilon(n,\delta)$. To avoid trivialities, we fix $n\in\mathbb{N}$ and $\delta\in(0,1)$ such that the set $\{\varepsilon>0:\ n(\varepsilon,\delta)\le n\}$ is nonempty.
The argument mirrors the proofs of Theorems (ref) and (ref) and relies on the same geometry as in Figure (ref). The key difference is that, for fixed $(n,\delta)$, the boundary tolerance $\varepsilon(n,\delta)$ is defined as an infimum and need not be attained. Accordingly, the result is stated with an arbitrary slack $\gamma>0$.
Conceptually, Theorem (ref) turns the sample--complexity map on its head: for a given sample size $n$ and confidence level $1-\delta$, it identifies a detectability frontier in terms of separation from $\mathcal{F}$. Any alternative whose distance from $\mathcal{F}$ exceeds $\varepsilon(n,\delta)$ by a strictly positive margin $\gamma$ is guaranteed to be detected with probability at least $1-\delta$ by any RP characterisation of $\mathcal{F}$. Thus, $\varepsilon(n,\delta)$ summarises the finite-sample resolution of the frequentist revealed-preference test: larger samples push the frontier down, while higher confidence shifts it up.
In what follows, we study the finite-sample power of revealed preference tests for static utility maximization (SUM), homotheticity (H), weak separability (WS), and expected utility (EU) against smooth violations of integrability. We consider a setup with four goods and, in the EU case, two equally probable states of the world. For the tests, we follow varian1983non and green1986. For weak separability, we only use necessary conditions since exact tests based on mixed-integer programming are NP-complete Cherchyeinteger2015.
We consider simulated demands given by $y(p, I) = x(p,I) + \varepsilon f(p,I)$, where \(x(p,I)\) are Cobb-Douglas demands, \(\varepsilon>0\) controls the deviation from rationality, and \(f(p,I)\) is a perturbation that preserves homogeneity of degree zero and budget neutrality, but violates Slutsky symmetry.\footnote{Although the perturbed demands may be negative, revealed preference theory still applies deb2023, so they are inconsequential for our simulations.} We normalize \(f\) such that the Frobenius norm of the skew-symmetric part of the Slutsky matrix equals one, ensuring comparability of \(\varepsilon\) across alternative hypotheses. See Appendix (ref) for full details.
We simulate 500 perturbed demands from Cobb-Douglas preferences drawn uniformly from the unit interval and scaled such as to sum up to one. Prices and income are drawn uniformly from $(0.01,10.00)$ and $(1.00,6.00)$, respectively. This level of price variation exceeds what is typically observed in empirical datasets. As a result, the RP tests are given favorable conditions for detecting deviations from rationality. For each value of $\epsilon \in [0.02, 0.03, \dots, 0.25]$, we compute the smallest sample size required to achieve a rejection probability of at least $90\%$. The resulting power curves are reported in Figure (ref).\footnote{We obtain qualitatively similar power curves with an alternative data generating process featuring a different pattern of deviations from Slutsky symmetry, indicating that the relative ranking of tests is robust across alternative hypotheses.}
The figure shows that for small deviations from rationality, SUM and WS require large sample sizes to achieve the target rejection rate. In contrast, H and EU only require modest sample sizes to achieve the target rejection rate. For large deviations from rationality, every RP test achieves the target rejection rate with small to moderate sample sizes.
Next, we summarize the relationship between the distance from rationality and the sample size needed to achieve 90% power using a log-log regression:
where $\overline{n}^m$ is the smallest sample size needed to achieve $90\%$ power for model $m \in \{\text{SUM}, \text{H}, \text{WS}, \text{EU}\}$, $\epsilon^{m}$ is the distance from rationality, and $\omega^{m}$ captures residual variation due to the simulation process. The intercepts ($\beta_0$) reflect baseline differences in the sample size needed to reach a given power level when the deviation from rationality is one. The slopes ($\beta_1$) reflect the elasticity of the required sample size with respect to deviations from rationality. Notice that by Lemma (ref), fixing $\delta$ to 0.1, the order of $\beta_1$ should be smaller than $2+l$. We find this empirically, and the results are presented in Table (ref).\footnote{Standard errors are omitted because the estimates are derived from Monte Carlo simulations rather than empirical sampling. We show two further regressions in Appendix (ref) for completeness.}
Table (ref) shows that the required sample size decreases progressively from SUM to H, then WS, and finally EU. The table also shows that SUM and EU have higher sensitivity of required sample sizes to deviations from rationality than H and WS. Since sample size is modeled on a log scale, EU has a steeper slope ($\beta_1$) in terms of relative changes even though its curve appears flatter in absolute sample size due to a lower baseline sample size ($\beta_0$) compared with WS.
The previous section applied our framework to analyze the power of nonparametric revealed-preference tests. This section extends the analysis to tests based on functional restrictions, which may increase power or prove useful when nonparametric characterisations are difficult to implement.
To develop functional tests, we first introduce functional restrictions that single out subclasses of preferences.
Suppose we want to describe a hypothesis test of $\mathcal{F}$ as in the preceding section: \[ \mathbf{H}_0: g \in \mathcal{F} \qquad\text{vs.}\qquad \mathbf{H}_1: g \in \mathcal{G}\setminus \mathcal{F}(\varepsilon), \] where \[ \mathcal{F}(\varepsilon) := \{ g \in \mathcal{G} : \inf_{g'\in\mathcal{F}}\mathsf{erf}_\mu(g,g') \leq \varepsilon \}, \]
as in Theorem (ref). As the examples below show, the definition of functional restriction imposes no finite-sample bounds on size and power without additional restrictions.
As the above examples illustrate, functional restrictions alone may be too weak unless they satisfy additional properties that control their behavior under small perturbations of the data. Intuitively, we want two properties: (i) if the true preference lies in $\mathcal{F}$, then any nearby preference should nearly satisfy the restriction (continuity), and (ii) if a preference is far from satisfying the restriction, then this separation should be detectable in terms of our error metric (regularity). We now formalise these requirements.
\
Putting everything together, a functional restriction is simply a condition $\mathcal{R}$ that defines a subclass $\mathcal{F}\subset\mathcal{G}$ and induces a natural hypothesis test. To make such tests well behaved in our setting, we impose uniform continuity and regularity. The former guarantees the stability of the restriction under small perturbations, while the latter guarantees detectability of violations with respect to our error metric.
The preceding subsection established the role of functional restrictions in defining well–behaved subclasses $\mathcal{F}\subset \mathcal{G}$ and showed how uniform continuity and regularity provide the discipline needed for asymptotic analysis. We now turn to the construction of an explicit testing algorithm. The idea is to exploit the functional restriction $\mathcal{R}$ to obtain a computable statistic that separates the null $\mathbf{H}_0:g\in\mathcal{F}$ from the alternative $\mathbf{H}_1:g\in \mathcal{G}\setminus \mathcal{F}(\varepsilon)$.
The algorithm below makes concrete the approach developed in Section (ref). Rather than treating revealed preference tests, it delivers a constructive procedure with explicit finite–sample power bounds. It also illustrates how functional restrictions can serve as a practical testing device in settings where algebraic characterisations of the null are either unknown or computationally prohibitive.
The algorithm operationalises functional restrictions by providing a statistic $\mathcal{R}(x)$ capable of distinguishing the null from the alternative. Step 2 ensures that the sample size is sufficiently large relative to the tolerance parameters so that the separation implied by regularity dominates the continuity margin. Step 3 then reduces the test to a simple decision rule based on whether the restriction is violated beyond the uniform–continuity threshold. In this way, the procedure turns the abstract properties of functional restrictions into a concrete finite–sample test. The next result uses this algorithm to construct a test based on functional restrictions.
Theorem (ref) follows directly from the properties of uniform continuity and regularity. Under the null, Algorithm 1 ensures that any rationalising preference $\succeq_n$ lies within $\varepsilon/t$ of the true preference with probability at least $1-\delta$. Uniform continuity then implies $\mathcal{R}(x_{n})\leq\gamma(\varepsilon/t)$, so rejection occurs with probability at most $\delta$. Under the alternative, the true preference is separated from $\mathcal{F}$ by at least $\varepsilon$, and regularity guarantees $\mathcal{R}(g)\geq\lambda(\varepsilon)$. With probability at least $1-\delta$, the rationalising preference $\succeq_n$ selected by the analyst lies within $\varepsilon/t$ of the true preference, so $\mathcal{R}(x_{\succeq_n})\geq \lambda(\varepsilon)-\gamma(\varepsilon/t)$, which exceeds the critical value by construction. Hence, rejection occurs with probability at least $1-\delta$. Taken together, Algorithm 1 and Theorem (ref) show how functional restrictions yield implementable frequentist tests with explicit finite–sample bounds.
This section outlines how to construct confidence intervals for smooth estimands of demand. To fix ideas, let $\mathcal{R}$ denote a functional restriction that maps a demand $x_\succeq$ into a real number summarising the change in a cost–of–living index induced by the preference $\succeq$. The analyst’s goal is to use the observed data to estimate the true welfare value $\Delta W^*=\mathcal{R}(g)$, where $g$ is the demand generated by the true underlying preference. The following result shows that our methodology allows one to obtain confidence intervals on welfare estimates.
Theorem (ref) shows that when $\mathcal{R}$ is uniformly continuous, $x$ is ensured to lie within $\varepsilon$ of $g$, such that the values of the functional differ by at most $\varepsilon$. With probability at least $1-\delta$, the true value of the estimand lies within $\gamma(\varepsilon)$ of the computed $\Delta W$. The attainable precision depends only on the modulus of continuity of $\mathcal{R}$ and on the PAC sample complexity $n(\cdot,\cdot)$ of the class $\mathcal{G}$.
Recall from Section (ref) that \(\mathcal{P}\times\mathcal{I}\) is compact and \(\mathcal{G}\) denotes the class of income-Lipschitz demands. In this section we restrict attention to the subclass \(\mathcal{G}^{R}\subseteq\mathcal{G}\) of (rational) models whose Marshallian demand \(x:\mathcal{P}\times\mathcal{I}\to Y\) is twice continuously differentiable in \((p,I)\). This additional regularity is used only to justify the derivative-based restriction operators \(\mathcal{R}\). The finite-sample size and power bounds from Sections (ref) are otherwise unchanged.
Throughout this section, we write $\mathsf{D}_p x(p,I)\in\mathbb{R}^{\ell\times\ell}$ for the Jacobian with respect to prices, with entries $(\mathsf{D}_p x)_{ij}(p,I)=\partial x_i(p,I)/\partial p_j$, and we write $\partial_I x(p,I)\in\mathbb{R}^{\ell}$ and $\partial_I^2 x(p,I)\in\mathbb{R}^{\ell}$ for the first and second derivatives with respect to income.\footnote{Whenever the dependence on $(p,I)$ is clear, we omit the arguments.} We also use the Slutsky matrix $S(x)(p,I):=\mathsf{D}_p x(p,I)+\big(\partial_I x(p,I)\big)x(p,I)^\top$, with entries $S_{ij}(x)(p,I)=(\mathsf{D}_px)_{ij}(p,I)+x_j(p,I)\cdot\partial_I x_i(p,I)$.
Homotheticity can be tested via the linearity of Engel curves. Under our assumptions, homothetic preferences \(\succeq\) are equivalent to linearity of Marshallian demand in income at every fixed price. Equivalently, the second derivative of demand with respect to income vanishes. This yields a simple derivative-based restriction.
Homotheticity admits a transparent derivative restriction: \(\mathcal{R}^{\mathrm{hom}}(x_\succeq):=\partial_I^2 x_\succeq\) must vanish under the null. Since \(\mathcal{R}^{\mathrm{hom}}\) is uniformly continuous (Proposition (ref)), it fits directly into our testing recipe: estimate \(\mathcal{R}^{\mathrm{hom}}\) from data on \((p,I,x)\), form the empirical score, and apply Algorithm 1 to obtain finite-sample size control and power against non-homothetic alternatives.
We now turn to weakly separable preferences across a given partition of goods. The revealed preference tests for this class are NP–hard, which motivates our functional approach. For simplicity, we define weak separability directly via a utility representation, as axiomatic characterisations in terms of preferences are well known (debreu60-2, Koopmans1972).
For a \(C^2\) demand \(x:\mathcal{P}\times\mathcal{I}\to Y\), the Slutsky matrix is $\mathcal{S}(x)\ :=\ \mathsf{D}_p x\ +\ x\,(\partial_I x)^\top$, with entries \(\mathcal{S}_{ij}(x)=\partial x_i/\partial p_j + x_j\,\partial x_i/\partial I\). Before stating the main result, we introduce some notation. For $(p,I)$, write the between–group Slutsky block \[ \mathcal{S}^{12}(x_{\succeq})(p,I) := \big(\mathcal{S}_{ij}(x_{\succeq})(p,I)\big)_{i\in G_1,\,j\in G_2} \in \mathbb{R}^{|G_1|\times|G_2|}, \] and the income–effect vectors \[ a(p,I):=\big(\partial_I x_{\succeq,i}(p,I)\big)_{i\in G_1}, \qquad b(p,I):=\big(\partial_I x_{\succeq,j}(p,I)\big)_{j\in G_2}. \] We also impose a two–sided income–effect nondegeneracy condition: for every $(p,I)$ there exist $i\in G_1$ and $j\in G_2$ with $\partial_I x_{\succeq,i}(p,I)\neq 0$ and $\partial_I x_{\succeq,j}(p,I)\neq 0$.
The proposition shows that, under the nondegeneracy condition on income effects, weak separability across $(G_1, G_2)$ is equivalent to the vanishing of the functional restriction $\mathcal{R}^{\mathrm{sep}}$. In other words, the structural property of weak separability is captured exactly by the condition that the between–group Slutsky block has rank one, with income effects determining the factorisation.
To see this more concretely, consider the case of three goods with $G_1=\{1\}$ and $G_2=\{2,3\}$. In this case, the between–group Slutsky block reduces to the row \[ \mathcal{S}^{12}(x_{\succeq})(p,I) = \big[\,\mathcal{S}_{12}(x_{\succeq})(p,I)\ \ \mathcal{S}_{13}(x_{\succeq})(p,I)\,\big]. \] The restriction operator becomes the scalar condition \[ \mathcal{R}^{\mathrm{sep}}(x_{\succeq})(p,I) = \mathcal{S}_{12}(x_{\succeq})(p,I)\,\partial_I x_{\succeq,3}(p,I) - \mathcal{S}_{13}(x_{\succeq})(p,I)\,\partial_I x_{\succeq,2}(p,I). \] Hence, weak separability across $(\{1\},\{2,3\})$ holds if and only if $\mathcal{R}^{\mathrm{sep}}(x_{\succeq})(p,I)=0$ for all $(p,I)$. Where both income effects $\partial_I x_{\succeq,2}$ and $\partial_I x_{\succeq,3}$ are nonzero, this condition is equivalent to \[ \frac{\mathcal{S}_{12}(x_{\succeq})(p,I)}{\mathcal{S}_{13}(x_{\succeq})(p,I)} = \frac{\partial_I x_{\succeq,2}(p,I)}{\partial_I x_{\succeq,3}(p,I)}. \] If one of the two income effects vanishes, the restriction reduces correspondingly to $\mathcal{S}_{12}=0$ or $\mathcal{S}_{13}=0$. This simplified case makes transparent the role of $\mathcal{R}^{\mathrm{sep}}$: it records the precise alignment between substitution effects and income effects across groups that characterises weak separability.
The notions of complementarity and substitutability have been formalized in several ways in the literature.\footnote{Early definitions based on the signs of second derivatives of utility functions (see auspitz2015untersuchungen, fisher1925mathematical, edgeworth1897teoria, pareto1909manuel) are not invariant to monotone transformations. For surveys, see samuelson1974complementarity and newman1987substitutes.} Two demand–based definitions are most prominent. First, goods $i$ and $j$ are said to be gross complements (substitutes) whenever \[ \frac{\partial x_i(p,I)}{\partial p_j} \;<\; 0 \quad \Big(\,>\,0\;\Big). \] Because this derivative includes income effects, the notion is not symmetric. Second, following hicksallen1934, goods $i$ and $j$ are net complements (substitutes) if the corresponding Hicksian demand satisfies \[ \frac{\partial h_i(p,u)}{\partial p_j} \;<\; 0 \quad \Big(\,>\,0\;\Big). \] By the Slutsky decomposition, this is equivalent to the sign of the Slutsky term \[ \mathcal{S}_{ij}(x)(p,I) = \frac{\partial x_i(p,I)}{\partial p_j} + x_j(p,I)\,\frac{\partial x_i(p,I)}{\partial I}. \] Unlike the gross notion, this condition is symmetric across goods.\footnote{For a recent approach that restores intuitive appeal to these definitions, see weinstein2022direct.}
Since these functional restrictions are constructed from derivatives of uniformly continuous functions, they inherit uniform continuity away from zero. Hence, as in the case of weak separability, they provide well–behaved finite–sample tests of strict complementarity and strict substitutability within our framework.
In this section, we adapt the functional form approach of kubler2014asset, who provide a characterization of expected utility preferences from contingent-claim demand.\footnote{See Theorem 3 in kubler2014asset.} Our objective is to extend this analysis to the broader betweenness class of preferences introduced by dekel1986axiomatic and chew1989axiomatic. Betweenness preferences generalise expected utility by requiring indifference curves to be straight lines, though not necessarily parallel, and admit an implicit representation of utility.
Formally, let there be $S$ states of nature, indexed by $s\in\{1,\ldots,S\}$. The decision maker has preferences over state–contingent consumption vectors $\mathbf{x}\in\mathbb{R}_{>0}^S$ and holds objective beliefs $\boldsymbol{\pi}=(\pi_1,\ldots,\pi_S)\in\Delta(S)$. Preferences are represented by a strictly increasing, thrice continuously differentiable, and strictly quasi–concave utility function $V(\mathbf{x};\boldsymbol{\pi})$. By dekel1986axiomatic, a preference belongs to the betweenness class if and only if there exists a Bernoulli index $u:\mathbb{R}_{>0}\times[0,1]\to\mathbb{R}$, increasing in the first argument and continuous in the second, such that
We assume throughout that $u$ is thrice continuously differentiable in both arguments, concave in consumption, and satisfies $u_{11}<0$.\footnote{Here and below, subscripts denote partial derivatives of the utility index $u$: for example $u_1(x,v)=\partial u(x,v)/\partial x$ and $u_{11}(x,v)=\partial^2 u(x,v)/\partial x^2$.}
As in kubler2014asset, we assume complete markets with strictly positive prices $\mathbf{p}\in\mathbb{R}_{>0}^S$ and income $I>0$. The Marshallian demand induced by a betweenness preference $\succeq$ is then defined by
For each fixed $s\geq 3$ define the map \[ H_s(\mathbf{p},I;\boldsymbol{\pi}):=\Big(x_{\succeq,1}(\mathbf{p},I,\boldsymbol{\pi}),\,x_{\succeq,2}(\mathbf{p},I,\boldsymbol{\pi}),\, k_s(\mathsf{p},\boldsymbol{\pi},\, x_{\succeq,s}(\mathbf{p},I,\boldsymbol{\pi})) \Big)\in \mathbb{R}^4, \] where $k_s=(\pi_s/\pi_1)(p_1/p_s)$ and $\boldsymbol{\pi}$ is treated as a parameter. Let $J_s(\mathbf{p},I)$ be the $4\times (S+1)$ Jacobian of $H_s$ with respect to $(\mathbf{p},I)$.
Proposition (ref) imposes a clear testable implication. Whenever we perturb prices and income so that, to first order, $x_1$, $x_2$, and the ratio $k_s=(\pi_s/\pi_1)(p_1/p_s)$ remain unchanged, then the demand for state $s$ must also remain unchanged. Equivalently, the only channels through which $x_s$ is allowed to move are movements in $x_1$, $x_2$, or $k_s$. There is no independent direction in $(\mathbf{p},I)$ that can change $x_s$ while holding these other quantities fixed. This is the differential analogue of the functional dependence in Proposition (ref): under betweenness, once $(x_1,x_2,k_s)$ are fixed, $x_s$ is pinned down.
The key difference with expected utility is that, under betweenness, $u_1(\cdot,V)$ depends jointly on consumption and the endogenous fixed point $V$. In expected utility, $u_1(\cdot)$ is univariate, which lets the FOCs collapse to a representation $x_s=f(x_1,k_s)$ that characterises the model. In contrast, for betweenness the FOCs imply that $x_s$ is a function of $(x_1,x_2,k_2,k_s)$, but this does not guarantee that a given $f$ arises from some $u(\cdot,\cdot)$ consistent with the fixed–point structure. Thus, necessity holds---any betweenness rationalisation must satisfy the rank condition--- but sufficiency need not hold in general, as not every solution to the rank condition corresponds to some betweenness utility.
The corollary makes precise the one–sided nature of the restriction: any demand violating $\mathcal{R}^{\mathrm{bet}}_{s}\equiv 0$ cannot come from betweenness, but some non–betweenness demands may also satisfy it.
Although $\mathcal{R}^{\mathrm{bet}}$ is not a full characterisation, it yields a valid non–rejection test. We test \[ \mathbf{H}_0:\ x_{\succeq}\in\mathcal{F}_{\mathrm{bet}} \qquad\text{vs}\qquad \mathbf{H}_1:\ x_{\succeq}\in\mathcal{G}\setminus \mathcal{F}_{\mathrm{bet}}(\varepsilon), \] where $\mathcal{F}_{\mathrm{bet}}(\varepsilon)$ is defined as in Section (ref). For each $s\ge 3$ and each evaluation point $(\mathbf{p},I)$, compute the Jacobian $J_s(\mathbf{p},I)$ of \[ H_s(\mathbf{p},I;\boldsymbol{\pi}) =\big(x_{\succeq,1},\,x_{\succeq,2},\,k_s,\,x_{\succeq,s}\big), \qquad k_s=(\pi_s/\pi_1)(p_1/p_s), \] with respect to $(\mathbf{p},I)$. Form \[ \mathcal{R}^{\mathrm{bet}}_{s}(x_{\succeq})(\mathbf{p},I,\boldsymbol{\pi}) =\sum_{M\in\mathscr{M}_{4\times 4}(J_s(\mathbf{p},I))}(\det M)^2, \] and aggregate across $s$ and $(\mathbf{p},I)$ via the $L^1_\mu$–integral of Section (ref), obtaining the sample analogue $\widehat{\mathcal{R}}^{\mathrm{bet}}$. Under $\mathbf{H}_0$, the restriction equals zero. By Proposition (ref) and compactness, the map $x\mapsto \mathcal{R}^{\mathrm{bet}}$ is uniformly continuous into $L^1_\mu$, so the finite–sample bounds established earlier apply.
Operationally, one can approximate the minors of $J_s$ by finite differences, using small perturbations $(\Delta\mathbf{p},\Delta I)$ chosen so that $x_1$, $x_2$, and $k_s$ are (approximately) held constant to first order. The restriction requires the induced change in $x_s$ to be (approximately) zero. Equivalently, the estimated $4\times 4$ minors evaluated at these perturbations should be close to zero under $\mathbf{H}_0$.
The sample complexity of both testing and estimation in our framework is inherited from the PAC--learnability of $\mathcal{G}$. This means that the attainable precision and confidence of welfare estimates are directly tied to the sample complexity of the underlying learning algorithm. In this section, we formalise this dependence and show how efficiency gains can be realised in two ways: by imposing additional structure on the class $\mathcal{G}$, and by allowing the analyst to select prices adaptively. The following proposition makes precise the link between improved sample complexity and tighter confidence intervals.
Proposition (ref) formalises the connection between sample complexity and inferential precision in our framework. The sample complexity $n(\varepsilon,\delta)$ captures how demanding it is to learn the class $\mathcal{G}$: smaller values mean that fewer observations are required to reach a given accuracy and confidence. Thus, the results states that improvements in sample complexity translate directly into tighter confidence bounds at the same $(n,\delta)$.
This paper develops a frequentist framework for conducting power analysis of revealed preference tests using finite choice data and tackles a central challenge in the literature: constructing parsimonious alternative hypotheses against which the discriminatory power of rationality tests can be assessed. Our procedure is grounded in the Probably Approximately Correct (PAC) learning framework and allows us to design tests that integrate functional and finite-data approaches. This method is broadly applicable, ranging from weak separability to choice under risk, highlighting its versatility for analyzing decision-making.
We make four main contributions. First, we show that PAC learnability results for Lipschitz demand functions provide the theoretical foundation for constructing well-defined alternative hypotheses. By exploiting the fact that any rationalizing demand function lies within a learnable neighborhood of the true demand with high probability, we can compute explicit bounds on the sample size required to achieve any desired level of statistical power against alternatives that are $\varepsilon$-separated from the null hypothesis.
Second, we extend our framework beyond traditional revealed preference characterizations to functional restrictions. This extension is particularly valuable for preference classes—such as weakly separable preferences—where exact RP tests are computationally intractable (NP-hard). We show that by tolerating a small but controlled size, one can construct polynomial-time tests with explicit finite-sample power guarantees.
Third, we demonstrate how to construct confidence intervals for smooth functionals of demand, such as equivalent variation and other welfare measures. The width of these intervals depends directly on the sample complexity of learning the underlying preference class, thereby linking the richness of the maintained hypothesis to the precision of welfare inference.
Fourth, our simulation results reveal an important insight: the apparent empirical success of GARP may partly reflect its limited ability to detect small deviations from rationality. In contrast, the frequent rejections of expected utility in experimental data may reflect its sensitivity to even negligible departures from the model. This suggests that comparing rejection rates across different preference classes requires careful attention to their differential statistical power.
Our approach shifts the focus from testing the dataset itself to testing the decision maker. Unlike RP tests that always reject a dataset if it cannot be rationalized by any preference from the class of interest, our tests are one-sided: they never reject a rationalizable dataset. Still, they may fail to reject a non-rationalizable one. However, the probability of failing to reject a non-rationalizable dataset decreases as more data are sampled. Our simulation exercises show that our tests perform well in finite samples and correctly reject demands that do not belong to specific classes of preferences.