EconBase
← Back to paper

Combining stated and revealed preferences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,334 characters · 17 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Combining stated and revealed preferences

abstractCan stated preferences inform counterfactual analyses of actual choice? This research proposes a novel approach to researchers who have access to both stated choices in hypothetical scenarios and actual choices, matched or unmatched. The key idea is to use stated choices to identify the distribution of individual unobserved heterogeneity. If this unobserved heterogeneity is the source of endogeneity, the researcher can correct for its influence in a demand function estimation using actual choices and recover causal effects. Bounds on causal effects are derived in the case, where stated choice and actual choices are observed in unmatched data sets. These data combination bounds are of independent interest. We apply numerical delta method to inference for the bounds and show its good performance in a simulation experiment.

Keywords: Stated and revealed preferences, data combination, optimal transport .

JEL codes: C19, C35, D19, D84.

Introduction

Social scientists use two different ways to learn about human behaviour: One possibility is to observe what people do or choose in real life, which economists call the revealed preference approach. This approach has the advantage of reflecting actual choices. Its main limitation is that choice relevant states and observed choices are driven by common unobserved underlying factors. Researchers have relied on randomised controlled trials or natural experiments to perform causal analysis. However, randomisation is not always feasible, for technical, political or ethical reasons, and when it is feasible, it is often plagued by imperfect compliance. The other possibility to learn about human behaviour is to ask people what they would do or choose in a hypothetical situation, in what is often called the stated preference approach. The main advantages of this approach are that the analyst can design experiments with rich variations in choice attributes, observe individual stated choice in counterfactual scenarios and explore preferences over policies never implemented before. Moreover, the analyst can exploit hypothetical choice scenarios to tackle endogeneity arising from omitted unobserved characteristics and identify causal effects of choice attributes. Indeed, even if the relevant choice characteristics are not all observed, exogenous variations on choice characteristics manipulated in a survey experiment are sufficient for causal inference. The obvious disadvantage of the stated preference approach is summarised by train2009: `What people say they will do is often not the same as what they actually do'. This discrepancy is known as the hypothetical bias.

Unlike other social sciences, economics has historically seen the hypothetical bias as a disqualifying argument against the stated preference approach, and until recently, economists have, by and large, relied solely on revealed preference data, except in the measurement of the non-use value of natural or cultural resources. This attitude is changing, as observed and advocated by Orazio Attanasio's 2024 Econometric Society presidential address, published in almaas2024.\footnote{The turn of the tide coincides with a surge of interest in new measures capturing individuals' subjective expectations as documented by manski2004 and almaas2024. See also mcfadden2017 for a historical perspective on stated preference data.} Stated preference analyses are increasingly used to describe individual preferences over choice attributes that are endogenous, hard to measure in observational data or hard to vary in randomized control trials. An early example is juster1964. Recent applications span many areas of economics, including education choices arcidiacono2020,delavande2019, mobility decisions gong2022, kocsar2022, meango2022, batista2025, health and long-term care investments kesternich2013, ameriks2020b,boyer2020, parental investments attanasio2019,almaas2024, marriage preference adams2019, low2024, occupational choices maestas2023, wiswall2015, wiswall2018,meango2024, and retirement decisions ameriks2020a, giustinelli2024. For recent reviews, see kocsar2023 and giustinelli2023. Despite the enthusiasm and encouraging evidence that carefully designed stated preference elicitation has predictive power for actual choices hurd2009, hainmueller2015, debresser2019, arcidiacono2020 and yields similar preference estimates as actual choice data mas2017, wiswall2018, concerns remain regarding systematic biases of stated preference data. See, for example, murphy2005 and hausman2012 and more recently haghani2021a,haghani2021b.

An alternative approach consists in combining stated preference data with revealed preference data, in order to mitigate their respective shortcomings (hypothetical bias on the one hand and endogenous choice attributes on the other). However, research in this direction has so far ignored the fact that revealed and stated preferences generally occur in different data sets, in which individuals can rarely be matched. In addition, research in this direction has so far relied on very strong structural assumptions on the relation between stated and revealed preference. morikawa2002, pantano2013, and giustinelli2024 restrict the difference between utility parameters that govern stated and revealed preferences; vanderklaauw2012 and wiswall2021 assume that any bias between statement and action comes solely from biased information; and bernheim2022 assume that the treatment affects actual choice only through the stated preference, so that the difference between stated and revealed preferences is stable across treatment groups. briggs2020 assume additively separable and scalar unobserved heterogeneity, a specific timing of the resolution of uncertainty, and rational expectations in their identification of marginal treatment effects from subjective expectations.\footnote{athey2025combining apply similar ideas to the combination of observational and experimental data.}

This paper proposes to combine revealed preference data with stated preference data, matched or unmatched, to analyze binary choice with endogenous choice attributes, without strong assumptions on the way the two are related. We propose a strategy to retrieve information about individual unobserved heterogeneity from stated preferences and we derive a new result on data combination to apply this strategy in the case, where revealed and stated preferences are observed in distinct, unmatched data sets. The fundamental idea is that stated preferences are useful, not because they necessarily match actual choices, but insofar as they provide valuable information on individual heterogeneity. Suppose that an analyst asks a job-seeker about their probability to take up a particular job. The probability is elicited for different scenarios varying wage, job security, and non-wage compensation. A respondent who reports a 90 percent chance of taking up the job in all scenarios is arguably different from a respondent who states a 90 percent chance if the wage is high, and a 10 percent chance otherwise. The first respondent might have a higher taste for work, lower disposable income or lower ability than the second respondent. Irrespective of the reason, and even if `what they say is often not what they actually do', their responses reveal important information about how they differ in their preferences about job attributes. This information helps to classify individuals into unobserved heterogeneity types. The dimension of unobserved heterogeneity that can be recovered depends on the richness of the stated choice experiment. Repeated elicitation in different scenarios allows the recovery of multiple dimensions of unobserved heterogeneity.

In our framework, agents make a binary decision $D$ based on an endogenous decision relevant attribute $X$. The parameter of interest is $\mu(x):=\mathbb E[D(x)]$, where $D(x)$ is the potential decision when the decision relevant attribute is exogenously set to $x$. We derive conditions under which a vector of stated preferences reports $\mathbf P$ (typically stated choice probabilities) can be used to identify the unobserved heterogeneity that causes endogeneity. More precisely, we derive conditions under which the parameter of interest is identified as $\mu(x)=\int\mathbb E[D\vert X=x,\mathbf P=\mathbf p]\;dF_{\mathbf P}(p)$, where $F_{\mathbf P}$ is the probability distribution of stated preference reports $\mathbf P$. For the common case, when actual choices $D$ and stated choice probabilities $\mathbf P$ are not observed in the same data set, we derive bounds on $\mu(x)$ based on knowledge of the distributions of $D\vert X$ from revealed preferences and of $\mathbf P\vert X$ from stated preferences.

horowitz1995identification, cross2002regressions and molinari2006generalization derive bounds on the conditional expectation $\mathbb E[D\vert X,\mathbf P]$ based on the distributions of $D\vert X$ and $\mathbf P\vert X$, when $\mathbf P$ has finite support. Here, however, $\mathbf P$ may be discrete or continuous, and we need bounds on the integral of $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ over $\mathbf p$, which involves constraints across values of $\mathbf p$ beyond the simple bounds in horowitz1995identification. cross2002regressions derive bounds for the function $p\mapsto \mathbb E[D\vert X=x,P=p]$. Those bounds are functional, hence they do take into account restrictions across $p$. More precisely, they show that the identified set for $p\mapsto \mathbb E[D\vert X=x,P=p]$ is the convex hull of a set of $J\!$ extreme points (where $J$ is the cardinality of the support of $P$). We derive simple bounds on the integral of this function over $p$, i.e., the parameter described in section 4 of cross2002regressions. Our bounds lend themselves to straightforward inference, and they remain valid when $P$ is continuous (or mixed). As such, our bounds complement the work of fan2014identifying (fan2014identifying,fan2016estimation), who also generalize the bounds of cross2002regressions and allow for mixed discrete and continuous covariates and outcomes. They derive bounds for the expectation $\mathbb E[g(Y_d)]$, as well as the counterfactual distributions, quantile and distributional treatment effects, under unconfoundedness $(Y_0,Y_1)\perp\!\!\!\!\perp D\vert Z$, where the distributions of $((Y_1D+Y_0(1-D)),D)$ and $(Z,D)$ are identified from two separate data sets.

Defining $m(x,\mathbf p)=\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$, we derive the bounds from the dual of the optimal transport problems

eqnarray*[eqnarray* omitted — 162 chars of source]

where the supremum is over all joint distributions for $(D,\mathbf P)\vert X=x$ with fixed marginals $D\vert X=x$ and $\mathbf P\vert X=x$. The application of ideas related to the theory of optimal transport in data combination problems can be traced back to horowitz1995identification and heckman1997. More recently, specific aspects of optimal transport theory were used to derive bounds in specific data combination problems in fan2014identifying (fan2014identifying,fan2016estimation) and lee2019identification for treatments effects under data combination, pacini2019two, gaillac2024linear (gaillac2024linear,gaillac2025partially) for partially linear regression and best linear prediction under data combination. The most closely related to ours is fan2025partial, which applies the `method of conditioning' from ruschendorf1991bounds to derive bounds on the parameter of a moment equality model under data combination. ichimura2005identification and bontemps2025functional use a different approach, in that they seek conditions for point identification of parameters of moment equality models despite data combination.

We show how to tighten the bounds under exclusion restrictions and we use the bootstrap procedure proposed by fang2019inference to the bounds, as the latter are Hadamard directionally differentiable functions of two density functions and a conditional expectation. We investigate the tightness of the bounds and the performance of the inference procedure in a simulation experiment.

The rest of the paper is organized as follows. Section (ref) lays out the econometric framework and the main identification and partial identification results. Sections (ref) and (ref) describe the inference procedure and the simulation experiment respectively. T he last section concludes.

Econometric framework and identification

Model specification for revealed preferences

Consider an economic agent with a vector $X$ of economic attributes with support $\mathcal X\subseteq \mathbb R^{d_x}$, who is facing binary decision $D\in\{0,1\}$. We denote $D(x)$ the potential decision, i.e., the decision the agent would make if their attributes were externally set to $X=x$. Both the vector of attributes $X$ and the actual decision $D:=D(X)$ are observed. Potential decision $D(x)$ is observed only when $X=x$ and not otherwise. The attribute $X$ is endogenous in the sense that $D(x)$ is not independent of $X$, and hence $\mathbb E[D\vert X=x]$, which is identified from the data, is generally not equal to $\mathbb E[D(x)]$, which is the parameter of interest. Finally, let $\eta$ be a vector of unobserved characteristics. The support $\mathcal H\subseteq\mathbb R^d$ of $\eta$ is independent of $X$. We posit that $\eta$ contains the source of endogeneity of $X$ in the sense of the following assumption.

assumption$( D(x) )_{x\in\mathcal X} \perp\!\!\!\!\perp X \; \vert \; \eta$.

Assumption (ref) is related to the control function approach pioneered by heckman1985alternative. Under this assumption, the structural choice function $\mathbb E[D(x)\vert\eta]$ satisfies

eqnarray*[eqnarray* omitted — 139 chars of source]

so that recovering $\eta$ would allow identification of $m(x,\eta)$ and hence of the following counterfactual choice parameters.

definitionThe average structural choice function is defined for all $x\in\mathcal X$ by \begin{eqnarray*} \mu(x) \; := \; \mathbb E[D(x)] \; = \; \int m(x,\eta) \; dF_\eta. \end{eqnarray*}
example[Retirement decision] Our pilot example is a model of retirement decisions based on health status. The retirement decision is governed by the following model: \begin{eqnarray*} D(x) & = & g(x,\eta,\nu), \end{eqnarray*} where the choice attribute $x$ is health status, which we assume here to be in $\mathcal X=\{0,1\}$ for simplicity, $g$ is an unknown function, $\eta\in\mathbb R$, again for simplicity, is an unobserved trait that influences preferences, such as risk aversion or health consciousness, and $\nu$ is a shock observed by the agent at the time of decision, but not by the analyst. Assume $\nu\perp\!\!\!\!\perp X \; \vert \; \eta$ so that assumption (ref) holds.

Model specification for stated preferences

Stated preferences are data collected from a survey of the agents prior to the time of decision. Agents are asked to give an assessment of their probability of making decision $D=1$ for some hypothetical values of $x$. Let $(x_1,\ldots,x_T)$ be a vector of hypothetical values $x_t\in\mathcal X$. The agent's response to the probability elicitation question is denoted $P_t$, $t=1,\ldots,T$.

We model $P_t$, for each $t=1,\ldots,T$ as

eqnarray*[eqnarray* omitted — 57 chars of source]

where the function $\tilde m(x,\eta)$ is called the stated choice function. In the special case where agents have rational expectations and no reporting bias, $P_t=\mathbb E[D(x)\vert \eta]$, so that $\tilde m(x_t,\eta)=m(x_t,\eta)$ for $t=1,\ldots,T$. In this case, the identification issue boils down into recovering $ m(x,\eta)$ for $x \not \in (x_1,\ldots,x_T)$. In general, the stated preference function $\tilde m(x_t,\eta)$ may differ from the structural choice function $m(x_t,\eta)$ for revealed preferences, due to the agent's perception and/or reporting biases. Crucially, the agent's report $P_t$ is driven by the same vector $\eta$ of unobserved characteristics. This provides the only link in the model between stated and revealed preferences.

continued[Example (ref) continued:] In the retirement example, agents are asked some time before retirement age what they assess their probability of retiring to be, given hypothetical health statuses. The stated preference report $P_t$ is modeled as the expectation of an anticipated decision variable \begin{eqnarray*} \tilde D(x_t) & = & \tilde g(x_t,\eta,\nu), \end{eqnarray*} where, as in the model for revealed preferences, $\eta$ is observed by the agent but not the analyst. However, $\nu$ is interpreted as resolvable uncertainty. It is not known by the agent at the time of stated preference elicitation, but is revealed at the time of decision. The agent is assumed to form their responses $P_t$ to stated preference elicitation by integrating $\nu$, so that \begin{eqnarray*} P_t \; = \; \tilde m(x_t,\eta) \; = \; \int \tilde g(x_t,\eta,\nu) \; d\tilde F_{\nu\vert x_t,\eta}, \end{eqnarray*} where the function $\tilde g$ may be different from its counterpart $g$ in the revealed preference model, and the perceived distribution $\tilde P_{\nu\vert x_t,\eta}$ of resolvable uncertainty $\nu$ may be different from the true distribution in the population. Note that the model above is observationally equivalent to a model stipulating $\tilde D(x_t)=\tilde g(x_t,\eta,\tilde\nu)$, where $\tilde\nu$, possibly different from $\nu$, is open to multiple interpretations.

Identification of revealed preferences using stated preferences

Identification of the parameters of interest relies on the ability to control for unobserved heterogeneity $\eta$ by matching on stated preference reports $\mathbf P=(P_1,\ldots,P_T)$. For this, we need to be able to recover unobserved heterogeneity from the stated preference reports of an individual.

assumptionThere is a subset $(x_1,\ldots,x_d)$ of the vector of stated preference scenarios such that $\eta\mapsto\tilde m(\eta):=(\tilde m(x_1,\eta),\ldots,\tilde m(x_d,\eta))$ is continuous and one-to-one.

Assumption (ref) guarantees that a unique $\eta$ can be recovered from the system $\mathbf P: = (P_1, \ldots, P_d)= (\tilde m(x_1,\eta),\ldots,\tilde m(x_d,\eta)):= \tilde m(\mathbf x,\eta)$. An example often used in empirical applications is a multiplicatively separable log-odd model, of which the model of blass2010 is a special case. It corresponds to:

eqnarray[eqnarray omitted — 102 chars of source]

where $\Gamma:\mathcal H\rightarrow[0,1]^d$ is one-to-one and $v(\mathbf x)$ is a full rank $d\times d$ matrix. In this case, assumption (ref) holds with $\eta = v(\mathbf x)^{-1} \left({\Gamma^{-1}(\mathbf P) - r(\mathbf x)}\right)$. More generally, sufficient conditions for assumption (ref) are given in Theorem 2 of mas1979. In model ((ref)), $\mathbf P$ has the same dimension as $\eta$. This need not be the case for assumption (ref) to hold, as long as the dimension of $\mathbf P$ is at least as large as the dimension of $\eta$. All that is needed, is that a subvector of $\mathbf P$ identifies $\eta$. This subvector need not be known, as long as it exists. Similarly, the stated preference function $\tilde m$ and the true dimension of $\eta$ need not be known. Assumption (ref), combined with assumption (ref), ensures that $D(x)\perp\!\!\!\!\perp X\vert \mathbf P$, which drives proposition (ref) below.

Assumption (ref) allows for any kind of hypothetical bias, providing that responses to hypothetical scenarios are sufficiently different to distinguish individuals who respond differently to choice variables in actual choices. If two individuals react differently to health shocks, possibly because they have a different taste for work, then Assumption (ref) holds if these individuals also respond differently to the hypothetical questions, one giving answers that are much more sensitive to the scenario than the other, even if both wildly overstate their willingness to work. Assumption (ref) requires the type $\eta$ of individuals to be fixed between scenarios and between stated and revealed preferences, and two individuals with different types to be distinguishable from their survey responses. Assumption (ref) would be violated for instance if an individual with low taste for work were induced by social desirability bias to give the same responses to the survey as an individual with high taste for work would.

continued[Example (ref) continued:] In the retirement example, we assume that the unobserved heterogeneity factor $\eta$ is scalar. Suppose the agent perceives the relative utility of option $D=1$ (retiring at $65$) with health status $x_0$ to be $\tilde S(x_0,\eta)-\nu$. Then, the perceived choice function is $\tilde g(x,\eta,\nu)=1\{\tilde S(x,\eta)\geq\nu\}$. If $\nu\perp\!\!\!\!\perp \eta$, $\tilde S(x_0,\eta)$ is increasing in $\eta$, and $\nu$ is absolutely continuous, then $\tilde m(x_0,\eta)=\tilde F_{\nu}(S(x_0,\eta))$ is increasing in $\eta$ and assumption (ref) holds. The average structural function $\mu(x)$ is identified as follows: \begin{eqnarray*} \mu(x) \; = \; \int \mathbb E[D(x)\vert \eta] \; dF_\eta \; = \; \int \mathbb E[D\vert X=x,\eta] \; dF_\eta \; = \; \int \mathbb E[D\vert X=x,P_0] \; dF_{P_0}, \end{eqnarray*} where the first equality is by definition, the second equality holds by assumption (ref) and the third equality holds by assumption (ref). Note also that when $\eta\mapsto\tilde m(x_0,\eta)$ is increasing, it is identified and $\eta$ can be recovered as $\eta=\tilde m^{-1}(x_0,P_0)$.

As illustrated in the example, under assumption (ref), stated preferences allow us to control for unobserved heterogeneity in the regression of decision $D$ on the choice attributes $x$, and to identify the parameters of interest.

proposition[Identification of the ASF] Under assumptions (ref) and (ref), the average structural function is identified as $\mu(x) = \int \mathbb E[ D\vert X=x,\mathbf P=\mathbf p ] \; dF_{\mathbf P}$, where $\mathbf P=(P_1,\ldots,P_T)$.
proof[Proof of proposition (ref)] Under assumption (ref), $D(x)\perp X\vert \eta \Rightarrow D(x)\perp X\vert \mathbf P$. Then, we have \begin{eqnarray*} \mu(x) & = & \mathbb E[D(x)] \\ & = & \int \mathbb E[D(x)\vert \mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p) \\ & = & \int \mathbb E[D(x)\vert X=x, \mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p) \\ & = & \int \mathbb E[D\vert X=x, \mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p), \end{eqnarray*} which completes the proof.

Proposition (ref) provides a constructive identification result. Inference on $\mu(x)$ can be based on the sample average of any estimator of the conditional expectation $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$.

Bounds under data combination

In most empirical environments, stated preferences and revealed preferences are collected in different data sets, and it is difficult or impossible to match individuals from the two data sets. Hence, the conditional distribution $F_{D\vert X}$ of $D$ given $X$ is identified from the revealed preference data set, and the conditional distribution $F_{\mathbf P\vert X}$ of $\mathbf P$ given $X$ is identified from the stated preference data set. However, the conditional distribution $F_{D,\mathbf P\vert X}$ of $(D,\mathbf P)$ given $X$ is not point identified. It can take any value in the set $\mathcal M(F_{D\vert X=x},F_{\mathbf P\vert X=x})$, where $\mathcal M(F,F^\prime)$ is the set of joint distributions with marginals $F$ and $F^\prime$.

For each conditional distribution $F_{D,\mathbf P\vert X}$ for $(D,\mathbf P)$ given $X$, proposition (ref) identifies a unique value for the structural function $\mathbb E[D\vert X,\mathbf P]$ and for the average structural function $\mu(x)$. However, since there are multiple distributions $F_{D,\mathbf P\vert X}$ compatible with the marginals $F_{D\vert X}$ and $F_{\mathbf P\vert X}$, both $\mathbb E[D\vert X,\mathbf P]$ and $\mu(x)$ are partially identified.

Sharp bounds on the structural function

We first derive sharp bounds for the conditional expectation $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$. Under assumption (ref), there is a one-to-one mapping between $\mathbf P$ and $\eta$. Hence, we abuse notation and call $m(x,\mathbf p)=\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ the structural function. The sharp identified set for the structural function $m(x,\mathbf p)$ contains all the conditional expectations $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ compatible with marginal distributions $F_{D\vert X}$ and $F_{\mathbf P\vert X}$.

definitionFor each $x\in\mathcal X$, the sharp identified set for $\mathbf p\mapsto m(x,\mathbf p)$ is the set of $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ such that $(D,\mathbf P)\vert X=x$ has distribution in $\mathcal M(F_{D\vert X=x},F_{\mathbf P\vert X=x})$ and $\mathbf p\mapsto \mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ is continuous.

We first give a characterization of the identified set in the form of sharp bounds on $m(x,\mathbf p)$. By definition of $m(x,\mathbf p)$, we have $\mathbb E[D-m(X,\mathbf P)\vert X=x,\mathbf P=\mathbf p]=0,$ which is equivalent to

eqnarray[eqnarray omitted — 97 chars of source]

for all continuous and integrable functions $h:\mathbb R^T\rightarrow\mathbb R$. The existence of a distribution in $\mathcal M(F_{D\vert X},F_{\mathbf P\vert X})$ such that ((ref)) holds is equivalent to

eqnarray[eqnarray omitted — 191 chars of source]

where the optimization is over $F\in\mathcal M(F_{D\vert X=x},F_{\mathbf P\vert X=x})$, and $h$ is any continuous integrable function of $\mathbf p$. The bounds are obtained through an application of optimal transport duality to the optimal transport problems on the left and right-hand sides of ((ref)). This yields the following proposition.

proposition[Sharp bounds on the structural function] A continuous function $m(x,\mathbf p)$ is in the identified set if and only if for all continuous and integrable function $h$, we have: \begin{eqnarray*} &&\sup_{\varphi\in\mathbb R} \left\{\mathbb E[\min(0,h(\mathbf P)-\varphi)\vert X=x]+\varphi\mathbb E[D\vert X=x]\right\} \\ && \hskip50pt \leq \; \mathbb E[h(\mathbf P)m(X,\mathbf P)\vert X=x] \\ && \hskip100pt \leq \; \inf_{\varphi\in\mathbb R} \left\{\mathbb E[\max(0,h(\mathbf P)+\varphi)\vert X=x]-\varphi\mathbb E[D\vert X=x]\right\}. \end{eqnarray*}
proof[Proof of proposition (ref)] Fix $x\in\mathcal X$. The existence of a distribution in $\mathcal M(F_{D\vert X},F_{\mathbf P\vert X})$ such that ((ref)) holds is equivalent to ((ref)). Both left and right-hand-sides of ((ref)) are solutions to optimal transport programs. By Theorem 1.3 in villani2021topics, the left-hand inequality is equivalent to the dual expression: \begin{eqnarray} \sup_{\varphi,\psi}\;\mathbb E[\varphi(D)\vert X=x] + \mathbb E[\psi(\mathbf P)\vert X=x] & \leq & 0, \end{eqnarray} where the supremum if over all integrable functions $(\varphi,\psi)$ such that \begin{eqnarray} \forall (d,\mathbf p), \;\; \varphi(d)+\psi(\mathbf p) \; \leq \; h(\mathbf p)(d-m(x,\mathbf p)). \end{eqnarray} Since $D\in\{0,1\}$, calling $\varphi:=\varphi(1)$, the constraint ((ref)) yields \begin{eqnarray*} \psi(\mathbf p) & = & -h(\mathbf p)m(x,\mathbf p)+\min \left( 0,h(\mathbf p)-\varphi \right). \end{eqnarray*} The latter can be plugged back into ((ref)) to yield the bound: \begin{eqnarray*} \mathbb E[h(\mathbf P)m(X,\mathbf P)\vert X=x] & \geq & \sup_{\varphi\in\mathbb R} \left\{\mathbb E[\max(0,h(\mathbf P)-\varphi)\vert X=x]+\varphi\mathbb E[D\vert X=x]\right\}. \end{eqnarray*} The right-hand side of ((ref)) can be treated similarly to yield the symmetric upper bound.

Although proposition (ref) is of interest in its own right, we present it here as a means to an end. A simple corollary provides bounds on the average structural function. Those bounds are simple and lend themselves to inference.

Bounds on the average structural function

To see how the bounds on the average structural function $\mu(x)$ are derived, notice that

eqnarray*[eqnarray* omitted — 254 chars of source]

with $h(\mathbf p):=f_{\mathbf P}(\mathbf p)/f_{\mathbf P\vert X=x}(\mathbf p)$, when the densities are well defined. Hence, proposition (ref) yields immediately the following bounds on the average structural function.

corollaryUnder assumptions (ref) and (ref), the average structural function satisfies \begin{eqnarray*} && \sup_\varphi \left\{ \int \min(\varphi f_{\mathbf P\vert X=x}(\mathbf p),f_{\mathbf P}(\mathbf p))\;d\mathbf p -\varphi\mathbb E[1-D\vert X=x]\right\} \\ && \hskip40pt \leq \; \mu(x) \; \leq \; \inf_\varphi \left\{ \int \max(-\varphi f_{\mathbf P\vert X=x}(\mathbf p),f_{\mathbf P}(\mathbf p))\;d\mathbf p +\varphi\mathbb E[1-D\vert X=x]\right\}, \end{eqnarray*} where $f_{\mathbf P}$ and $f_{\mathbf P\vert X}$ denote either count densities or Lebesgue densities depending on whether $\mathbf P$ is discrete or continuous.

The bounds of corollary (ref) are best understood in the special case, where $\mathbf P$ takes only two values $\mathbf P\in\{\underline{\mathbf p},\bar{\mathbf p}\}$. Hence, there are only two types in the population. In that case, the average structural function can be written in the following way:

eqnarray*[eqnarray* omitted — 289 chars of source]

Compare the previous expression with the identified expectation:

eqnarray*[eqnarray* omitted — 276 chars of source]

In the expressions above, all the terms are identified except $\mu(x)$ (the parameter of interest) and $\mathbb P[D=1\vert X=x,\mathbf P=\underline{\mathbf p}]$, $\mathbb P[D=1\vert X=x,\mathbf P=\bar{\mathbf p}]$. The bounds are attained when one of the latter takes an extreme value. As explained in cross2002regressions, the two extreme values for the pair $(\mathbb P[D=1\vert X=x,\mathbf P=\underline{\mathbf p}],\mathbb P[D=1\vert X=x,\mathbf P=\bar{\mathbf p}])$ are

eqnarray*[eqnarray* omitted — 481 chars of source]

obtained when all type $\underline{\mathbf p}$ individuals choose $D=0$ or all type $\bar{\mathbf p}$ individuals choose $D=1$, and

eqnarray*[eqnarray* omitted — 488 chars of source]

obtained when all type $\underline{\mathbf p}$ individuals choose $D=1$ or all type $\bar{\mathbf p}$ individuals choose $D=0$.

The resulting bounds on $\mu(x)$ are

eqnarray*[eqnarray* omitted — 380 chars of source]

These bounds can be obtained directly from the simple expression in corollary (ref), which remain valid for general $\mathbf P$, including mixed discrete and continuous, and lends itself to statistical inference.

Exclusion restrictions

The bounds of corollary (ref) can be tightened for more informative inference under additional covariates and exclusion restrictions. Suppose the vector of observable characteristics is now $(X,Z)$, where $X$ is the vector of choice attributes, and $Z$ is a vector of additional covariates. Define potential decision $D(x,z)$ as before, and extend the conditional independence assumption to $(X,Z)$:

assumptionp{(ref)$'$} $( D(x,z) )_{(x,z)\in\mathcal X\times\mathcal Z} \perp\!\!\!\!\perp (X, Z) \; \vert \; \eta$.

Assume that covariate $Z$ is excluded in the sense that $Z$ doesn't affect potential outcome $D(x)$ on average.

assumption[Exclusion restriction] For all $(x,z)\in\mathcal X\times\mathcal Z$, $\mathbb E[D(x,z)]=\mathbb E[D(x)]$.
corollaryUnder assumption (ref), (ref), and (ref), the average structural function satisfies \begin{eqnarray*} && \sup_{z\in\mathcal Z}\sup_{\varphi\in\mathbb R} \left\{ \int \min(\varphi f_{\mathbf P\vert X=x, Z=z}(\mathbf p),f_{\mathbf P}(\mathbf p))\;d\mathbf p -\varphi\mathbb E[1-D\vert X=x,Z=z]\right\} \\ && \hskip5pt \leq \; \mu(x) \; \leq \; \inf_{z\in\mathcal Z}\inf_{\varphi\in\mathbb R} \left\{ \int \max(-\varphi f_{\mathbf P\vert X=x,Z=z}(\mathbf p),f_{\mathbf P}(\mathbf p))\;d\mathbf p +\varphi\mathbb E[1-D\vert X=x,Z=z]\right\}, \end{eqnarray*} where $f_{\mathbf P}$ and $f_{\mathbf P\vert X}$ denote either count densities or Lebesgue densities depending on whether $\mathbf P$ is discrete or continuous.
proof[Proof of proposition (ref)] Under assumption (ref), $D(x,z)\perp (X,Z)\vert \eta \Rightarrow D(x,z)\perp (X,Z)\vert \mathbf P$. Then, we have \begin{eqnarray*} \mu(x) & = & \mathbb E[D(x)] \\ & = & \mathbb E[D(x,z)] \\ & = & \int \mathbb E[D(x,z)\vert \mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p) \\ & = & \int \mathbb E[D(x,z)\vert X=x, Z=z,\mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p) \\ & = & \int \mathbb E[D\vert X=x, Z=z, \mathbf P=\mathbf p]\;dF_{\mathbf P}(\mathbf p). \end{eqnarray*} Corollary (ref) applies for every value of $z\in\mathcal Z$, hence the result.

Inference

Inference with matched data

The average structural function is equal to $\mu(x)=\int \mathbb E[D\vert X=x,\mathbf P=\mathbf p]\;dF_{\mathbf P}$ from proposition (ref). Let $\hat{\mathbb E}[D\vert X=x,\mathbf P=\mathbf p]$ be an estimator of the conditional expectation. Then, the sample analogue

eqnarray[eqnarray omitted — 112 chars of source]

can be used as an estimator for $\mu(x)$. If a kernel estimator is used for the conditional expectation, then the standard nonparametric bootstrap delivers valid standard errors for $\hat\mu(x)$.

Inference with unmatched data

Fix the value $x$ of interest. We consider inference based on the bounds with exclusion restrictions from corollary (ref). Assume that the excluded variable $Z\in\mathcal Z=\{z_1,\ldots,z_J\}$ takes a finite number of values and the densities $f_{\mathbf P}$ and $f_{\mathbf P\vert X=x,Z=z}$ are continuous functions of $\mathbf p$ on $[0,1]^d$. We derive a one-sided confidence bound for the upper bound:

eqnarray*[eqnarray* omitted — 284 chars of source]

for some large $K>0$. The lower bound can be treated symmetrically. Define the infinite dimensional parameter $\theta_x:=(\theta_{x,z})_{z\in\mathcal Z}$ with

eqnarray*[eqnarray* omitted — 124 chars of source]

as well as a standard choice of estimator $\hat\theta_{x}:=(\hat\theta_{x,z})_{z\in\mathcal Z}$ with

eqnarray*[eqnarray* omitted — 144 chars of source]

where $\hat f_{\mathbf P\vert X=x,Z=z},$ $\hat f_{\mathbf P}$ and $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ are kernel estimators of $f_{\mathbf P\vert X=x,Z=z},$ $f_{\mathbf P}$ and $\mathbb E[1-D\vert X=x,Z=z]$ respectively. Such estimators are available in all standard software packages with automatic bandwidth procedures. Finally, define the bootstrapped estimator $\theta^\ast_x:=(\theta^\ast_{x,z})_{z\in\mathcal Z}$ with

eqnarray*[eqnarray* omitted — 144 chars of source]

where $f^\ast_{\mathbf P\vert X=x,Z=z},$ $f^\ast_{\mathbf P}$ and $\mathbb E^\ast[1-D\vert X=x,Z=z]$ are standard nonparametric bootstrapped versions of $\hat f_{\mathbf P\vert X=x,Z=z},$ $\hat f_{\mathbf P}$ and $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ respectively.

We apply the bootstrap procedure of fang2019inference in this context. Since kernel estimators do not follow a functional CLT, validity would require discretization, or possibly a fixed bandwidth design, which is beyond the scope of this work.

Define the function $\bar\phi$ of parameter $\theta$ as follows:

eqnarray*[eqnarray* omitted — 186 chars of source]
proposition[Directional differentiability] The function $\bar\phi:\theta_x\mapsto\mbox{UB}_{x}$ is Hadamard directionally differentiable on $ l_\infty([0,1]^d)^{J+1}\times([0,1]^d)^J$.
proof[Proof of proposition (ref)] We verify the conditions of theorem 4.12 page 272 of bonnans2013perturbation. Fix $z\in\mathcal Z$. Call \begin{eqnarray*} U(\varphi,\theta) & := & \int \max(-\varphi\theta_1,\theta_2)\;d\mathbf p+\varphi\theta_3, \end{eqnarray*} with $\theta_1:=f_{\mathbf P\vert X=x,Z=z}\in l_\infty[0,1]^d$, $\theta_2:=f_{\mathbf P}\in l_\infty[0,1]^d$ and $\theta_3:=\mathbb E[1-D\vert X=x,Z=z]\in[0,1]$. \begin{enumerate} • $U(\varphi,\theta)$ is continuous in all its variables. • The “inf-compactness” condition holds since for all $\theta$, $\{\varphi:U(\varphi,\theta)\leq 1\}$ is included in $[-K,K]$ which is compact. • For all $\varphi\in[-K,K]$, the function $\theta\mapsto U(\varphi,\theta)$ is directionally differentiable on $ l_\infty([0,1]^d)^2\times[0,1]$. Indeed, the only non linear feature is the $\max$. • Condition (iv) of theorem 4.12 page 272 of bonnans2013perturbation also holds by convexity of $\theta\mapsto U(\varphi,\theta)$ for all $\varphi$. See the proof of theorem 4.16 page 276 of bonnans2013perturbation. \end{enumerate}

Following hong2018numerical, estimate the directional derivative of function $\bar\phi$ with $\widehat{\bar\phi}^\prime_n$ defined for all $h\in l_\infty([0,1]^d)^{J+1}\times([0,1]^d)^J$ by

eqnarray*[eqnarray* omitted — 142 chars of source]

We use $\widehat{\bar\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ to approximate the distribution of $r_n(\bar\phi(\hat\theta_x)-\mbox{UB}_x)$, where $r_n$ is the rate of convergence of $\hat\theta_x$. Symmetrically, defining the function $\underline\phi:\theta_x\mapsto\mbox{LB}_x$, $\widehat{\underline\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ provides a valid approximation of the distribution of $r_n(\underline\phi(\hat\theta_x)-\mbox{LB}_x)$. These approximations yield the $(1-\alpha)$-level confidence region

eqnarray*[eqnarray* omitted — 162 chars of source]

where $\underline q_{1-\alpha/2}$ and $\bar q_{\alpha/2}$ are quantiles of the distribution of $\widehat{\underline\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ and $\widehat{\bar\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ respectively.

Simulations

Data generating model

We run simulation experiments to assess the informativeness of the data combination bounds and the performance of the proposed inference procedure. The data generating process is based on the retirement decision model of example (ref). Suppose the agent types are normally distributed $\eta\sim N(0,1)$. Health status $X$ at the time of decision and $Z$ at the time of elicitation take values in $\{0 (\mbox{healthy}),1 (\mbox{unhealthy})\}$. The potential retirement decision follows

eqnarray*[eqnarray* omitted — 60 chars of source]

where $\nu\sim U[0,1]$ and $\Phi$ is the cumulative distribution function of the standard normal distribution. Assume that agents have rational expectations and no reporting bias, so that the stated probability of retiring given good health is

eqnarray*[eqnarray* omitted — 77 chars of source]

Finally, health status at time of elicitation and decision are determined by

eqnarray*[eqnarray* omitted — 99 chars of source]

where $\nu$, $\nu_x$ and $\nu_z$ are independently and identically distributed. Note that the model is consistent with the interpretation that high types $\eta$ are more health conscious, hence healthier and more likely to retire at either health status.

In this model, $\tilde m(x,\eta)=m(x,\eta)=1-\Phi((x+1)\eta)$. Moreover, $\mathbf P=m(1,\eta)=1-\Phi(\eta),$ so that $\eta=\Phi^{-1}(1-\mathbf P)$. The true value of the average structural function is

eqnarray*[eqnarray* omitted — 68 chars of source]

The bounds of corollary (ref) can be easily computed in this model. Since $\mathbf P=1-\Phi(\eta),$ we have $f_{\mathbf P}(\mathbf p)=1$, and $f_{\mathbf P\vert X=x,Z=z}(\mathbf p)=3\mathbf p^2$ (resp. $6\mathbf p(1-\mathbf p)$, $6\mathbf p(1-\mathbf p)$, and $3(1-\mathbf p)^2$) if $(x,z)=(1,1)$ (resp. $(1,0)$, $(0,1)$, $(0,0)$). Since $D=XD(1)+(1-X)D(0)$, we have $\mathbb E[1-D\vert X=x,Z=z]=\mathbb P(\Phi(2\Phi^{-1}(U))\leq \nu\vert U\leq \nu_x,U\leq \nu_z)=e_1\approx0.8270$ (resp. $\mathbb P(\Phi(2\Phi^{-1}(U))\leq \nu\vert U\leq \nu_x,U\geq \nu_z)=e_0\approx0.4993$, $\mathbb P(U\leq \nu\vert U\geq \nu_x,U\leq \nu_z)=1/2$, and $\mathbb P(U\leq \nu\vert U\geq \nu_x,U\geq \nu_z)=1/4$) if $(x,z)=(1,1)$ (resp. $(1,0)$, $(0,1)$, $(0,0)$), where $U,\nu,\nu_x,\nu_z$ are four independent uniform random variables. Hence, the bounds are

eqnarray*[eqnarray* omitted — 929 chars of source]

for $x=0$, i.e., $0.371\leq\mu(0)\leq0.654$, and

eqnarray*[eqnarray* omitted — 893 chars of source]

for $x=1$ i.e., $0.346\leq\mu(1)\leq0.559$.

Inference details

In each simulation instance, we derive a sample of $n$ independent values of $\eta_i\sim N(0,1)$, and $\nu_i,\nu_{ix},\nu_{iz}$ iid $U[0,1]$, for $i=1,\ldots,n$. We compute $X_i$, $Z_i$, $\mathbf P_i$ and $D_i=X_iD_i(1)+(1-X_i)D_i(0)$ according to the data generating model in the subsection (ref).

Call $n_{xz}$ the size of the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$. Call $s$ the sample standard error for $\mathbf P_i$ and $s_{xz}$ the sample standard error for $\mathbf P_i$ in the subsample of individuals with $X_i=x$ and $Z_i=z$.

Let $\hat f_{\mathbf P}$ be the kernel density estimator for $f_{\mathbf P}$ with the STATA defaults, i.e., the Epanechikov kernel and bandwidth $h:=s n^{-1/5}$. Similarly, let $\hat f_{\mathbf P\vert X=x,Z=z}$ be the kernel density estimator for $f_{\mathbf P}$ in the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$, with the Epanechikov kernel and bandwidth $h_{xz}:=s_{xz} n_{xz}^{-1/5}$. Finally, let $\hat{\mathbb E}[D\vert X=x,Z=z]$ be the sample average of $D_i$ in the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$.

Draw $B$ bootstrap size $n$ resamples $(D_i^b,\mathbf P_i^b,X_i^b,Z_i^b)_i$ from the initial sample $(D_i,\mathbf P_i,X_i,Z_i)_i$, and for each bootstrap sample, compute $f^b_{\mathbf P}$, $f^b_{\mathbf P\vert X=x,Z=z}$ and $\mathbb E^b[D\vert X=x,Z=z]$ from sample $(D_i^b,\mathbf P_i^b,X_i^b,Z_i^b)_i$ exactly as $\hat f_{\mathbf P}$, $\hat f_{\mathbf P\vert X=x,Z=z}$ and $\hat{\mathbb E}[D\vert X=x,Z=z]$ were computed from sample $(D_i,\mathbf P_i,X_i,Z_i)_i$. Compute the numerical Hadamard directional derivative $\bar\phi^b_x$ (resp. $\underline\phi^b_x$) of the upper (resp. lower) bound at the bootstrap distribution $r_n(\theta^\ast-\hat\theta)$, where $r_n$ is the rate of convergence of the kernel estimator, i.e., $r_n=n^{2/5}$, and the tuning parameter $\xi_n$ satisfies $1/\xi_n+\xi_nr_n\rightarrow\infty$ (for instance $\xi_n=n^{-3/10}$):\footnote{Note that we neglect the variability of $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ because its rate of convergence is faster than that of $\hat f_{\mathbf P}$ and $\hat f_{\mathbf P\vert X=x,Z=z}$.}

eqnarray*[eqnarray* omitted — 1,681 chars of source]

Hence, $\bar\phi^b_x$ approximates a draw from the distribution of $r_n(\bar\phi(\hat\theta_x)-UB_x)$, and $\underline\phi^b_x$ approximates a draw from the distribution of $r_n(\underline\phi(\hat\theta_x)-LB_x)$.

Call $\bar\phi^{(b)}_x$ and $\underline\phi^{(b)}_x$, $b=1,\ldots,B$, the order statistics, where $\bar\phi^{(1)}_x$ and $\underline\phi^{(1)}_x$ are the largest values. Then, denoting $\lfloor \cdot \rfloor$ and $\lceil \cdot \rceil$ the floor and ceiling functions respectively,

eqnarray*[eqnarray* omitted — 211 chars of source]

is the bootstrap confidence region for the true value of $\mu(x)$ et the level of significance $\alpha$.

Simulation results

For each simulation exercise, we repeat the procedure in subsection (ref) $1,000$ times and report the coverage rate of the true value and the average length of the confidence interval relative to the identified set. We repeat this for sample sizes $n=500,$ $n=1,000$ and $n=2,000$, and tuning parameter values $\xi_n=0.5n^{-3/10}$, $\xi_n=0.75n^{-3/10}$, $\xi_n=n^{-3/10}$, and $\xi_n=1.5n^{-3/10}$. Table (ref) reports the rate of coverage of the true value of $\mu(0)$ and the excess length of the confidence region, relative to the identified set. Coverage rates close to $1$ are expected, since the true value is an interior point in the identified set. The confidence bands are narrow and decline both with the sample size, as expected, and somewhat also with the value of the tuning parameter $\xi_n$.

table[table omitted — 1,024 chars of source]

Discussion

In this paper, we consider agents making binary decisions based on an endogenous choice attribute. We propose to use stated choice probabilities to correct for the endogeneity. We eschew structural assumptions on the relation between stated choice probabilities and actual choice, and instead assume that stated choice probabilities reveal the unobserved heterogeneity that also governs actual choices. For the common case, where stated choice probabilities and actual choices are observed in different data sets, we derive new bounds for counterfactual choice under data combination. These bounds are useful beyond the combination of stated and revealed preferences considered in this work. Although we study binary choice, most of the ideas and results extend to discrete and continuous choice, but the data combination bounds would take a more complex form. The inference procedure proposed here is best suited to empirical environments with scalar or at most bivariate stated probability reports. For empirical environments with elicited preferences from a rich set of counterfactual scenarios, we propose, in ongoing research, to identify a finite number of latent types from elicited preferences using methods from the large factor model and group fixed effects literature.