Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,334 characters · 17 sections · 60 citation commands
Combining stated and revealed preferences
Keywords: Stated and revealed preferences, data combination, optimal transport .
JEL codes: C19, C35, D19, D84.
Social scientists use two different ways to learn about human behaviour: One possibility is to observe what people do or choose in real life, which economists call the revealed preference approach. This approach has the advantage of reflecting actual choices. Its main limitation is that choice relevant states and observed choices are driven by common unobserved underlying factors. Researchers have relied on randomised controlled trials or natural experiments to perform causal analysis. However, randomisation is not always feasible, for technical, political or ethical reasons, and when it is feasible, it is often plagued by imperfect compliance. The other possibility to learn about human behaviour is to ask people what they would do or choose in a hypothetical situation, in what is often called the stated preference approach. The main advantages of this approach are that the analyst can design experiments with rich variations in choice attributes, observe individual stated choice in counterfactual scenarios and explore preferences over policies never implemented before. Moreover, the analyst can exploit hypothetical choice scenarios to tackle endogeneity arising from omitted unobserved characteristics and identify causal effects of choice attributes. Indeed, even if the relevant choice characteristics are not all observed, exogenous variations on choice characteristics manipulated in a survey experiment are sufficient for causal inference. The obvious disadvantage of the stated preference approach is summarised by train2009: `What people say they will do is often not the same as what they actually do'. This discrepancy is known as the hypothetical bias.
Unlike other social sciences, economics has historically seen the hypothetical bias as a disqualifying argument against the stated preference approach, and until recently, economists have, by and large, relied solely on revealed preference data, except in the measurement of the non-use value of natural or cultural resources. This attitude is changing, as observed and advocated by Orazio Attanasio's 2024 Econometric Society presidential address, published in almaas2024.\footnote{The turn of the tide coincides with a surge of interest in new measures capturing individuals' subjective expectations as documented by manski2004 and almaas2024. See also mcfadden2017 for a historical perspective on stated preference data.} Stated preference analyses are increasingly used to describe individual preferences over choice attributes that are endogenous, hard to measure in observational data or hard to vary in randomized control trials. An early example is juster1964. Recent applications span many areas of economics, including education choices arcidiacono2020,delavande2019, mobility decisions gong2022, kocsar2022, meango2022, batista2025, health and long-term care investments kesternich2013, ameriks2020b,boyer2020, parental investments attanasio2019,almaas2024, marriage preference adams2019, low2024, occupational choices maestas2023, wiswall2015, wiswall2018,meango2024, and retirement decisions ameriks2020a, giustinelli2024. For recent reviews, see kocsar2023 and giustinelli2023. Despite the enthusiasm and encouraging evidence that carefully designed stated preference elicitation has predictive power for actual choices hurd2009, hainmueller2015, debresser2019, arcidiacono2020 and yields similar preference estimates as actual choice data mas2017, wiswall2018, concerns remain regarding systematic biases of stated preference data. See, for example, murphy2005 and hausman2012 and more recently haghani2021a,haghani2021b.
An alternative approach consists in combining stated preference data with revealed preference data, in order to mitigate their respective shortcomings (hypothetical bias on the one hand and endogenous choice attributes on the other). However, research in this direction has so far ignored the fact that revealed and stated preferences generally occur in different data sets, in which individuals can rarely be matched. In addition, research in this direction has so far relied on very strong structural assumptions on the relation between stated and revealed preference. morikawa2002, pantano2013, and giustinelli2024 restrict the difference between utility parameters that govern stated and revealed preferences; vanderklaauw2012 and wiswall2021 assume that any bias between statement and action comes solely from biased information; and bernheim2022 assume that the treatment affects actual choice only through the stated preference, so that the difference between stated and revealed preferences is stable across treatment groups. briggs2020 assume additively separable and scalar unobserved heterogeneity, a specific timing of the resolution of uncertainty, and rational expectations in their identification of marginal treatment effects from subjective expectations.\footnote{athey2025combining apply similar ideas to the combination of observational and experimental data.}
This paper proposes to combine revealed preference data with stated preference data, matched or unmatched, to analyze binary choice with endogenous choice attributes, without strong assumptions on the way the two are related. We propose a strategy to retrieve information about individual unobserved heterogeneity from stated preferences and we derive a new result on data combination to apply this strategy in the case, where revealed and stated preferences are observed in distinct, unmatched data sets. The fundamental idea is that stated preferences are useful, not because they necessarily match actual choices, but insofar as they provide valuable information on individual heterogeneity. Suppose that an analyst asks a job-seeker about their probability to take up a particular job. The probability is elicited for different scenarios varying wage, job security, and non-wage compensation. A respondent who reports a 90 percent chance of taking up the job in all scenarios is arguably different from a respondent who states a 90 percent chance if the wage is high, and a 10 percent chance otherwise. The first respondent might have a higher taste for work, lower disposable income or lower ability than the second respondent. Irrespective of the reason, and even if `what they say is often not what they actually do', their responses reveal important information about how they differ in their preferences about job attributes. This information helps to classify individuals into unobserved heterogeneity types. The dimension of unobserved heterogeneity that can be recovered depends on the richness of the stated choice experiment. Repeated elicitation in different scenarios allows the recovery of multiple dimensions of unobserved heterogeneity.
In our framework, agents make a binary decision $D$ based on an endogenous decision relevant attribute $X$. The parameter of interest is $\mu(x):=\mathbb E[D(x)]$, where $D(x)$ is the potential decision when the decision relevant attribute is exogenously set to $x$. We derive conditions under which a vector of stated preferences reports $\mathbf P$ (typically stated choice probabilities) can be used to identify the unobserved heterogeneity that causes endogeneity. More precisely, we derive conditions under which the parameter of interest is identified as $\mu(x)=\int\mathbb E[D\vert X=x,\mathbf P=\mathbf p]\;dF_{\mathbf P}(p)$, where $F_{\mathbf P}$ is the probability distribution of stated preference reports $\mathbf P$. For the common case, when actual choices $D$ and stated choice probabilities $\mathbf P$ are not observed in the same data set, we derive bounds on $\mu(x)$ based on knowledge of the distributions of $D\vert X$ from revealed preferences and of $\mathbf P\vert X$ from stated preferences.
horowitz1995identification, cross2002regressions and molinari2006generalization derive bounds on the conditional expectation $\mathbb E[D\vert X,\mathbf P]$ based on the distributions of $D\vert X$ and $\mathbf P\vert X$, when $\mathbf P$ has finite support. Here, however, $\mathbf P$ may be discrete or continuous, and we need bounds on the integral of $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ over $\mathbf p$, which involves constraints across values of $\mathbf p$ beyond the simple bounds in horowitz1995identification. cross2002regressions derive bounds for the function $p\mapsto \mathbb E[D\vert X=x,P=p]$. Those bounds are functional, hence they do take into account restrictions across $p$. More precisely, they show that the identified set for $p\mapsto \mathbb E[D\vert X=x,P=p]$ is the convex hull of a set of $J\!$ extreme points (where $J$ is the cardinality of the support of $P$). We derive simple bounds on the integral of this function over $p$, i.e., the parameter described in section 4 of cross2002regressions. Our bounds lend themselves to straightforward inference, and they remain valid when $P$ is continuous (or mixed). As such, our bounds complement the work of fan2014identifying (fan2014identifying,fan2016estimation), who also generalize the bounds of cross2002regressions and allow for mixed discrete and continuous covariates and outcomes. They derive bounds for the expectation $\mathbb E[g(Y_d)]$, as well as the counterfactual distributions, quantile and distributional treatment effects, under unconfoundedness $(Y_0,Y_1)\perp\!\!\!\!\perp D\vert Z$, where the distributions of $((Y_1D+Y_0(1-D)),D)$ and $(Z,D)$ are identified from two separate data sets.
Defining $m(x,\mathbf p)=\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$, we derive the bounds from the dual of the optimal transport problems
where the supremum is over all joint distributions for $(D,\mathbf P)\vert X=x$ with fixed marginals $D\vert X=x$ and $\mathbf P\vert X=x$. The application of ideas related to the theory of optimal transport in data combination problems can be traced back to horowitz1995identification and heckman1997. More recently, specific aspects of optimal transport theory were used to derive bounds in specific data combination problems in fan2014identifying (fan2014identifying,fan2016estimation) and lee2019identification for treatments effects under data combination, pacini2019two, gaillac2024linear (gaillac2024linear,gaillac2025partially) for partially linear regression and best linear prediction under data combination. The most closely related to ours is fan2025partial, which applies the `method of conditioning' from ruschendorf1991bounds to derive bounds on the parameter of a moment equality model under data combination. ichimura2005identification and bontemps2025functional use a different approach, in that they seek conditions for point identification of parameters of moment equality models despite data combination.
We show how to tighten the bounds under exclusion restrictions and we use the bootstrap procedure proposed by fang2019inference to the bounds, as the latter are Hadamard directionally differentiable functions of two density functions and a conditional expectation. We investigate the tightness of the bounds and the performance of the inference procedure in a simulation experiment.
The rest of the paper is organized as follows. Section (ref) lays out the econometric framework and the main identification and partial identification results. Sections (ref) and (ref) describe the inference procedure and the simulation experiment respectively. T he last section concludes.
Consider an economic agent with a vector $X$ of economic attributes with support $\mathcal X\subseteq \mathbb R^{d_x}$, who is facing binary decision $D\in\{0,1\}$. We denote $D(x)$ the potential decision, i.e., the decision the agent would make if their attributes were externally set to $X=x$. Both the vector of attributes $X$ and the actual decision $D:=D(X)$ are observed. Potential decision $D(x)$ is observed only when $X=x$ and not otherwise. The attribute $X$ is endogenous in the sense that $D(x)$ is not independent of $X$, and hence $\mathbb E[D\vert X=x]$, which is identified from the data, is generally not equal to $\mathbb E[D(x)]$, which is the parameter of interest. Finally, let $\eta$ be a vector of unobserved characteristics. The support $\mathcal H\subseteq\mathbb R^d$ of $\eta$ is independent of $X$. We posit that $\eta$ contains the source of endogeneity of $X$ in the sense of the following assumption.
Assumption (ref) is related to the control function approach pioneered by heckman1985alternative. Under this assumption, the structural choice function $\mathbb E[D(x)\vert\eta]$ satisfies
so that recovering $\eta$ would allow identification of $m(x,\eta)$ and hence of the following counterfactual choice parameters.
Stated preferences are data collected from a survey of the agents prior to the time of decision. Agents are asked to give an assessment of their probability of making decision $D=1$ for some hypothetical values of $x$. Let $(x_1,\ldots,x_T)$ be a vector of hypothetical values $x_t\in\mathcal X$. The agent's response to the probability elicitation question is denoted $P_t$, $t=1,\ldots,T$.
We model $P_t$, for each $t=1,\ldots,T$ as
where the function $\tilde m(x,\eta)$ is called the stated choice function. In the special case where agents have rational expectations and no reporting bias, $P_t=\mathbb E[D(x)\vert \eta]$, so that $\tilde m(x_t,\eta)=m(x_t,\eta)$ for $t=1,\ldots,T$. In this case, the identification issue boils down into recovering $ m(x,\eta)$ for $x \not \in (x_1,\ldots,x_T)$. In general, the stated preference function $\tilde m(x_t,\eta)$ may differ from the structural choice function $m(x_t,\eta)$ for revealed preferences, due to the agent's perception and/or reporting biases. Crucially, the agent's report $P_t$ is driven by the same vector $\eta$ of unobserved characteristics. This provides the only link in the model between stated and revealed preferences.
Identification of the parameters of interest relies on the ability to control for unobserved heterogeneity $\eta$ by matching on stated preference reports $\mathbf P=(P_1,\ldots,P_T)$. For this, we need to be able to recover unobserved heterogeneity from the stated preference reports of an individual.
Assumption (ref) guarantees that a unique $\eta$ can be recovered from the system $\mathbf P: = (P_1, \ldots, P_d)= (\tilde m(x_1,\eta),\ldots,\tilde m(x_d,\eta)):= \tilde m(\mathbf x,\eta)$. An example often used in empirical applications is a multiplicatively separable log-odd model, of which the model of blass2010 is a special case. It corresponds to:
where $\Gamma:\mathcal H\rightarrow[0,1]^d$ is one-to-one and $v(\mathbf x)$ is a full rank $d\times d$ matrix. In this case, assumption (ref) holds with $\eta = v(\mathbf x)^{-1} \left({\Gamma^{-1}(\mathbf P) - r(\mathbf x)}\right)$. More generally, sufficient conditions for assumption (ref) are given in Theorem 2 of mas1979. In model ((ref)), $\mathbf P$ has the same dimension as $\eta$. This need not be the case for assumption (ref) to hold, as long as the dimension of $\mathbf P$ is at least as large as the dimension of $\eta$. All that is needed, is that a subvector of $\mathbf P$ identifies $\eta$. This subvector need not be known, as long as it exists. Similarly, the stated preference function $\tilde m$ and the true dimension of $\eta$ need not be known. Assumption (ref), combined with assumption (ref), ensures that $D(x)\perp\!\!\!\!\perp X\vert \mathbf P$, which drives proposition (ref) below.
Assumption (ref) allows for any kind of hypothetical bias, providing that responses to hypothetical scenarios are sufficiently different to distinguish individuals who respond differently to choice variables in actual choices. If two individuals react differently to health shocks, possibly because they have a different taste for work, then Assumption (ref) holds if these individuals also respond differently to the hypothetical questions, one giving answers that are much more sensitive to the scenario than the other, even if both wildly overstate their willingness to work. Assumption (ref) requires the type $\eta$ of individuals to be fixed between scenarios and between stated and revealed preferences, and two individuals with different types to be distinguishable from their survey responses. Assumption (ref) would be violated for instance if an individual with low taste for work were induced by social desirability bias to give the same responses to the survey as an individual with high taste for work would.
As illustrated in the example, under assumption (ref), stated preferences allow us to control for unobserved heterogeneity in the regression of decision $D$ on the choice attributes $x$, and to identify the parameters of interest.
Proposition (ref) provides a constructive identification result. Inference on $\mu(x)$ can be based on the sample average of any estimator of the conditional expectation $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$.
In most empirical environments, stated preferences and revealed preferences are collected in different data sets, and it is difficult or impossible to match individuals from the two data sets. Hence, the conditional distribution $F_{D\vert X}$ of $D$ given $X$ is identified from the revealed preference data set, and the conditional distribution $F_{\mathbf P\vert X}$ of $\mathbf P$ given $X$ is identified from the stated preference data set. However, the conditional distribution $F_{D,\mathbf P\vert X}$ of $(D,\mathbf P)$ given $X$ is not point identified. It can take any value in the set $\mathcal M(F_{D\vert X=x},F_{\mathbf P\vert X=x})$, where $\mathcal M(F,F^\prime)$ is the set of joint distributions with marginals $F$ and $F^\prime$.
For each conditional distribution $F_{D,\mathbf P\vert X}$ for $(D,\mathbf P)$ given $X$, proposition (ref) identifies a unique value for the structural function $\mathbb E[D\vert X,\mathbf P]$ and for the average structural function $\mu(x)$. However, since there are multiple distributions $F_{D,\mathbf P\vert X}$ compatible with the marginals $F_{D\vert X}$ and $F_{\mathbf P\vert X}$, both $\mathbb E[D\vert X,\mathbf P]$ and $\mu(x)$ are partially identified.
We first derive sharp bounds for the conditional expectation $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$. Under assumption (ref), there is a one-to-one mapping between $\mathbf P$ and $\eta$. Hence, we abuse notation and call $m(x,\mathbf p)=\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ the structural function. The sharp identified set for the structural function $m(x,\mathbf p)$ contains all the conditional expectations $\mathbb E[D\vert X=x,\mathbf P=\mathbf p]$ compatible with marginal distributions $F_{D\vert X}$ and $F_{\mathbf P\vert X}$.
We first give a characterization of the identified set in the form of sharp bounds on $m(x,\mathbf p)$. By definition of $m(x,\mathbf p)$, we have $\mathbb E[D-m(X,\mathbf P)\vert X=x,\mathbf P=\mathbf p]=0,$ which is equivalent to
for all continuous and integrable functions $h:\mathbb R^T\rightarrow\mathbb R$. The existence of a distribution in $\mathcal M(F_{D\vert X},F_{\mathbf P\vert X})$ such that ((ref)) holds is equivalent to
where the optimization is over $F\in\mathcal M(F_{D\vert X=x},F_{\mathbf P\vert X=x})$, and $h$ is any continuous integrable function of $\mathbf p$. The bounds are obtained through an application of optimal transport duality to the optimal transport problems on the left and right-hand sides of ((ref)). This yields the following proposition.
Although proposition (ref) is of interest in its own right, we present it here as a means to an end. A simple corollary provides bounds on the average structural function. Those bounds are simple and lend themselves to inference.
To see how the bounds on the average structural function $\mu(x)$ are derived, notice that
with $h(\mathbf p):=f_{\mathbf P}(\mathbf p)/f_{\mathbf P\vert X=x}(\mathbf p)$, when the densities are well defined. Hence, proposition (ref) yields immediately the following bounds on the average structural function.
The bounds of corollary (ref) are best understood in the special case, where $\mathbf P$ takes only two values $\mathbf P\in\{\underline{\mathbf p},\bar{\mathbf p}\}$. Hence, there are only two types in the population. In that case, the average structural function can be written in the following way:
Compare the previous expression with the identified expectation:
In the expressions above, all the terms are identified except $\mu(x)$ (the parameter of interest) and $\mathbb P[D=1\vert X=x,\mathbf P=\underline{\mathbf p}]$, $\mathbb P[D=1\vert X=x,\mathbf P=\bar{\mathbf p}]$. The bounds are attained when one of the latter takes an extreme value. As explained in cross2002regressions, the two extreme values for the pair $(\mathbb P[D=1\vert X=x,\mathbf P=\underline{\mathbf p}],\mathbb P[D=1\vert X=x,\mathbf P=\bar{\mathbf p}])$ are
obtained when all type $\underline{\mathbf p}$ individuals choose $D=0$ or all type $\bar{\mathbf p}$ individuals choose $D=1$, and
obtained when all type $\underline{\mathbf p}$ individuals choose $D=1$ or all type $\bar{\mathbf p}$ individuals choose $D=0$.
The resulting bounds on $\mu(x)$ are
These bounds can be obtained directly from the simple expression in corollary (ref), which remain valid for general $\mathbf P$, including mixed discrete and continuous, and lends itself to statistical inference.
The bounds of corollary (ref) can be tightened for more informative inference under additional covariates and exclusion restrictions. Suppose the vector of observable characteristics is now $(X,Z)$, where $X$ is the vector of choice attributes, and $Z$ is a vector of additional covariates. Define potential decision $D(x,z)$ as before, and extend the conditional independence assumption to $(X,Z)$:
Assume that covariate $Z$ is excluded in the sense that $Z$ doesn't affect potential outcome $D(x)$ on average.
The average structural function is equal to $\mu(x)=\int \mathbb E[D\vert X=x,\mathbf P=\mathbf p]\;dF_{\mathbf P}$ from proposition (ref). Let $\hat{\mathbb E}[D\vert X=x,\mathbf P=\mathbf p]$ be an estimator of the conditional expectation. Then, the sample analogue
can be used as an estimator for $\mu(x)$. If a kernel estimator is used for the conditional expectation, then the standard nonparametric bootstrap delivers valid standard errors for $\hat\mu(x)$.
Fix the value $x$ of interest. We consider inference based on the bounds with exclusion restrictions from corollary (ref). Assume that the excluded variable $Z\in\mathcal Z=\{z_1,\ldots,z_J\}$ takes a finite number of values and the densities $f_{\mathbf P}$ and $f_{\mathbf P\vert X=x,Z=z}$ are continuous functions of $\mathbf p$ on $[0,1]^d$. We derive a one-sided confidence bound for the upper bound:
for some large $K>0$. The lower bound can be treated symmetrically. Define the infinite dimensional parameter $\theta_x:=(\theta_{x,z})_{z\in\mathcal Z}$ with
as well as a standard choice of estimator $\hat\theta_{x}:=(\hat\theta_{x,z})_{z\in\mathcal Z}$ with
where $\hat f_{\mathbf P\vert X=x,Z=z},$ $\hat f_{\mathbf P}$ and $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ are kernel estimators of $f_{\mathbf P\vert X=x,Z=z},$ $f_{\mathbf P}$ and $\mathbb E[1-D\vert X=x,Z=z]$ respectively. Such estimators are available in all standard software packages with automatic bandwidth procedures. Finally, define the bootstrapped estimator $\theta^\ast_x:=(\theta^\ast_{x,z})_{z\in\mathcal Z}$ with
where $f^\ast_{\mathbf P\vert X=x,Z=z},$ $f^\ast_{\mathbf P}$ and $\mathbb E^\ast[1-D\vert X=x,Z=z]$ are standard nonparametric bootstrapped versions of $\hat f_{\mathbf P\vert X=x,Z=z},$ $\hat f_{\mathbf P}$ and $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ respectively.
We apply the bootstrap procedure of fang2019inference in this context. Since kernel estimators do not follow a functional CLT, validity would require discretization, or possibly a fixed bandwidth design, which is beyond the scope of this work.
Define the function $\bar\phi$ of parameter $\theta$ as follows:
Following hong2018numerical, estimate the directional derivative of function $\bar\phi$ with $\widehat{\bar\phi}^\prime_n$ defined for all $h\in l_\infty([0,1]^d)^{J+1}\times([0,1]^d)^J$ by
We use $\widehat{\bar\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ to approximate the distribution of $r_n(\bar\phi(\hat\theta_x)-\mbox{UB}_x)$, where $r_n$ is the rate of convergence of $\hat\theta_x$. Symmetrically, defining the function $\underline\phi:\theta_x\mapsto\mbox{LB}_x$, $\widehat{\underline\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ provides a valid approximation of the distribution of $r_n(\underline\phi(\hat\theta_x)-\mbox{LB}_x)$. These approximations yield the $(1-\alpha)$-level confidence region
where $\underline q_{1-\alpha/2}$ and $\bar q_{\alpha/2}$ are quantiles of the distribution of $\widehat{\underline\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ and $\widehat{\bar\phi}^\prime_n(r_n(\theta^\ast_x-\hat\theta_x))$ respectively.
We run simulation experiments to assess the informativeness of the data combination bounds and the performance of the proposed inference procedure. The data generating process is based on the retirement decision model of example (ref). Suppose the agent types are normally distributed $\eta\sim N(0,1)$. Health status $X$ at the time of decision and $Z$ at the time of elicitation take values in $\{0 (\mbox{healthy}),1 (\mbox{unhealthy})\}$. The potential retirement decision follows
where $\nu\sim U[0,1]$ and $\Phi$ is the cumulative distribution function of the standard normal distribution. Assume that agents have rational expectations and no reporting bias, so that the stated probability of retiring given good health is
Finally, health status at time of elicitation and decision are determined by
where $\nu$, $\nu_x$ and $\nu_z$ are independently and identically distributed. Note that the model is consistent with the interpretation that high types $\eta$ are more health conscious, hence healthier and more likely to retire at either health status.
In this model, $\tilde m(x,\eta)=m(x,\eta)=1-\Phi((x+1)\eta)$. Moreover, $\mathbf P=m(1,\eta)=1-\Phi(\eta),$ so that $\eta=\Phi^{-1}(1-\mathbf P)$. The true value of the average structural function is
The bounds of corollary (ref) can be easily computed in this model. Since $\mathbf P=1-\Phi(\eta),$ we have $f_{\mathbf P}(\mathbf p)=1$, and $f_{\mathbf P\vert X=x,Z=z}(\mathbf p)=3\mathbf p^2$ (resp. $6\mathbf p(1-\mathbf p)$, $6\mathbf p(1-\mathbf p)$, and $3(1-\mathbf p)^2$) if $(x,z)=(1,1)$ (resp. $(1,0)$, $(0,1)$, $(0,0)$). Since $D=XD(1)+(1-X)D(0)$, we have $\mathbb E[1-D\vert X=x,Z=z]=\mathbb P(\Phi(2\Phi^{-1}(U))\leq \nu\vert U\leq \nu_x,U\leq \nu_z)=e_1\approx0.8270$ (resp. $\mathbb P(\Phi(2\Phi^{-1}(U))\leq \nu\vert U\leq \nu_x,U\geq \nu_z)=e_0\approx0.4993$, $\mathbb P(U\leq \nu\vert U\geq \nu_x,U\leq \nu_z)=1/2$, and $\mathbb P(U\leq \nu\vert U\geq \nu_x,U\geq \nu_z)=1/4$) if $(x,z)=(1,1)$ (resp. $(1,0)$, $(0,1)$, $(0,0)$), where $U,\nu,\nu_x,\nu_z$ are four independent uniform random variables. Hence, the bounds are
for $x=0$, i.e., $0.371\leq\mu(0)\leq0.654$, and
for $x=1$ i.e., $0.346\leq\mu(1)\leq0.559$.
In each simulation instance, we derive a sample of $n$ independent values of $\eta_i\sim N(0,1)$, and $\nu_i,\nu_{ix},\nu_{iz}$ iid $U[0,1]$, for $i=1,\ldots,n$. We compute $X_i$, $Z_i$, $\mathbf P_i$ and $D_i=X_iD_i(1)+(1-X_i)D_i(0)$ according to the data generating model in the subsection (ref).
Call $n_{xz}$ the size of the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$. Call $s$ the sample standard error for $\mathbf P_i$ and $s_{xz}$ the sample standard error for $\mathbf P_i$ in the subsample of individuals with $X_i=x$ and $Z_i=z$.
Let $\hat f_{\mathbf P}$ be the kernel density estimator for $f_{\mathbf P}$ with the STATA defaults, i.e., the Epanechikov kernel and bandwidth $h:=s n^{-1/5}$. Similarly, let $\hat f_{\mathbf P\vert X=x,Z=z}$ be the kernel density estimator for $f_{\mathbf P}$ in the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$, with the Epanechikov kernel and bandwidth $h_{xz}:=s_{xz} n_{xz}^{-1/5}$. Finally, let $\hat{\mathbb E}[D\vert X=x,Z=z]$ be the sample average of $D_i$ in the subsample of individuals $i$ with $X_i=x$ and $Z_i=z$.
Draw $B$ bootstrap size $n$ resamples $(D_i^b,\mathbf P_i^b,X_i^b,Z_i^b)_i$ from the initial sample $(D_i,\mathbf P_i,X_i,Z_i)_i$, and for each bootstrap sample, compute $f^b_{\mathbf P}$, $f^b_{\mathbf P\vert X=x,Z=z}$ and $\mathbb E^b[D\vert X=x,Z=z]$ from sample $(D_i^b,\mathbf P_i^b,X_i^b,Z_i^b)_i$ exactly as $\hat f_{\mathbf P}$, $\hat f_{\mathbf P\vert X=x,Z=z}$ and $\hat{\mathbb E}[D\vert X=x,Z=z]$ were computed from sample $(D_i,\mathbf P_i,X_i,Z_i)_i$. Compute the numerical Hadamard directional derivative $\bar\phi^b_x$ (resp. $\underline\phi^b_x$) of the upper (resp. lower) bound at the bootstrap distribution $r_n(\theta^\ast-\hat\theta)$, where $r_n$ is the rate of convergence of the kernel estimator, i.e., $r_n=n^{2/5}$, and the tuning parameter $\xi_n$ satisfies $1/\xi_n+\xi_nr_n\rightarrow\infty$ (for instance $\xi_n=n^{-3/10}$):\footnote{Note that we neglect the variability of $\hat{\mathbb E}[1-D\vert X=x,Z=z]$ because its rate of convergence is faster than that of $\hat f_{\mathbf P}$ and $\hat f_{\mathbf P\vert X=x,Z=z}$.}
Hence, $\bar\phi^b_x$ approximates a draw from the distribution of $r_n(\bar\phi(\hat\theta_x)-UB_x)$, and $\underline\phi^b_x$ approximates a draw from the distribution of $r_n(\underline\phi(\hat\theta_x)-LB_x)$.
Call $\bar\phi^{(b)}_x$ and $\underline\phi^{(b)}_x$, $b=1,\ldots,B$, the order statistics, where $\bar\phi^{(1)}_x$ and $\underline\phi^{(1)}_x$ are the largest values. Then, denoting $\lfloor \cdot \rfloor$ and $\lceil \cdot \rceil$ the floor and ceiling functions respectively,
is the bootstrap confidence region for the true value of $\mu(x)$ et the level of significance $\alpha$.
For each simulation exercise, we repeat the procedure in subsection (ref) $1,000$ times and report the coverage rate of the true value and the average length of the confidence interval relative to the identified set. We repeat this for sample sizes $n=500,$ $n=1,000$ and $n=2,000$, and tuning parameter values $\xi_n=0.5n^{-3/10}$, $\xi_n=0.75n^{-3/10}$, $\xi_n=n^{-3/10}$, and $\xi_n=1.5n^{-3/10}$. Table (ref) reports the rate of coverage of the true value of $\mu(0)$ and the excess length of the confidence region, relative to the identified set. Coverage rates close to $1$ are expected, since the true value is an interior point in the identified set. The confidence bands are narrow and decline both with the sample size, as expected, and somewhat also with the value of the tuning parameter $\xi_n$.
In this paper, we consider agents making binary decisions based on an endogenous choice attribute. We propose to use stated choice probabilities to correct for the endogeneity. We eschew structural assumptions on the relation between stated choice probabilities and actual choice, and instead assume that stated choice probabilities reveal the unobserved heterogeneity that also governs actual choices. For the common case, where stated choice probabilities and actual choices are observed in different data sets, we derive new bounds for counterfactual choice under data combination. These bounds are useful beyond the combination of stated and revealed preferences considered in this work. Although we study binary choice, most of the ideas and results extend to discrete and continuous choice, but the data combination bounds would take a more complex form. The inference procedure proposed here is best suited to empirical environments with scalar or at most bivariate stated probability reports. For empirical environments with elicited preferences from a rich set of counterfactual scenarios, we propose, in ongoing research, to identify a finite number of latent types from elicited preferences using methods from the large factor model and group fixed effects literature.