Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,372 characters · 9 sections · 56 citation commands
Using Probabilistic Stated Preference Analyses to Understand Actual Choices
Keywords: Stated preference; revealed preference; unobserved heterogeneity.\\
JEL codes: C30, C33, D84.
Stated preference analyses typically present agents with discrete alternatives in several hypothetical scenarios and elicit their intended choice in each. Recent applications are wide-spanning and include publications in leading journals in economics: among other topics, education and degree choices delavande2019, arcidiacono2020, mobility decisions gong2022, kocsar2022, health and long-term care investments kesternich2013, ameriks2020b,boyer2020, parental investments almaas2023, voting behaviour delavande2015, debresser2019, marriage preferences adams2019, occupational choices wiswall2015, maestas2018, wiswall2018, retirement decisions ameriks2020a and irregular migration bah2018,meango2022. See also a recent literature review by kocsar2023.\footnote{A larger literature on stated choice analysis is referenced in kocsar2022, footnotes 6 and 7. almaas2023 provides an historical perspective on the debate over the suitability of stated preference and revealed preference analyses.} almaas2023, taking inspiration from Orazio Attanasio's presidential address to the Econometric society, provides compelling arguments for empirical strategies integrating stated and revealed preferences. They note: `recent developments indicate that the [economics] profession is moving towards using choice data and directly observable variables in combination with stated preferences and answers to hypothetical questions.'
The increasing interest in the stated preference approach stems from its numerous advantages over the revealed preference approach (i.e. an approach using actual choices only): it avoids assumptions about agents' belief formation and equilibrium allocation mechanisms, and is a more natural approach to questions pertaining to ex ante perceptions. Using probabilistic stated choices, that is, the chance of choosing an alternative on a scale from 0 to 100, rather than a binary answer, provides the means for respondents to express uncertainty about their future choice and for researchers to learn about this uncertainty blass2010, meango2023. One key advantage of stated preference analyses is that the analyst can exploit hypothetical choice scenarios to tackle the pervasive problem of endogeneity arising from omitted variable bias and identify causal effects of choice attributes.\footnote{ameriks2020b use hypothetical choices elicited in strategic survey question to estimate some preference parameters in a fully specified structural model. The models addressed in this paper are nonparametric by nature.} Even if not all relevant choice characteristics are observed, exogenous variations on observed characteristics induced in a survey experiment are sufficient for causal inference.
Given the increasing importance of stated preference analyses, it is essential to address an ever-recurring question: Can stated choices serve to understand actual behaviour or causal effects in actual choice environments? There is growing evidence that a carefully designed stated preference elicitation often has predictive power for actual choices in several contexts hurd2009, bah2018, wiswall2018, debresser2019, arcidiacono2020 but concerns remain regarding systematic biases of this method. See, for example, murphy2005 and more recently haghani2021a,haghani2021b. The conventional approach to the above question, when one has access to both hypothetical and actual choices, remains unsophisticated. It involves assessing the correlation between stated and actual outcomes, or measuring the size of the biases between stated and average choice in sub-populations defined by observed characteristics. Some studies have combined elicited and actual choices in structural model estimation, finding some gains in estimation precision.\footnote{See, for example, vanderklaauw2012, kesternich2013, zafar2013 and an earlier literature in transport and environmental economics, surveyed in whitehead2008.} However, counterfactual analyses using subjective choices have largely remained separate from causal inference on objective choices, since there is no consensus on remedies to biases in stated preference analyses. Two notable exceptions are briggs2020 and bernheim2022 that are discussed below.
This research proposes a novel and simple approach to researchers who have access to both stated choices in hypothetical scenarios and actual choices in a binary choice environment, and wish to estimate counterfactual parameters. The key idea is that probabilistic stated choices not only allow for learning about respondents' uncertainty, but also provide the means to learn about the distribution of individual unobserved heterogeneity. To motivate with a simple example, suppose that an analyst has access to data on intended migration choices in different hypothetical scenarios, and subsequent migration behaviour. A respondent who reports a high probability of migration in most scenarios is arguably different from a respondent who states a low probability of migration in most scenarios. Thus, the stated preferences contain some information about individual `types', and can be harnessed to discriminate among types in the population. Once these types have been identified, they may serve as control variables in a demand function estimation using individual's actual choices.
This paper provides a general framework where the distribution of multi-dimensional unobserved heterogeneity can be identified from probabilistic stated choices in different hypothetical scenarios. The approach allows for respondents to make mistakes and/or be biased about the future, as long as (1) the observable and unobservable attributes that matter in the actual choice also influence the stated choice, and, if the analyst is interested in causal inference on the demand function, (2) the unobserved heterogeneity in the stated choice contains a sub-vector that is the only source of endogeneity. Hence, stated preferences can serve as control functions. These requirements depart from the previous literature that combines both stated preference and revealed preference data while assuming that expectations about future behavior accurately portray optimal future behaviour conditional on current information vanderklaauw2012, or from the conditions in bernheim2022 that the difference between stated and revealed demand function is stable across treatment groups.
Section (ref) presents the econometric framework that justifies matching individuals on their stated preferences before performing a regression of actual choices on choice characteristics. It shows identification for the mean squared bias that gives an accurate assessment of the bias between stated and actual choices, counterfactual distribution functions in the spirit of chernozhukov2013, and the average structural function blundell2004, that serves for causal inference of the effect of choice attributes on actual choice. The case of an uni-dimensional heterogeneity serves as a motivating example. In this case, a probabilistic stated choice in only one hypothetical scenario is required. When the heterogeneity is multidimensional, say $d>1$, at least $d$ scenarios are required for identifying the unobserved heterogeneity.
Section (ref) incorporates (classical) measurement error. It shows that even with a small number of hypothetical scenarios, the objects of interest are identified. The result rests on kotlarski1967's Lemma, as the identification results for nonseparable panel models in evdokimov2010 and meango2023, respectively.
Section (ref) considers estimation. It adapts the Two-Step Group Fixed Effects estimator (TSGFE) proposed by bonhomme2022. TSGFE exploits auxiliary moments to classify individuals into a finite number of types, based on a continuous, low-dimensional unobserved heterogeneity. The recovered types are then used in the estimation of the parameters of interest. In this context, TSGFE is a natural approach to use, as probabilistic stated choices generate identifying moments for individual unobserved heterogeneity and serve to classify individuals. The identified types correct for the influence of the unobserved heterogeneity on the parameters of interest.
Section (ref) discusses the importance of the result for practitioners, and its link to the proposals from briggs2020 and bernheim2022. Section (ref) concludes. Appendix (ref) considers the question of whether a given survey experiment exhausts the information about the unobserved heterogeneity and derives testable implications for restrictions on the dimension of unobserved heterogeneity.
Consider an economic agent $i$, a binary choice alternative $0$ or $1$, and two consecutive periods: a time of preference elicitation and a time of decision. At the time of decision, $i$ chooses between option $0$ and option $1$ based on a threshold-crossing rule:
The notation borrows from the potential outcome framework, as $D_i(x)$ represents $i$'s choice, when the choice attributes are characterised by a vector $x$, and $D_i$ is the actual choice, as observed by the analyst at the time of decision. The potential outcomes notation emphasises the analyst's interest in `treatment effects' and counterfactual policies. The choice is binary and can be, for example, between migrating or not, following a STEM degree or not, retiring early or late. The vector-valued variable $X_i$ represents choice characteristics that are observed by the analyst at the time of decision and can be manipulated in a choice experiment. In the example of a migration decision, it could include the average wage at origin and destination and the pecuniary cost of moving. In the example of major choice, it could include average wages in STEM and non-STEM occupations, average female representation, or average success rates.
The variable $\eta_i$ subsumes unobserved preferences over choice attributes that influence $i$'s decision, for example, $i$'s taste for migration or innate ability in math-related subjects. The dimension of $\eta$ is finite and the reader should think of $\eta$ as having a low dimension. The variable $\nu_i^r$ captures the resolved uncertainty, a realisation of what is known as the resolvable uncertainty. It summarises choice characteristics that are unknown to $i$ at the time of elicitation, but are revealed at the time of decision blass2010. The implicit assumption of Equation ((ref)) that $\nu_i^r$ is separable and univariate is standard in the stated choice literature and most discrete choice models. It implies that the unobserved (dis-)utility caused by preference shocks does not vary with the observed choice characteristics and the unobserved heterogeneity. This is equivalent to the assumption of a univariate, non-separable resolvable uncertainty such that the mapping $\nu \mapsto S^r(x,n,\nu)$ is strictly increasing for each $(x,n) \in \mathcal{X} \times \mathcal{H}$, the joint support of $X_{i}$ and $\eta_i$ vytlacil2002. The characteristics $\eta_i$ is thought of as being stable, that is, it does not change between the time of elicitation and the time of decision, but may be correlated with the resolved uncertainty. Finally, $S^r \left({X_i,\eta_i}\right)-\nu_i^r$ can be interpreted as $i$'s perceived returns on choosing option 1 over option 0.
Prior to their decision and at the time of elicitation, $i$ is asked to state their preference over the binary choice alternative $0$ or $1$ in a hypothetical scenario $t$ where $X_{it} = x_{it}$ and reports:
where $F_{\nu \vert X_i, \eta_i}(v \vert x_{it},\eta_i)$ is the perceived distribution of resolvable uncertainty at the time of elicitation. $S \left({x_{it},\eta_i,\nu_i}\right) $ is $i$'s average perceived returns to choosing option 1 over option 0 at the time of elicitation, when the choice characterisitics are $x_{it}$. As above, $\nu_i$ is assumed to be univariate and additively separable.
Beyond the fact that they have different supports, there are two main differences between probabilistic stated choices and the actual choice. First, $F_{\nu \vert X_i, \eta_i}$ is not necessarily the same as the distribution of $\nu^r$, that is, individuals can be mistaken about the distribution of future realisations of the resolvable uncertainty. Second, $S(\cdot)$ and $S^r(\cdot)$ are not necessarily the same, that is, at the time of elicitation, individuals can be mistaken about their perception of returns at the time of decision, for example if they are subject to present bias or hyperbolic discounting.
The framework will not allow distinguishing between the two sources of bias without further assumptions. Furthermore, because of the binary nature of the choice, it will not be possible to make progress on the identification of $S^r(.)$ without a normalisation of the distribution on $\nu_i^r$. Therefore, without loss of generality, one can write:
where $m^r( x,\eta_i):= F_{\nu^r \vert \eta} \Bigl({S^r \left({x,\eta_i}\right)}\Bigr)$ and $U_i^r \vert \eta_i \sim \mathcal{U}[0,1]$. $m(x,\eta)$ is the take-up rate of or the objective demand function for option 1 for individuals with unobserved attributes $\eta$, when observed attributes are exogenously set to $x$. Similarly, define:
the take-up rate or the stated demand function in the stated choice analysis.
Both Equations ((ref)) and ((ref)) share one important feature: they both depend on $\eta_i$. Crucially, the unobserved heterogeneity affecting the actual choice needs to be a sub-vector of the unobserved heterogeneity influencing stated choices. To simplify the exposition, the same vector is used in both cases. The presence of $\eta_i$ in both stated and actual demand functions is, I propose, the key link between stated and actual choices, and one of the main goals in eliciting choice probabilities in hypothetical scenarios. In fact, the identification results proposed in the next section would remain valid with any survey experiment that would guarantee that (1) individual choices within the experiment are affected by the same unobserved heterogeneity vector $\eta_i$ as in the actual choice, and (2) this individual heterogeneity can be recovered from the survey experiment. The case for using stated preferences over the hypothetical scenarios mimicking the actual choice environment is both intuitive and compelling. Even if they only replicate actual choices imperfectly, carefully designed stated preference analyses induce respondents to consider the variables that influence their decision in a real environment. Ideally, these variables are also reflected in their stated choice. The next section shows how to recover the unobserved heterogeneity from elicited choice probabilities.
In both this and the next section, the issue of measurement error is intentionally omitted for simplicity; however, Section (ref) discusses identification with measurement error. Section (ref) below discusses the main identification result.
The difference $m^r(X_i,\eta_i) - m(X_i,\eta_i)$ measures the bias between the actual and the stated choice for individuals with choice characteristic $(X_i,\eta_i)$. More generally, one can define a mean-squared bias (MSB) as:
MSB provides a more thorough assessment of the accuracy of stated choice than comparing subgroups with observed characteristics $X_i=x$ via the difference $\mathbb{E}(D_i\vert X_i=x) - \mathbb{E}(P_{it} \vert X_i=x)$.
The analyst may also be interested in a counterfactual analysis in the spirit of chernozhukov2013. For example, assume there are two groups in the population $g_1$ and $g_2$. $m_g^r(\cdot,\cdot)$ is the demand function in group $g \in \{g_1,g_2\}$ and $F_{{X,\eta}\vert g}$ is the distribution of $(X_i,\eta_i)$ in group $g$. The analyst is interested in the counterfactual distribution:
which gives the demand that would prevail if members of group $g_1$ had the same distribution of characteristics, $(X_i,\eta_i)$, as members of $g_2$.
Finally, the analyst may be interested in features of the structural function $x \mapsto m^r(x,.)$, the objective demand function, such as the average structural function:
The key identifying assumption for this counterfactual quantity is the following:
Assumption (ref) imposes that the unobserved heterogeneity $\eta_i$ is the source of endogeneity of $X_i$ in a regression of $D_i$ on $X_i$. Controlling for $\eta_i$ would purge the estimated demand function from the omitted variable bias. Thus, Assumption (ref) requires that, when reporting their intended choice, respondents account for factors that would create a statistical dependence between observed choice attributes and the resolved uncertainty. For example, if both future unobserved shocks on the utility of a STEM degree and STEM degree attributes depend on individual's ability in math-related subjects, a respondent should factor-in their innate ability in reporting their intended choice.
In the rest of the paper, the analyst is assumed to elicit stated preferences in $T+1$ hypothetical scenarios. Hypothetical scenarios refer to exogenously set values of the variable $X$. The first scenario denoted $0$ serves for identification of the stated demand function. In this scenario, $X_{i0}$ spans the support of $X_i$. The stated demand function is needed for identification of $MSB(x)$ only. The remaining $T$ scenarios serve to learn about the unobserved heterogeneity $\eta$. The reader should think of them as either drawing random values of $X_i$, or setting a common counterfactual for all individuals.
To understand the intuition of identification, it is easier to proceed first with the case where $d:= \text{dim}(\eta) =1$. Assume further that the function $\eta \mapsto m(x,\eta)$ is strictly monotone. Given the flexibility for the researcher to analyse counterfactual scenarios, assume that $X_{i1}=\bar{x}$ for any individual $i$, that is, one counterfactual elicits the intended choice in a common hypothetical scenario. It entails that $P_{i1} = m(\bar{x},\eta_i)$. Because the counterfactual $\bar{x}$ is the same for all individuals, there is a one-to-one mapping between the stated choice and the unobserved heterogeneity. Observing $P_{i1}$ is equivalent to observing $\eta_i$. Hence:
where $\mathcal{H}$ is the support of $\eta$. The first equality is by definition, the second equality follows from the fact that the conditional distribution of $U^r$ is uniform. Equality (3) uses Assumption (ref). Equality (4) is by definition. Equality (5) follows from the fact that there is a one-to-one mapping between $P_{i1}$ and $\eta_i$ for all those such that $X_{i1}=\bar{x}$.
A similar result holds for the MSB:
and $F_{D\langle {j,k}\rangle}$:
In sum, the fundamental idea is to match individuals on their stated preferences. The logic is the same for an heterogeneity in higher dimensions.
Assumption (ref).(1) imposes that the analyst observes at least as many scenarios as the dimension of $\eta$. In hypothetical scenarios, observed choice attributes are either (i) randomly generated from a set support, or (ii) fixed for everyone in the population. Note that the main objective of these hypothetical scenarios is to learn about the distribution of unobserved heterogeneity and not necessarily about the stated demand function.
Assumption (ref).(2) is key. It imposes that given $\boldsymbol{X}\in \mathcal{C}^d$, such that $\textrm{rank}(X)=T=d$, a unique $\eta$ can be recovered from $\boldsymbol{P} = m(\boldsymbol{X},\eta).$ One example often used in empirical applications is a multiplicatively separable log-odd model, of which the model of blass2010 is a special case. It corresponds to:
where $\Gamma(x) = \exp(x)/(1 + \exp(x))$, and $K(\boldsymbol{X})$ is $d \times d$ matrix with full rank that spans $\mathcal {X}^d$. In this case, $\eta = K(X)^{-1} \left({\Gamma^{-1}(\boldsymbol{P}) - v(\boldsymbol{X})}\right)$.
Global invertibility conditions is used, for example, in matzkin2008 to show global identification for general nonparametric simultaneous equation models. A detailed derivation of sufficient conditions on individual preferences is beyond the scope of this paper. However, the next paragraphs offer some comments.
In the context of consumer choice, sufficient conditions on preferences for invertibility between demand and nonseparable individual heterogeneity are discussed, for example, in beckert2008. Their framework is akin to that of this paper, and the derived conditions rest on the theorems of gale1965 and mas1979 for the existence of (global) homeomorphisms.\footnote{beckert2008 consider the demand, $\boldsymbol{x}$ for $d$ goods, characterised by a vector of prices $\boldsymbol{p}$, with $y$ being the agent income, and $\eta$ being the unobserved heterogeneity. $\boldsymbol{p}$ and $y$ are independent of $\eta$. The reduced form demand is characterised by a system $\boldsymbol{x} = m(\boldsymbol{p},y,\eta)$.} Following their discussion, the homeomorphism property must be deduced from properties of the Jacobian, $\nabla_{\eta} m(\boldsymbol{X},\eta)$.
Take for example an agent who perceives a utility $U_j(x_j,\eta) + \nu_j$ in option $j \in \{0,1\}$, for a resolved uncertainty $\nu_j$, with $U_j$ having the usual continuity and strict concavity properties. The utility gain from option 1 is then: $U_1(x_1,\eta) - U_0(x_0,\eta) + \nu_1 - \nu_0$. The stated preferences over $d$ hypothetical scenarios can therefore be summarised by the system:
where $\nu:= \nu_0-\nu_1$, $F_{\nu \vert \eta} $ is the conditional distribution of the resolvable uncertainty as perceived by the agent, $S(\boldsymbol{X},\eta)=U_1(\boldsymbol{X}_1,\eta) - U_0(\boldsymbol{X}_0,\eta)$, and $\boldsymbol{X}= (\boldsymbol{X}_0, \boldsymbol{X}_1)$. The Jacobian matrix can be written:
where $f_{\nu \vert \eta} $ is the probability distribution function of the resolvable uncertainty. Assuming that $m$ is continuous in $\eta$, a Jacobian matrix with full rank is sufficient for a local homeomorphism, that is local invertibility. A local homeomorphism is a sufficient condition in the case where scenarios are set at fixed value for the whole population. In addition, if $\eta$ is an interior solution and the Jacobian has a positive determinant, then $m$ is a global homeomorphism mas1979.
When $\eta$ is unidimensional and the resolvable uncertainty does not depend on $\eta$, this requires $\eta \mapsto S(\boldsymbol{X},\eta)$ to be strictly monotone for any $X$. In other words, $\eta$ should have a monotonic effect on the differential of utility between option 0 and 1. Because the resolvable uncertainty may also depend on $\eta$, the monotonicity of $S$ in $\eta$ needs to be strong enough to compensate for any converse effect of $\eta$ on the distribution of the resolvable uncertainty. The global homeomorphism property generalises the monotonicity requirements.
The next proposition is the main result of paper.
The proof follows the same steps as Equation ((ref)), mutatis mutandis. Proposition (ref) summarises the main insight of the paper. By matching individuals on their stated preferences, the analyst can control for their unobserved heterogeneity. This allows assessing average forecast error by using a finer definition of heterogeneity or by performing counterfactual analyses. In addition, if conditioning on the unobserved heterogeneity ensures that the resolved uncertainty and the observed choice attribute are independent, the analyst can learn about causal effects of choice attributes.
The previous section ignored measurement error to simplify the exposition. It is possible to allow for some form of classical measurement error while preserving the main identification result. This section deals with measurement errors that arise from respondents reporting inaccurately their true intended choice, possibly due to inattention, misunderstanding the survey instrument, or lack of effort. Thus, the measurement error is viewed as a source of randomness across scenarios. The case where measurement error would be systematic over all scenarios amounts to consider an additional dimension of unobserved heterogeneity, say $\xi$. For the purpose of causal inference, Assumption (ref) should be adapted to state that $\nu_i^r \protect\mathpalette{\protect\independenT}{\perp} X_i, \xi_i \vert \eta_i$.
Assumption (ref) summarises the structure of the data available to the researcher.
$P_{it}^*$ deviates from the true stated choice because of the measurement error $\epsilon_{it}$. Assumption (ref) imposes that $\epsilon_{it}$ is a classical measurement error, that is, independent across scenarios, unrelated to $\eta_i$, but possibly related to $X_{it}$. The main identification idea is to use kotlarski1967's Lemma to show that the joint distribution of $(D_i,X_i,X_{i1},\ldots, X_{it}, P_{i1}, \ldots, P_{it})$,is identified from the joint distribution of $(D_i,X_i,X_{i0},X_{i1},\ldots, X_{it}, P_{i0}^*, P_{i1}^*,\ldots,P_{it}^*)$. The argument is reminiscent of evdokimov2010 for the case $d=1$ and meango2023, for the case where $d>1$.
Suppose that $d=1$, $t \in \{0,1\}$. The main goal is to show that the conditional distribution of $P_{i1}$ given $(D_i, X_i, X_{i1})$ is identified from the data. Indeed, since the joint distribution of $(D_i, X_i, X_{i1})$ is identified from the data, by Bayes rule, one can recover the quantity $\mathbb{E}(D_i\vert X_{i}, X_{i1}, P_{i1})$ that is needed for identification (see Equation ((ref))). Note that for any $i$ and an arbitrary $x$ such that $X_{i0} = X_{i1} = \bar{x}$:
Lemma 1 of evdokimov2010 implies that the conditional characteristic distribution of $\epsilon_{it}$, say $\phi_{\epsilon_{t} \vert X_{it}}(s \vert \bar{x})$ is identified from the joint distribution of $(X_{i0},P_{i0}^*,X_{i1},P_{i1}^*)$ on the set such that $\{x \in \mathcal{X}: X_{i0}=X_{i1}=x\}$. Assume for simplicity that this set covers $\mathcal{C}$.\footnote{The flexibility of belief elicitation means that the researcher can purposefully induce a repetition of the same scenario at any point of the support.} Given identification of the characteristic function of measurement error, we can identify the characteristic functions:
Knowledge of the characteristic function being equivalent to knowledge of the distribution function, this shows the result for $d=1$. For the case $d=2$, $T=2$, one can repeat the same steps and note that: \[
\] The same development can be repeated for $d>2$ mutatis mutandis. Proposition (ref) summarises the result and updates Proposition (ref) for the case of measurement error.
Proposition (ref) shows that even a small number of scenarios allows identifying the joint distribution of unobserved heterogeneity. $MSB(x)$ is identified on $x \in \mathcal{C}$, because $P_{i0}^*$ is observed with error. The error can be corrected only on the set such that: $\{x \in \mathcal{X}: X_{i0}=X_{it}=x\} = \mathcal{C}$. Thus, the joint distribution of $P_0,X_i$ is only identified on $\mathcal{C}$. Identification on the full support requires $\mathcal{C}= \mathcal{X}$.
The identification result being constructive, it can serve for a three-step estimation as in evdokimov2010 or meango2023: first, a deconvolution estimator for the joint distribution of $\{P_{t}: t = 1, \ldots,T\}$ from the joint distribution of $\{P_{t}^*: t = 0,1, \ldots,T\}$. Second, an estimation of the joint distribution of $(D_i,X_i, P_1, \ldots, P_T)$. Finally, estimation of the parameters of interest by integration. However, these steps can be involved and practitioners are unfamiliar with such tools. The next section proposes instead to assume that $T$ is large relative to $d$ and use the Two-Step Group Fixed Effects estimator proposed by bonhomme2022.
This section proposes to use the Two-step Group Fixed-Effect (TSGFE) methodology of bonhomme2022 for estimating $MSB(\cdot)$, $F_{D\langle{\cdot,\cdot}\rangle}$ and $\mu_r(\cdot)$, with data that satisfy Assumptions (ref) and (ref). Under the latter, their crucial Assumption 2 of existence of injective moments is satisfied, for $m(\cdot,\cdot)$ being Lipschitz-continuous and the variance of $\epsilon_{it}$ being finite. The procedure consists in two steps: the first step involves a classification that uses auxiliary moments to classify individuals into types, based on a finite (low) dimensional unobserved heterogeneity. The second step uses the estimated classification to estimate the parameters of interest. In the setting of this paper, TSGFE is a natural approach to exploit the link between probabilistic stated choices and actual choices through the unobserved heterogeneity $\eta$. This section assumes that $T>\dim(\mathcal{X})$ is large.
\paragraph{First-step: Classification by kmeans clustering.} The procedure relies on individual-specific moments $h_i$. Identification in our framework requires at least $d$ moments. It is as follows:\\ Estimate consistently:
for each individual $i$; generate (at least) $d$ moments:
where the rank of the matrix that stacks the vectors $\bar{x}_j$ is (at least) $d$; finally, partition individuals into $K$ groups, corresponding to group indicators $\widehat{k}_i \in \left\{{1,\ldots,K}\right\}$ by computing
where $\{k_i\}$ are partitions of $\{1,\ldots,N\}$ into $K$ groups, and $\tilde{h}(K)$ is a vector of dimension $d$. bonhomme2022 provide a data-driven selection rule for the choice of $K$.
\paragraph{Second-step: (Parametric) estimation.} The TSGFE uses a parametric form for Equation ((ref)). Assume that there exists a finite dimension parameter $\theta_0$, such that $m(X_i,\eta_i) = m(X_i,\eta_i; \theta_0)$. In the second step, the estimator maximises the likelihood function $\ln \Pr(D_i \vert X_{i}, \eta_i;\theta_0)$ with respect to the common parameter $\theta$ and group-specific effects, that is:
Estimates of $\mu^r(\cdot)$, $F_{D\langle{\cdot,\cdot}\rangle}$ and $MSB(\cdot)$ follow naturally by taking empirical means. For example, if $\widehat{p}_k$ is the proportion of individuals in group $k$, the estimator for $\mu_r(x)$ is given by:
bonhomme2022 discuss regularity conditions and asymptotic properties of TSGFE in appropriate length. Appendix (ref) provides simulation results that demonstrates the usefulness of the method in our context.\footnote{A routine easily adaptable is also available at the following \href{https://www.dropbox.com/scl/fi/ri53cxa76lt5bkgd7qowy/simulation_RM.m?rlkey=voz2bz979rq1tgjh21pv1zbeg&dl=0}{link} }
The proposed strategy is of practical importance to empirical researchers for two types of analyses: (1) causal inference, and (2) heterogeneity analyses.
Causal inference often requires exogenous variations that may not be available in some contexts. Even when a suitable instrument is available, it recovers treatment effects for sub-populations that may differ from that of the entire population. One example is a situation where the researcher would only have access to administrative data, which records individual choices and a limited set of observed characteristics. The above strategy suggests complementing these data with a survey experiment, where choice probabilities are elicited and the main treatment of interest is varied along with other relevant characteristics. Provided that the unobserved heterogeneity recovered is the main source of bias, the analyst can estimate (marginal) treatment effects for the entire population. The latter hypothesis is of course untestable, as are the critical assumptions of the other existing strategies. Still, the researcher can test whether the survey experiment captures all the information on the unobserved heterogeneity available from the stated choice (See Appendix (ref)). If a surrogate for an experiment exists (exogenous variation or discontinuity at a threshold), the researcher can validate the obtained results using the stated preference approach. It suffices to compare them with the results from using the surrogate experiment, on the sub-population where the latter permits identification of a causal effect.
The discussion above implies that the proposed strategy should not necessarily be viewed as a substitute for existing strategies, but as a complement. The ability to recover unobserved heterogeneity is valuable to (i) increase the credibility of an instrument, (ii) control for confounding factors that are not necessarily separable from treatment status, or (iii) assess balance across a threshold. In the case of randomised control trials, identifying the unobserved heterogeneity may help to understand treatment effect heterogeneity, or improve the precision of the estimator. Hence, even in situations where usual inference strategies are available, a stated preference analysis may aid inference. Thus, stated preference can complement and enhance the applied researcher toolbox.
To the best of my knowledge, two recent contributions exploit the idea that a stated preference analysis can help in recovering some information on individual heterogeneity and performing causal inference: briggs2020 and bernheim2022. briggs2020 adapt the generalized Roy model to show that marginal treatment effects are identified when data on stated preferences are available. Translated into our framework, $D$ would be a sector choice (the treatment), and the effect of interest would pertain to a third variable $Y$. Their framework imposes additive separability of $\eta$ and $X$, and, in allowing $\eta$ to be considered as univariate, they correct for the selection bias from the observation of one single intended choice. However, point identification breaks down in the case of measurement error. This paper complements their work and shows that, with repeated observations, matching on a range of stated preferences corrects for the selection bias, even in the presence of nonseparable heterogeneity and measurement error.
bernheim2022 characterises the relation between stated preferences and realised outcomes. Under the assumption that this relation is stable across treatment, the treatment effect revealed from the stated preference analysis can be used to infer the treatment effect on the actual choice. A key advantage of their methodology is that it does not require jointly recording choices and stated preferences in the same population. Furthermore, stated choices can be binary and $\eta$ can remain unrestricted. The key exclusion restriction is that treatment can only affect actual choice through the stated preference. In contrast, the assumptions in this paper on the relationship between stated and actual choice are significantly milder. The main requirement is for the unobserved heterogeneity in the actual choice to be a subset of the unobserved heterogeneity in the choice experiment. Thus, the strategy used here complements their approach in the case where the exclusion assumption fails.
Stated choices can serve to understand actual behaviour, not by matching actual choices, but rather by providing valuable information on individual heterogeneity. This paper shows how to harness stated preference analyses to recover the distribution of individual unobserved heterogeneity (i.e. agent's types). Once recovered, the types serve to evaluate biases in prediction for subgroups. If they are the main source of endogeneity, their introduction corrects for biases in demand function estimations.
Up to now, stated and revealed preference analyses have mostly been conducted separately due to the strong assumptions needed to combine them. The hope is that this research will convince practitioners that stated preference analyses can become part of their toolbox for valid counterfactual analyses of actual choices.