Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,503 characters · 23 sections · 58 citation commands
Identifying causal effects with subjective ordinal outcomes
Many survey questions ask respondents to choose from a set of two or more ordered categories that lack clear definitions, leaving the interpretation of those categories to the respondent. Examples include self-reported health status (SRHS), product or service ratings, job satisfaction, and questions gauging satisfaction with life overall. Individuals' responses are then often used as an outcome variable in research, frequently as a proxy for some underlying latent variable of interest (e.g. true health in the case of SRHS).\footnote{A broad class of this type of survey questions that use so-called Likert scales: e.g. allowing responses such as “strongly agree”, “agree” \dots “strongly disagree” to indicate agreement with a given statement, or to categorize quantities such as frequencies (“often”, “sometimes”, \dots “almost never”). hamermesh2004 discusses the use of such outcomes in economics.}
A key question for this practice is how “reporting functions”---the way that individuals map that latent variable into one of the available response categories---impact conclusions drawn from the data.\footnotemark bondandlang influentially show that even if individuals share a common reporting function (but it is not ex-ante known to the researcher), averages of the latent variable cannot be meaningfully compared between groups using their responses, absent strong restrictions on the latent variable's unobserved distribution. More fundamentally, if the response categories lack objective definitions, reporting functions might vary between individuals, potentially confounding any attempt to study relationships between explanatory variables and the latent variable.
This paper shows that the observed categorical responses can nevertheless be informative about causal relationships in which this latent variable is the outcome, despite the dual threats of reporting functions being both i) unknown to the researcher and ii) heterogeneous across respondents. Taking the perspective of bondandlang that the latent variable driving individuals' responses is the researcher's ultimate outcome of interest, I decompose differences in the observed joint distribution of responses and covariates to the causal effects of those covariates on the latent variable. I do so by strengthening the familiar selection-on-observables assumption that one or more explanatory variables are statistically independent of potential outcomes, adding to it that explanatory variables are also independent of heterogeneity in reporting functions (with both independence assumptions made conditional on observed control variables). Under this assumption I show how the estimand that arises from the common practice of regressing categorical response numbers on explanatory variables can be interpreted in terms of the causal effects of those regressors on the latent variable of interest. \footnotetext{The use of the term “reporting function” for subjective data appears to have first appeared in the economics literature in oswaldletters, though the general concept predates its discussion in economics (e.g. psycho).}
Concretely, I consider a general model of ordered response taking the form:
where $H_i \in \mathbb{R}^K$ reflects a set of unobserved latent variables, and $R_i$ an observed response mapped to a real number in some set $\mathcal{R}$. For example, $\mathcal{R} = \{0,1\}$ for a binary yes/no question, or $\mathcal{R} = \{0,1,2,3,4\}$ for a question with five ordered response categories. I focus primarily on the case of a scalar latent variable $H \in \mathbbm{R}$, and later generalize to $K > 1$.
The function $h_i(x)$ in (ref) denotes the potential outcomes of the latent variable for individual $i$, indicating the value of $H$ that would occur if a vector of observed explanatory variables $X$ took each counterfactual value $x$. The function $r_i(h)$ represents individual $i$'s reporting function, which I assume to be weakly increasing in $h$ for each $i$. The random vectors $U_i$ and $V_i$ parameterize heterogeneity across individuals in potential outcomes and reporting functions, respectively. The main statistical assumption of the model is that $X_i \perp\!\!\!\!\perp (U_i,V_i)$, which I relax to conditional independence given control variables. The researcher's objective is to learn how $h_i(x)$ varies with $x$, observing only $R_i$ and $X_i$.
One of the key implications of my results is that if $X_{1i}$ and $X_{2i}$ reflect two continuously distributed components of the vector $X_i$, and $\mathcal{R}$ is associated with a set of integers, then
where $\tilde{\beta}_j$ reflects a convex weighted average across individuals of the causal effect of a small change in the $j^{th}$ component of $X$ on $H$. In particular, $\tilde{\beta}_j$ averages the causal partial derivative $\partial_{x_j} h(X_i,U_i)$ over individuals $i$ who are on the margin between two response categories $r-1$ and $r$ for any $r \in \mathcal{R}$.\footnote{$\partial_{x_j} h(X_i,U_i)$ denotes $\partial_{x_j} h(x,U_i)$ with $x_j$ the $j^{th}$ component of $x$, evaluated at $x=X_i$ (and similarly for $\partial_{x_j}\mathbbm{E}[R_{i}|X_i]$)} If the conditional expectation $\mathbbm{E}[R_i|X_i]$ happens to be linear, then the average derivative quantity $\mathbbm{E}[\partial_{x_j}\mathbbm{E}[R_{i}|X_i]]$ in the LHS of (ref) is simply the coefficient on $X_{j}$ in a linear regression of $R$ on $X$. In this case, Eq. (ref) affords a causal interpretation to the ratio of OLS regression coefficients for two continuous treatments.
Throughout the paper, I discuss results through an application to survey questions that ask respondents about their overall satisfaction with life, and for ease of exposition refer to the latent variable $H$ as “happiness”.\footnote{This simplified language ignores e.g. distinctions between hedonic, affective and evaluative notions of well-being DEATON201818,helliwellcpbl.} For example, the popular Cantril Ladder question asks individuals to describe their satisfaction with life on an eleven point scale from $0$ to $10$.\footnote{A popular version of the Cantril ladder question asks: Please imagine a ladder with steps numbered from zero at the bottom to ten at the top. Suppose we say that the top of the ladder represents the best possible life for you and the bottom of the ladder represents the worst possible life for you. If the top step is 10 and the bottom step is 0, on which step of the ladder do you feel you personally stand at the present time? gallup.} Questions like this about general well-being motivate treating the latent variable $H$ as an outcome of normative interest, drawing on the notion of cardinal utility as a measure of welfare Fleming1952,Harsanyi1955. With this interpretation, the marginal rates of substitution between treatment variables are a key input for normative analysis, suggesting trade-offs that would be welfare improving for individuals. However, my results are also applicable to other outcomes elicited on ordered scales, e.g. general or mental health status, job satisfaction, product or service ratings, and other settings in which ordered response models might be employed with individual-specific heterogeneity in the thresholds between response categories.
Despite a growing trend in papers leveraging natural experiments with subjective outcome data,\footnote{Some prominent examples include cardetal,aermrs,lindqvistetal2020,aerperez,pnasdwyerdunn.} empiricists have lacked formal results such as Eq. (ref) to interpret precisely what is estimated by regressions in which subjectively-defined ordinal responses $R$ are used as the dependent variable. This paper helps to fill the gap by showing that when the selection-on-observables research design is extended to include reporting-function heterogeneity, derivatives of the conditional expectation function of integer category numbers on $X$ reveal positive aggregations of the local causal effects of $X$ on $H$.\footnote{I also show that when the researcher is interested in establishing correlations rather than causation, the same results capture changes to the conditional quantile function of the underlying latent variable, rather than causal effects.} The weights in this aggregation have an intuitive form but are not under the researcher's control. This illuminates the limits for identification of overall unweighted means of causal effects, which correspond to the parameter analyzed by bondandlang. My results show that mean regression can nonetheless remain a useful tool for analyzing more general weighted averages of effects, without assuming cardinality or interpersonal comparability of $H$.
I apply my formal results to revisit the influential study of luttmer2005, who considers the effects of household income as well as the incomes of one's neighbors on satisfaction with life. Using a selection-on-observables strategy and linear regression adjustment, luttmer2005 finds a positive coefficient on own-income along with a negative coefficient on neighbor income, suggesting that relative income comparisons are important for subjective well-being. My nonparametric identification results corroborate this interpretation under the maintained exogeneity assumptions, but without assuming cardinality or interpersonal comparability of individuals' responses to the well-being question. Empirically, I first report distributional regressions of $\mathbbm{1}(R_i \le r)$ on $X$ for each $r$. The patterns suggest that differences in the coefficients across $r$ are driven by the unknown distribution of the underlying latent variable, underscoring the theoretical observation that coefficients must be compared between variables to be quantitatively meaningful. The heterogeneity in coefficients across $r$ in fact cancels in the ratio, and I cannot reject equality across $r$ of the local marginal rates of substitution between own and neighbor income (among respondents on the threshold between $r$ and $r+1$). I also estimate these “marginal” respondents to be similar to inframarginal respondents in terms of gender and education. The empirical results overall are consistent with simple models of heterogeneity in potential outcomes and/or response functions that render the effects for marginal respondents somewhat typical of the population, in this particular setting.
When the treatment variables of interest are discrete, rather than continuous as above, I find that comparisons of magnitude become more complicated. First, I show that when one compares the mean of $R$ between two fixed values $x$ and $x'$ of the vector $X$:
where $\Delta_i = h(x',U_i)-h(x,U_i)$ is the treatment effect of changing $X$ from $x$ to $x'$ on outcome $H$ for individual $i$. The “weight” $\bar{f}(\Delta_i,v_i,x_i)$ is unknown but positive for all $i$, and Eq. ((ref)) thus implies that if the sign of the treatment effect $\Delta_i$ is the same for all individuals, then the sign of $\mathbbm{E}[R_i|X_i=x']-\mathbbm{E}[R_i|X_i=x]$ will be the same as that of the causal effect. However, the magnitude of $\mathbbm{E}\left[\bar{f}(\Delta_i,V_i,x)\right]$ can in general depend on the values $x$ and $x'$ being compared, and quantitative comparisons of regression coefficients can be misleading (even if the regression is correctly specified) if one or more of the treatment variables being considered is discrete and treatment effects are not small.\footnote{The function $\bar{f}$ is defined in Sec. (ref), and no longer depends upon $\Delta$ as $x' \rightarrow x$ and the difference becomes a derivative.}
Eq. (ref) provides a new perspective on the key point made by bondandlang, who argue that the conditional distributions $R_i|X_i=x'$ and $R_i|X_i=x$ are generally uninformative about the sign of $\mathbbm{E}[H_i|X_i=x']-\mathbbm{E}[H_i|X_i=x]$, even if it is assumed that all individuals share a common reporting function. Knowing the sign of the difference in means of $H_i$ between groups $x$ and $x'$ from observations of $R_i$ generally requires that the quantile functions of $H_i|X_i=x'$ and $H_i|X_i=x$ do not cross, i.e. that the latent variable distribution for one group stochastically dominates that of the other. This assumption cannot be verified from the data $(R,X)$, and may be implausible if $x$ and $x'$ are two populations (men vs. women, two countries, etc.), that are each quite heterogeneous in themself and differ from one another across many dimensions. However, when $x'$ and $x$ differ in a single component representing treatment in e.g. a quasi-experimental setting, it may be possible to argue that treatment effects $\Delta_i$ are not too heterogeneous.\footnote{Indeed, the much stronger assumption of complete homogeneity in treatment effects is often made implicitly to motivate a causal interpretation of regression models. For example, $h(x,u) = g(x)+u$ yields the regression $H_i = g(X_i)+U_i$ but implies that $\Delta_i = g(x')-g(x)$ for all $i$.} If $\Delta_i$ has the same sign for all $i$, that sign is equal to that of $\mathbbm{E}[R_i|X_i=x']-\mathbbm{E}[R_i|X_i=x]$. Thus while the argument made by bondandlang is compelling for generic comparisons between two groups, it may have less bearing on settings where a clear research design is leveraged to interpret differences in $R_i$ causally.
Implications of my results for regression analysis using subjective ordinal outcomes are threefold. First, the focus on finding natural experiments popular in modern applied work yields a previously unrecognized benefit for subjective outcomes: reporting functions may become uncorrelated with treatment variables of interest $X$, affording inference on the direction of causal effects of $X$ on unobserved $H$. Second, researchers can move beyond interpretations of the sign of average effects and consider magnitudes only when multiple valid treatment variables are available. Third, such comparisons of magnitude are most informative when the two variables being compared are continuous rather than discrete. An implication is that identification in experimental work with subjective outcome variables would benefit from randomizing the quantitative “doses” of multiple treatments.\\
Outline of paper: In Section (ref) I propose a general nonparametric model of ordered response with nonseparable heterogeneity: it allows each respondent to have their own response function, but takes treatment variables $X$ to be conditionally independent of all unobserved heterogeneity. Section (ref) establishes my main identification result when there is continuous variation in $X$, which provides a generalization of Eq. (ref). I outline assumptions under which Eq. (ref) in turn reveals a local average marginal rate of substitution between two continuous treatment variables. Section (ref) applies these results to revisit the findings of luttmer2005 relating the effect of one's own income and one's neighbors' incomes on life satisfaction.
In Section (ref), I turn to identification with a discrete treatment variable. After showing that ratios of regression coefficients involving one or more discrete regressors lack the guarantee of a simple quantitative interpretation like Eq. (ref), I describe how one can obtain bounds on the ratio of the total weight that the conditional expectation function applies to causal effects when comparing continuous to discrete variation in $X$. The analytic results suggest that when there are many response categories and individual reporting functions are approximately linear, discrete contrasts will tend to overstate causal effects relative to regression derivatives, by a factor that is upper bounded by two. I assess this implication through simulations with a variety of assumed distributions of the latent variable, and only find evidence of appreciable distortion when treatment effects are made implausibly large in the DGP.
Appendix (ref) provides an extended discussion of how my results relate to bondandlang. Appendix (ref) relates my general model of ordered response to ones previously considered in the literature. Appendix (ref) considers several extensions to my baseline model, such as using instrumental variables rather than selection-on-observables for identification, or allowing for a multivariate latent variable. Appendices (ref) and (ref) develop some supporting theoretical results for the paper. Appendix (ref) provides additional results for the empirical application, while Appendix (ref) expands on the implications of my results for practical regression analysis and presents a numerical illustration.
Suppose that there exists a meaningful latent value $H_i$ for each individual which the researcher is ultimately interested in as an outcome. With the life satisfaction example in mind, I will often refer to $H_i$ as $i's$ underlying “happiness”, which the researcher aims to learn about given those individuals' responses $R_i$.\footnote{The model extends naturally to a setting in which the definition of “$H$” is itself subjective, in the sense that different individuals use different latent variables when constructing their responses. The key requirement is that these subjectively defined latent variables in turn reflect increasing transformations of an objective variable of interest. See Appendix (ref).} Section (ref) discusses the interpretation of $H_i$ as a measure of utility. In the body of this paper I take $H_i$ to be a scalar, but Appendix (ref) extends results to the vector case.
The researcher observes a sample of $(R_i,X_i,W_i)$ across individuals $i$ generated as:
where $r_i(h)$ is in individual-specific function mapping happiness $h$ to the space of possible responses $\mathcal{R}$. The above model indexes heterogeneity in $r_i(\cdot)$ by a heterogeneity parameter $V_i \in \mathcal{V} \subseteq \mathbbm{R}^{d_V}$. Since no constraints are placed on $d_V$, this is without loss of generality and the model is compatible with each individual having their own reporting function $r_i(h)$. Figure (ref) depicts two examples of reporting functions when $\mathcal{R} = \{0,1,2\}$.
For each individual there is a function $h_i(\cdot)$ mapping values of a vector of $J$ explanatory variables $X$ into a value of $H$ via ((ref)), where heterogeneity in the function $h_i(\cdot)$ is represented by parameter $U_i \in \mathcal{U} \subseteq \mathbbm{R}^{d_U}$. The primary interpretation of the function $h_i(x)$ is that it denotes potential outcomes for individual $i$ as a function of counterfactual values of $x$, in some set of possible treatments $\mathcal{X} \subseteq \mathbbm{R}^{J}$.\footnote{An alternative interpretation of $h(x,u)$ is always also available and requires no causal assumptions, which is that $h$ represents the conditional quantile function of $H_i$ given $X_i$, with $U_i \in [0,1]$ a scalar indicating $i$'s rank in a distribution of their peers. In particular, let $\theta_i:=F_{H|XVW}(H_i|X_i,V_i,W_i)$ be $i$'s “rank” in the conditional happiness distribution of individuals sharing their value of $X,V$ and $W$, where $F_{H|XVW}$ denotes a cumulative distribution function of $H$. Now let $U_i = (\theta_i, V_i,W_i)^T$, and define $h(x,u):=Q_{H|XVW}(\theta|x,v,w)$ for any $u=(\theta,v,w)^T$, where $Q_{H|XVW}$ denotes the conditional quantile function of $H$ given $X,V$ and $W$. Eq. ((ref)) now follows from these definitions. See Appendix (ref) for details. This representation is helpful when causal effects are not the target, and the researcher is instead interested in uncovering statistical features of the joint distribution between $H_i$ and $X_i$.} Since the dimension $d_U$ is again left unrestricted, the above model places no restriction on heterogeneity in potential outcomes and causal effects across individuals.
Finally, $W_i \in \mathcal{W} \subseteq \mathbbm{R}^{d_W}$ is a vector of additional observed variables to be used as control variables in the analysis. These $W_i$ can be thought of as variables that matter for happiness but are not necessarily manipulable (e.g. race), as components of $U_i$ that are potentially correlated with $X_i$ but are observable, or as correlates of $X_i$ that proxy for reporting function heterogeneity $V_i$. In settings with stratified randomization, $W_i$ isolates the experimental strata.
The function $h(x,u)$ is our main object of interest: how it varies with $x$ holding $u$ fixed yields the causal effect of that change on $H$. For example, $h(x',U_i)-h(x,U_i)$ denotes the “treatment effect” for unit $i$ of moving between two counterfactual values $x$ and $x'$ of the vector $X$. I consider the identification of such discrete treatment effects in Section (ref).
For most of the analysis, I will consider small changes in one or more components of $x$ that are continuously distributed. Letting $\partial_{x_j}$ denote a partial derivative with respect to $x_j$, the function $\partial_{x_j} h(x,U_i)$ for a given individual $i$ characterizes the effect of a small change in the $j^{th}$ component of $X$ on $H$, when $X=x$. An average of the value of this derivative across individuals $i$ provides a summary of the marginal effect of $x_j$ on $H$ when $X=x$. More generally, such averages can employ weights $\rho_i$ that depend on the individual-level observables $(X_i,W_i)$ and unobserved heterogeneity parameters $(U_i,V_i)$. For example, for a given function $\rho(u,v,x,w)$, we might consider a weighted average of the form:
where $\rho$ is chosen such that $\rho_i:=\rho(U_i,V_i,X_i,W_i)$ is positive with probability one and satisfies $\mathbbm{E}[\rho_i] =1$. Many results of this paper represent, intuitively, limits of parameters of the form $\tilde{\beta}_j$ for a sequence of such weighting functions $\rho(\cdot)$.\footnote{For example, the average derivative $\mathbbm{E}[\partial_{x_j} h(x,U_i)|h(x,U_i)=h]$ that conditions on a single value $h$ for $h_i(x)$ represents the limit of $\tilde{\beta}_j$ for the function $\rho(U_i,V_i)=\frac{\mathbbm{1}(h(x,U_i) \in [h, h+\epsilon])}{\mathbbm{E}[\mathbbm{1}(h(x,U_i) \in [h, h+\epsilon])]}$, as $\epsilon \rightarrow 0$. See also discussion after proof of Theorem (ref).}
If we interpret $H$ as a measure of “utility”, then $h_i(\cdot)=h_i(\cdot,U_i)$ can be thought of as $i$'s utility function, and $H_i=h_i(X_i)$ as their realized utility (evaluated at $i$'s actual $X_i$). Under this interpretation the ratio of two derivatives of $h_i(x)$ represents a local marginal rate of substitution of $X_1$ for $X_2$ when $X=x$, for individual $i$, e.g. $$MRS_i(x):=\frac{\partial_{x_2}h(x,U_i)}{\partial_{x_1} h(x,U_i)}$$ Note that this interpretation only requires $h_i(\cdot)$ to represent utility in an ordinal sense: $MRS_i(x)$ yields the slope of the indifference curve for $i$ that passes through the point $x$. Among individuals for whom $X_i=x$, the quantity $MRS_i(x)$ yields a marginal rate of substation at their actual value of $X_i$. Weighted averages of this realized marginal rate of substitution across individuals take the form $\widetilde{MRS}:=\mathbbm{E}\left[\rho_i\cdot MRS_i(X_i)\right]$ for $\rho_i=\rho(U_i,W_i,X_i,W_i)$ defined as following Eq. (ref), or a limit of $\widetilde{MRS}$ for a sequence of such functions.
Finally, this paper will consider weighted averages of discrete treatment effects between two fixed values of $X$, i.e. $\Delta_i:=h(x',U_i)-h(x,U_i)$ for some $x,x' \in \mathcal{X}$. Weighted averages of treatment effects take the form: $$\widetilde{\Delta}:=\mathbbm{E}\left[\rho(U_i,W_i,X_i,W_i)\cdot \Delta_i\right]$$ with $\rho_i:=\rho(U_i,V_i,X_i,W_i)$ as above, or the limit of $\widetilde{\Delta}$ for a sequence of such functions $\rho$.
Note that model ((ref))-((ref)) embeds an exclusion restriction: $X$ does not directly enter in the equation for $R$, and only affects reports through $H$. This is important for drawing inferences about the relationship between $H$ and $X$ from the observable joint distribution of $R$ and $X$. The model can be generalized slightly to allow reporting behavior to depend directly on observables, as described in Appendix (ref).
The following two subsections introduce the two key identifying assumptions of the model: first, that reporting functions are weakly increasing in $h$; and second, that the researcher as exogenous variation in some components of $X_i$. These assumptions are, under suitable regularity conditions, sufficient for the main results of this paper. The basic model is therefore more general than existing models of ordered response, which typically couple parametric assumptions with an assumption that there is no heterogeneity in $v$. Appendix (ref) shows how the model nests models previously considered in the literature.
The main assumption that I make about the reporting functions $r_i(\cdot)$ themselves is that they are increasing in $H_i$:
Appendix (ref) extends Assumption MONO to the case in which $H_i$ is a random vector, assuming that $r_i(\cdot)$ is weakly increasing in each component of $H_i$. Note that MONO does not assume the effect of $X$ on $H$ to be monotonic or uniform across individuals.
The first part of Assumption MONO rules out cases in which individuals would report a lower value of $R$ if $H$ were increased. The left-continuity assumption of MONO is essentially a normalization, since any weakly increasing function of bounded variation is continuous except at isolated points within its support.\footnote{Hence a reporting function that is, say, right continuous rather than left continuous could be made left continuous by modifying the function on a set of Lebesque measure zero.}
The following lemma shows that Assumption MONO is equivalent to there being a set of “thresholds” $\tau_v(r)$ that separate the ordered categories in $\mathcal{R}$. This characterization is useful in developing the formal results to come.
As an illustration of Lemma (ref), suppose that $\mathcal{R} = \{0,1,\dots \bar{R}\}$ for some integer $\bar{R}$. Then Lemma (ref) implies that any given reporting function $r(h,v)$ can be written as:
Remark: Assumption MONO does not require that respondents are motivated only by “honesty” when choosing $R_i$. Instead, they may have direct preferences for certain response categories. Consider a utility maximization model in which $ r(h,v) = \textrm{argmax}_{r \in \mathcal{R}} \textrm{ }u(r,h,v)$, with utility $u$ depending not only on happiness $h$, but also directly on the response category $r$. As an example, let us further assume that the utility function takes the form $u(r,h,v) = \phi_v(r)-|h^*_v(r)-h|$ where individuals of type $v$ obtain utility $\phi_v(r)$ from giving a response of $r$, but also value giving an answer close to a value $h^*_v(r)$ they perceive to correspond to response $r$. Provided that $h^*_v(r)$ is strictly increasing in $r$ (i.e. higher responses are subjectively associated with higher values of happiness), then $u$ satisfies the property of increasing differences (cf. shannonmilgrom) in $(r,h)$, which in turn implies MONO.\footnote{Note that heterogeneity $v$ in this form for utility need not be additively separable from quantities that depend on $x$ (i.e. $h$). Such separability is shown by allenrehbeck to admit important identification results for latent utility.}\footnote{MONO also allows there to be individuals with preferences that only depend on $r$, giving the same response regardless of their $H_i$. Such individuals will not contribute to regression derivatives and differences of $R$ on $X$ under EXOG.}
The final piece of the model is a conditional independence assumption for variation in $X$. In particular, I suppose that conditional on $W$, the treatments $X$ are as-good-as-randomly assigned in the following sense:
A sufficient condition for Assumption EXOG is that:
Eq. ((ref)) provides a natural foundation for EXOG and is simpler to motivate, but is technically stronger than the results require.\footnote{Eq. ((ref)) can equivalently be expressed as $\{(U_i,V_i) \perp\!\!\!\!\perp X_{j,i}\}| (X_{-j,i},W_i)$ for all $j$, where $X_{-j,i}$ denotes the elements of $X_i$ apart from $X_{j,i}$. Assumption EXOG can be re-expressed similarly.} For causal inference, an assumption like $\{X \perp\!\!\!\!\perp U\}|W$ is generally already necessary for identification: one needs some kind of experiment or natural experiment providing exogenous variation in $X$. Eq. ((ref)) then simply requires this natural experiment to also render $X$ (conditionally) independent of $V$. Note that under EXOG, $U$ and $V$ may be arbitrarily correlated with one another (e.g. if happier individuals have more optimistic reporting functions).\footnote{This is a feature that distinguishes my approach from the treatment of measurement error by hausmanabreyava, who assume (in my notation) that $R \perp\!\!\!\!\perp X | H$, which amounts to $V \perp\!\!\!\!\perp U | H$. They also restrict the model functionally, with a linear index structure for $h$ and scalar errors with monotonicity.} In Appendix (ref), I relax EXOG to consider identification using instrumental variables. Appendix (ref) illustrates through an example how violations of EXOG can affect results.
The assumption that response behavior is independent of a treatment variable may be restrictive in many contexts, especially in the absence of a credible research design. Appendix (ref) describes one specific threat to Assumption EXOG, that reporting functions might themselves be affected by the treatment variables $X$. I show there that EXOG can be relaxed slightly, and in fact tested under additional structural assumptions. Whether reporting functions might themselves be affected by a given treatment variable must be considered on a case-by-case basis.\footnote{Other approaches to allowing for reporting-function heterogeneity that do not require EXOG rely on particular models of that heterogeneity (e.g. CPBL) or auxiliary data sources. Such sources include “anchoring vignettes” kingetal, kapetyn,molina,jeboreversing,stantchevaannrev, memories of past life satisfaction kaiserjebo, calibration questions adjustingforscaleuse and survey response times happytimes.}
Given the model outlined in the last section, let us consider what can be identified by looking at responses given variation in $X$. In this section, I suppose that at least one component of $X$ is continuously distributed.
Denote by $f_H(h|x,v,w)$ the density of $H_i$ at $h$, conditional on $X_i=x$, $V_i=v$ and $W_i=w$, and assume the following:
Assumption REG reflect standard regularity conditions, as described in hoderleinmammen2007. The only substantive modification above is that I take the conditions to hold conditional on each reporting function type $V_i=v$.
Let $P(R_{i} \le r|x,w):=P(R_{i} \le r|X_i=x,W_i=w)$ denote the observed distribution of responses $R_i$ given values $x$ of treatments $X_i$ and $w$ of the control variables $W_i$. For brevity, I will often use this type of shorthand in long expressions.
Theorem (ref) shows that the derivative of $P(R_i \le r|X_i=x,W_i=w)$ with respect to changes in $x_j$ provides a positively-weighted linear combination of the causal response in $H$ due to $X_j$: “marginal” causal effects $\partial_{x_j} h(x,U_i)$ due to a small change in $X_j$. The proof of Theorem (ref) relates the derivative of the conditional CDF of $R$ to a mixture of (infeasible) quantile regressions that condition on response type $V_i$ (Lemma (ref)), and then makes use of a connection between quantile regressions and local average structural derivatives hoderleinmammen2007,sasaki_2015. As an intermediate step in establishing Theorem (ref), Lemma (ref) in Appendix (ref) shows establishes the connection between $\partial_{x_j} P(R_i \le r|x,w)$ and the conditional quantiles of $H$ given $X$.\\
Example: Theorem (ref) generalizes the well-known formula for “marginal effects” in the probit model: $\partial_{x_j}P(R_i=1|X_i=x) = \sigma^{-1}\phi(x^T \beta/\sigma)\cdot \beta_j$, where $\phi$ is the standard normal probability density function. In the probit model, $v$ is degenerate and the single threshold $\tau_v(0)=0$, while $h(x,u) = x^T \beta + u$ and $H_i|X_i=x \sim \mathcal{N}(x^T\beta, \sigma^2)$. Thus, $f_H(\tau(0)|x)=f_H(0|x)=1/\sigma\cdot \phi(-x'\beta/ \sigma)=\phi(x'\beta)$.\\
Normalization: It is well-known that $\beta$ in the probit model is only identified up to an overall scale normalization, often achieved by fixing the variance of the error distribution $\sigma^2=1$. Similarly, we lack from Theorem (ref) the ability to pin down the overall scale of derivatives of the structural function $\partial_{x_j}h(x,U_i)$. The inner expectation in Theorem (ref) (indicated by square brackets [ ]) is over heterogeneity in causal effects $U_i$, while the outer expectation (indicated by curly brackets \{ \}) is over heterogeneity $V_i$ in reporting functions. Expanding this second expectation out, we have
The weights $dF_{V|W}(v|w) \cdot f_H(\tau_{v}(r)|x,v,w)$ that multiply the conditional expectation do not necessarily integrate to one---indeed all that we can say about $\mathbbm{E}[f_H(\tau_{V_i}(r)|x,V_i,w)|W_i=w]=\int dF_{V|W}(v|w) \cdot f_H(\tau_{v}(r)|x,v,w)$ is that it is positive. However, considering the ratio of two derivatives cancels out the dependence on this unknown scale:
where $\omega_r(x,v,w):= f_H(\tau_v(r)|x,v,w)/\mathbbm{E}[f_H(\tau_{V_i}(r)|x,V_i,w)|W_i=w]$. The function $\omega_r$ yields weights that are positive and integrate to one, i.e. $\mathbbm{E}[\omega_r(x,V_i)|X_i=x,W_i=w]=1$. To contrast this with the positive but non-normalized integration measure that appears in ((ref)), I refer to weights such as the $\omega_r$ appearing in ((ref)) as “convex”. Note that the convex weight applied to each group characterized by $H_i=\tau_v(r),X_i=x,V_i=v,W_i=w$ is exactly the same in both the numerator and denominator of ((ref)). Eq. (ref) can be seen as a ratio $\tilde{\beta}_2/\tilde{\beta}_1$ of two parameters of the form $\tilde{\beta}_j=\mathbbm{E}[\rho_i \cdot \partial_{x_j} h(X_i,U_i)]$ described in Section (ref), where $\rho_i$ picks out individuals with $H_i$ close to $\tau_{V_i}(r)$ (and $X_i=x,W_i=w$).\footnote{In particular, let $\rho_i:=\rho(U_i,V_i,X_i,V_i)=\frac{\mathbbm{1}\left\{h(x,U_i) \in B_\epsilon^1(\tau_{V_i}(r)), X_i \in B_\epsilon^J(x), W_i \in B_\epsilon^{d_W}(w)\right\}}{\mathbbm{E}\left[\mathbbm{1}\left\{h(x,U_i) \in B_\epsilon^1(\tau_{V_i}(r)), X_i \in B_\epsilon^J(x), W_i \in B_\epsilon^{d_W}(w)\right\}\right]}$, where $B_\epsilon^d(p)$ denotes an open ball of radius $\epsilon$ centered around $p \in \mathbbm{R}^d$, e.g. $B_\epsilon^1(p) = (p-\epsilon,p+\epsilon)$. Then consider the limit of $\tilde{\beta}_j$ as $\epsilon \rightarrow 0$. See the end of the proof of Theorem (ref) for more details about this limit.}
Note that the practice sometimes seen in applied work of reporting standardized or “beta” coefficients (in which each regressor $X_j$ is normalized against it's standard deviation) would break this important property of ((ref)). In that case, the ratio of the total weights appearing in the numerator and denominator would become $sd(X_{2})/sd(X_{1})$ rather than unity. By contrast, rescaling regression coefficients only by the standard deviation of the outcome $R_i$ leaves (ref) unchanged.\\
Intuition for Theorem (ref): By Eq. ((ref)), the “weight” in the observable $\partial_{x_j} P(R_i \le r|x,w)$ placed on an individual with happiness close to $\tau_v(r)$ is positive and proportional to $dF_{V|W}(v|w) \cdot f_H(\tau_{v}(r)|x,v,w)$. Figure (ref) provides intuition for this particular weighting.
Suppose for simplicity there are no controls $w$. By the law of iterated expectations, we can write $\partial_{x_j} P(R_i \le r|X_i=x)$ as a weighted average of $\partial_{x_j} P(R_i \le r|X_i=x,V_i=v)$ across the various reporting functions $v$ in the population. For a given $v$, $\partial_{x_j} P(R_i \le r|X_i=x,V_i=v)$ captures the “flow” of individuals over the threshold $\tau_{v}(r)$ due to a small change in $x_j$, in one direction or the other. Some of these individuals can have negative effects: $\partial_{x_j} h(x,U_i)<0$, denoted by arrows to the left in Figure (ref). Others can have positive effects $\partial_{x_j} h(x,U_i)>0$, indicated by rightward arrows in Figure (ref). The net effect captured by $\partial_{x_j} P(R_i \le r|X_i=x,V_i=v)$ depends on the average derivative $\mathbbm{E}\left[\partial_{x_j} h(x,U_i)|H_i=\tau_{v}(r),x,v\right]$ local to the threshold. Since the derivative $\partial_{x_j}$ considers an infinitesimal change in $X$, any such “flow” over the threshold requires a positive density there: $f_H(\tau_{v}(r)|x,v)>0$.\footnote{The quantity $f_H(h|x,v) \cdot \mathbbm{E}\left[\partial_{x_j} h(x,U_i)|H_i=h,x,v\right]$ at a given $h$ is sometimes referred to as a “flow density”, and appears in kasy2022, goff2022 and in the physics of fluids, where it arises from the conservation of mass.}
While Theorem (ref) is specific to a fixed value of $X_i=x$ (and $W_i=w$), Corollary (ref) to come shows that averaging back over the distribution of $X_i,Z_i$ yields a simpler formula for the average derivative: $\mathbbm{E}[\partial_{x_j} P(R_i \le r|X_i,W_i)] = -f_{H-\tau_{V}(r)}(0) \cdot \mathbbm{E}\left[\partial_{x_j} h(X_i,U_i)|H_i=\tau_{V_i}(r)\right]$. This again captures a an average causal response among respondents who are located at their individual-specific threshold $\tau_{V_i}(r)$, up to a non-identified but positive scale factor. The estimand of Theorem (ref) is a more disaggregated parameter, representing a more fundamental identification result.
I refer to individuals with $H_i=\tau_{V_i}(r)$ for some $r$ as “marginal”, or “indifferent” between response categories. Theorem (ref) shows that local derivatives of the distribution of $R_i$ conditional on $X_i$ and $R_i$ only average causal effects among these marginal respondents. These marginal respondents averaged over in the RHS of Theorem (ref) cannot be individually identified, since neither $H_i$ nor $\tau_{V_i}(r)$ are observed for a given $i$. However, I show in Appendix (ref) that if the sign of causal effects is assumed to be common across individuals, average characteristics of the marginal respondents can be identified (Section (ref) provides an implementation). Appendix (ref) shows that reporting function heterogeneity can have a counter-intuitive benefit: if the heterogeneous thresholds $f_{V_i}(r)$ are so spread out that they are approximately uniform across the support of $H_i$, then $\partial_{x_j}P(R_i \le r|x,w)$ is proportional to $\mathbbm{E}\left[\partial_{x_j} h(x,U_i)|X_i=x,W_i=w\right]$, which averages over both the marginal and infra-marginal respondents having $X_i=x$ and $W_i=w$.
Beyond the case of binary survey questions, researchers do not typically estimate regressions of response the CDF evaluated at a fixed category $r$, as contemplated by Theorem (ref). However, the result allows us to study the more common practice of modeling the conditional mean of $R_i$ given $X_i$. To see this, suppose that $\mathcal{R}$ consists of integers $\{0,1, \dots, \bar{R}\}$ for some $\bar{R}$. Note that the following identity holds for all $i$:
From this it then follows that for any $x$: $\mathbbm{E}[R_i|X_i=x] = \sum_{r =0}^{\bar{R}-1} P(r<R_i|X_i=x)=\bar{R}-\sum_{r =0}^{\bar{R}-1} P(R_i \le r|X_i=x)$. Then, applying Theorem (ref):
For brevity, I use the shorthand $\sum_r$ for the definite sum $\sum_{r =0}^{\bar{R}-1}$. Collecting ((ref)) across all continuous regressors, we can summarize as:
Remark: if instead of the integers, the researcher associates alternative numerical values $r_j$ with the ordered responses $\mathcal{R}$, where $r_0 < r_1 < \dots < r_R$, then instead of ((ref)) we have $R_i = r_0+\sum_{j=0}^{R-1} (r_{j+1}-r_j)\cdot \mathbbm{1}(r_j < R_i)$. The above results thus generalize with $f_H(\tau_{v}(r_j)|x,v,w)$ upweighted by the positive factor $(r_{j+1}-r_{j})$. This implies that different labeling schemes could be used in estimation to achieve different weightings over local causal effects, though the most information one could learn is by simply repeating Theorem (ref), one $r$ at a time. When considering mean regression, using integer category labels is natural in that it weighs each threshold in proportion to its occupancy, as demonstrated in Eq. (ref) below. Corollary (ref) also holds unchanged so long as $\mathcal{R}$ reflects any set of consecutive integers, with $\Sigma_r$ denoting a sum over all but the highest integer in $\mathcal{R}$.\\
Another way to express Corollary (ref) is to let $\tau_v := \{\tau_v(r)\}_{r \in \mathcal{R}}$ denote the set of all thresholds for individuals with reporting function $v$. Then for each $j$ that satisfies $REG_j$:
where $\rho(x,v,w):=\sum_r f_H(\tau_v(r)|x,v,w)$ we assume that $\lim_{h\rightarrow \infty}f_H(h|x,v,w)=0$ and that for each $v \in \mathcal{V}$, the $\tau_v(r)$ are all distinct.\footnote{Let $A$ and $B$ be random variables, where $B$ is absolutely continuous and let $\mathcal{B}$ be a finite set of distinct values. Assume that $\mathbbm{E}[A|B=b]$ is continuous in $b$, so we can then define $\mathbbm{E}[A|B \in \mathcal{B}]$ simply as $\lim_{\epsilon \downarrow 0} \mathbbm{E}[A|\min_{b \in \mathcal{B}} |B-b| < \epsilon]$ which works out to $\sum_{b \in \mathcal{B}} \frac{f_{B}(b)}{\sum_{b' \in \mathcal{B}}f_{B}(b')} \cdot \mathbbm{E}[A|B=b]$.} This expression shows that $ \partial_{x_j} \mathbbm{E}[R_i|x,w]$ averages over all units having $X_i=x$ (and $W_i=w$), located at any of their individual-specific happiness thresholds, with (positive but not convex) weights $\rho(X_i,V_i,X_i)$.
As in ((ref)), if we consider the ratio of such regression derivatives for two continuous treatment variables $X_1$ and $X_2$, the “total” weight cancels out:
where $\tilde{\beta}_j(x,w):=\mathbbm{E}\left[\omega(x,V_i,w)\cdot \partial_{x_j} h(x,U_i)|H_i \in \tau_{V_i},X_i=x,W_i=w\right]$ and $ \omega(x,v,w):=\rho(x,v,w)/\mathbbm{E}\left[\rho(x,V_i,w)|H_i \in \tau_{V_i},X_i=x,w\right]$. The quantity $\tilde{\beta}_j(x,w)$ is thus a convex combination of causal effects with respect to $X_j$ across individuals in the population, in the sense described in Sections (ref) and (ref). Note that the weights $\omega$ appearing in the numerator and denominator are the same for any $(x,w)$.
Theorem (ref) shows how observable derivatives $\partial_{x_j}P(R_i \le r|x,w)$ and $\partial_{x_j}\mathbbm{E}[R_i|x,w]$ can be interpreted in terms of average causal effects among individuals $i$ who are marginal between response categories and for whom $X_i=x$, $W_i=w$. These local derivatives at a given $x,w$ are identified by a non-parametric regression of $\mathbbm{1}(R_i\le r)$ or $R_i$ on $X_i$ and $W_i$, respectively.
Corollary (ref) shows furthermore that if one averages these local regression derivatives across the observable distribution of $X_i,W_i$, one obtains an average causal effect that remains “local” to individuals who are on the margin between response categories, but is no longer specific to individuals having a particular value of $X_i$ and $W_i$:
The density $f_{H-\tau_{V}(r)}(0)$ is not identified by the data, but it does not depend on $j$. Thus we have as in ((ref)) that this unidentified density cancels out in ratios, i.e. $\frac{\mathbbm{E}[\partial_{x_2} P(R_i \le r|X_i,W_i)]}{\mathbbm{E}[\partial_{x_1} P(R_i \le r|X_i,W_i)]} = \frac{\mathbbm{E}\left[\partial_{x_2} h(X_i,U_i)|H_i=\tau_{V_i}(r)\right]}{\mathbbm{E}\left[\partial_{x_1} h(X_i,U_i)|H_i=\tau_{V_i}(r)\right]}$, and if $\mathcal{R} = \{0,1, \dots \bar{R}\}$ we have similarly for the mean that:
where $\tilde{\beta}_j:=\sum_{r =0}^{\bar{R}-1} \omega_r\cdot \mathbbm{E}[\partial_{x_j} h(X_i,U_i)|H_i = \tau_{V_i}(r)]$ and $ \omega_r:=\frac{f_{H-\tau_{V}(r)}(0)}{\sum_{r' =0}^{\bar{R}-1}f_{H-\tau_{V}(r')}(0)}$ are positive weights that sum to one. Response thresholds $r$ that are more “populated” in the sense that $f_{H-\tau_{V}(r)}(0)$ is larger, receive higher weight, in such a way that $\tilde{\beta}_j = \mathbbm{E}[\partial_{x_j} h(X_i,U_i)|H_i \in \tau_{V_i}]$.\footnote{i.e. $\sum_{r' =0}^{\bar{R}-1} f_{H-\tau_{V}(r)}(0) \cdot \mathbbm{E}\left[\partial_{x_j} h(X_i,U_i)|H_i=\tau_{V_i}(r)\right] = \rho \cdot \mathbbm{E}[\partial_{x_j} h(X_i,U_i)|H_i \in \tau_{V_i}]$, with $\rho = \sum_{r' =0}^{\bar{R}-1}f_{H-\tau_{V}(r')}(0)$.} In the case with no control variables $W_i$, we then obtain Eq. (ref) stated in the introduction.
If the conditional mean function $\mathbbm{E}[R_i|X_i=x,W_i=w]$ happens to be linear in $x$ and $w$, then the quantity $\mathbbm{E}[\partial_{x_j}\mathbbm{E}[R_{i}|X_i,W_i]]$ on the LHS of Eq. (ref) and (ref) is simply the coefficient $\gamma_j$ from the OLS regression
where the vector of control variables $W$ includes a constant. While specification (ref) is the standard in empirical practice, Appendix (ref) discusses the implications of this practice when the functional form is misspecified, i.e. when $\mathbbm{E}[R_i|X_i=x,W_i=w]$ is not actually linear but the researcher proceeds in estimating (ref) anyways. Such issues are generally a concern when selection on observables identification arguments are implemented via linear regression, and are not specific to the use of subjective ordinal outcome variables.
Equation (ref) shows that a ratio of regression derivatives at $X=x$ identifies the ratio of a conditional average causal effect of $X_2$ on $H$ to the same conditional average of the effect of $X_1$ on $H$, among individuals for whom $X=x$. Similarly, (ref) shows that a ratio of average regression derivatives (or simply OLS regression coefficients in the case of a linear conditional mean) has a similar interpretation, but averaging over $x$. luttmer2005 and ditellaetal represent two prominent empirical studies in which the relative magnitude of regression coefficients (with subjective well-being as the dependent variable) is interpreted as yielding the implicit trade-off between two goods.
In general, a ratio of averages is not the same as an average of ratios, and thus neither (ref) nor (ref) immediately yields an average marginal rate of substitution parameter $\widetilde{MRS}$ of the form introduced in Section (ref). For example, Equation (ref) does not immediately yield an average of $MRS_i(X_i)$. A sufficient condition however is that $Cov\left(\left.MRS_i(X_i), \partial_{x_1} h(X_i,U_i)\right|H_i \in \tau_{V_i}\right) = 0$. In this case
capturing the average marginal rate of substitution between $X_1$ and $X_2$, among respondents who are marginal at any threshold, i.e. $H_i = \tau_{V_i}(r)$ for some $r$. This covariance condition says that heterogeneity in $MRS_i(X_i)$ across individuals is uncorrelated with heterogeneity in the magnitude of the marginal effect of $X_1$ alone.
Proposition (ref) of Appendix (ref) also shows how a similar result to Eq. (ref) holds using the estimand $\frac{\partial_{x_2}\mathbbm{E}[R_{i}|X_i=x,W_i=w]}{\partial_{x_1}\mathbbm{E}[R_{i}|X_i=x,W_i=w]}$ of (ref), which fixes a value of $x$. In this case note that variation in $MRS_i(x)$ conditional on $H_i$ and $X_i$ comes from $U_i$ alone. Thus if $U_i$ is degenerate conditional on $X_i$ and the value of $H_i$ (e.g. if $h$ is invertible in a scalar $u$), then the needed covariance condition holds automatically. Proposition (ref) generalizes this with a covariance restriction similar to the above, which conditions on $X_i$ and $W_i$. Proposition (ref) also shows how one can obtain a one-sided bound on the RHS of (ref) by relaxing this to assume a known sign of the correlation between $MRS_i(X_i)$ and $\partial_{x_1} h(X_i,U_i)$.
Below I consider two particular cases in which assuming some natural structure for the function $h(x,u)$ is sufficient to interpret $\frac{\partial_{x_2}\mathbbm{E}[R_{i}|X_i=x,W_i=w]}{\partial_{x_1}\mathbbm{E}[R_{i}|X_i=x,W_i=w]}$ as a marginal rate of substitution, and the ratio of averages of such derivatives as an average of such marginal rates of substitution across individuals.
We say that the potential outcomes function $h(x,u)$ is weakly separable between $x$ and $u$ when
i.e. some function $g: \mathcal{X}\rightarrow \mathbbm{R}$ aggregates over the treatments $X$ into a scalar $g(x)$, which is then combined through $\texttt{h}$ with heterogeneity $u$ in a way that may or may not be additively separable. For example, a linear model $h(x,u) = x^T\beta + u$ sets $g(x) = X^T\beta$ and $\texttt{h}(g,u) = g+u$, combining a linear causal response with an additive scalar error term. For any $g(x)$, the additively separable form $\texttt{h}(g,u) = g+u$ is equivalent to imposing that the causal effect of changing between any treatment values $x$ and $x'$ is the same for all individuals. This assumption is implicit in much empirical work employing regressions with subjective outcome data.
When ((ref)) holds, Eq. (ref) yields
where the highlighted factors cancel out in the numerator and denominator, since the derivatives of $g(x)$ do not depend on $v$. In the weakly separable model the marginal rate of substitution between $X_1$ and $X_2$ when $X=x$ is the same for all individuals and equal to $\frac{\partial_{x_2} g(x)}{\partial_{x_1} g(x)}$. Thus the ratio of local regression derivatives at a point $X=x$ identifies the MRS at that point $x$.\footnote{Weakly separable models for ordered response in which $u$ is a scalar have been studied by matzkin1994. Appendix (ref) discusses how Eq. (ref), which does not require $u$ to be a scalar, relates to that body of work.} A testable implication of the weakly separable model is therefore that the LHS of Eq. (ref) does not depend on the value of the controls $w$.\footnote{Another testable implication is that $\partial_{x_2}P(R_{i} \le r|x,w)/\partial_{x_1}P(R_{i} \le r|x,w)$ does not depend on $r$. Appendix (ref) applies this insight to test the assumption that $X$ does not directly affect reporting functions. DHAULTFOEUILLE2024105075 consider testable restrictions of a similar weakly-separable structure in certain IV models, while relaxing exclusion.}
Suppose that $h$ represents preferences and for each individual, these preferences are quasi-linear in $X_1$ such that $h(x,u) = x_1 + h(x_2, \dots x_J,u)$.\footnote{Any preference relation that is quasi-linear in $X_1$, continuous, and strictly “increasing” in $X_1$ admits of a representation $h(x,u) = x_1 + h(x_2, \dots x_J,u)$ rubinstein. In the other direction, we can see that if $h(x,u)$ is a representation of quasi-linear preferences that is strictly increasing in $x_1$ and differentiable in $x_2 \dots x_J$, then it must be the case that $h(x,u) = \phi(x_1 + h(x_2, \dots x_J,u))$ where $\phi_u$ is a strictly increasing function.} Quasi-linear utility is widely used in economics to simplify welfare analysis (see e.g. FENG2025105927).\footnote{Although quasi-linearity is a property of preferences, we can think of $h(x,u) = x_1 + h(x_2, \dots x_J,u)$ as a cardinalization of these ordinal preferences (which will generally differ by individual $i$) in which a unit increase in $x_1$ has equal weight for any individual in population expectations involving $h$. Under this normalization, $\mathbbm{E}[h(x,U_i)]$ for example represents a utilitarian social welfare function whose value is unaffected by transfers of $x_1$ between individuals.} When the elements of $X$ are priced, quasilinearity in $X_1$ can also deliver demand functions for the remaining goods that do not depend on income (see e.g. qljet).
In this case the condition $Cov\left(\left.MRS_i(X_i), \partial_{x_1} h(X_i,U_i)\right|H_i \in \tau_{V_i}\right) = 0$ is satisfied trivially, because $\partial_1 h(x,U_i) = 1$ with probability one. Thus we have that $\frac{\mathbbm{E}[\partial_{x_2}\mathbbm{E}[R_i|X_i,W_i]]}{\mathbbm{E}[\partial_{x_1}\mathbbm{E}[R_i|X_i,W_i]]}= \mathbbm{E}\left[\left.MRS_i(X_i)\right|H_i \in \tau_{V_i}\right]$. Further $\frac{\partial_{x_2}\mathbbm{E}[R_i|X_i=x,W_i=w]}{\partial_{x_1}\mathbbm{E}[R_i|X_i=x,W_i=w]} = \mathbbm{E}\left[\left.MRS_i(x)\right|H_i \in \tau_{V_i}, X_i=x,W_i=w\right]$, which can be derived as a special case of Proposition (ref) given in Appendix (ref).
In a prominent paper, luttmer2005 studies the effects of absolute and relative income on life satisfaction, investigating whether individuals draw on social comparisons in assessing their personal well-being. To do so, luttmer2005 merges data from the 1987 and 1992 waves of the U.S. National Survey of Families and Households (NSFH)---which contains a question on self-reported satisfaction with life along with self-reported socioeconomic data---to information on the local average earnings for a given household constructed from the Current Population Study and the 1990 Census.
In the notation of the present paper, let $i$ denote the primary respondent of an individual household in the NSFH. We consider two treatment variables $X_i=(X_{1i},X_{2i})$, where $X_{1i}$ denotes the log of household income for $i$'s household (self-reported in the NSFH) and $X_{2i}$ denotes average predicted log earnings in the Public Use Microdata Area (PUMA) in which $i$ lives. The construction of this variable is described in detail in luttmer2005. $R_i$ denotes $i$'s response to the question “taking things all together, how would you say things are these days?”, reported on a one to seven Likert-type scale in which a response of one indicates “very unhappy” and seven “very happy”.\footnote{The intermediate values 2-6 do not have associated descriptions in the survey (e.g. “somewhat happy”), and are labeled by integers only.} Finally, $W_i$ represents a vector of control variables that includes home size/type/value, employment, education, gender, marriage, race religion, state fixed effects and PUMA characteristics.
I follow luttmer2005 and focus on households in which the main respondent was married in both waves of the NSFH. Details on the sample construction are provided in Appendix (ref). While I let $i$ denote the main respondent for a household, the primary specification of luttmer2005 averages values of $R_i$ and $W_i$ between the main respondent and their spouse, finding very similar results. I focus on the individual-level specification for two reasons: i) it affords a more straightforward interpretation through the lens of the results of this paper, given that the main respondent and their spouse may have different reporting functions; and ii) I explore departures from linear models, where averaging across observations within a household does not affect the functional form of the regression. Nevertheless, results with this averaging are provided in Appendix (ref).
The main results of luttmer2005 exploit a selection-on-observables strategy, estimating an OLS regression of $R_i$ on $X_i$ and $W_i$, i.e. Eq (ref):
and ascribing a causal interpretation to the coefficients $\gamma_1$ and $\gamma_2$. Luttmer uses fixed effects regressions as well as data on movers between PUMAs to argue that selection due to neighborhood choice is not a major concern in this context. Luttmer further argues that individuals' definitions of “very happy” or “very unhappy” are not affected by $X$, by replicating the qualitative results with other outcome variables that are expected to be less prone to this threat. I refer the reader to sections IV.B and IV.C of luttmer2005 for details. These arguments motivate making Assumption EXOG in this context.
Luttmer finds that an increase in household earnings increases subjective well-being $\gamma_1 > 0$, while an increase in the earnings of one's neighbors decreases subjective well-being $\gamma_2 < 0$. This provides evidence that well-being is influenced not only by one's absolute income, but also one's relative income compared with the reference group of one's neighbors.\footnote{This finding has since been replicated using experimental variation in beliefs about relative income jansens.} While luttmer2005 also reports estimates that instrument for own-income to overcome potential measurement error, I focus on magnitudes from the benchmark OLS regression (ref). In this specification, the positive coefficient on own income has about half the magnitude as the negative coefficient on PUMA (neighbors') income. That is, if one's PUMA were to go up by 1%, one's own income would need to go up by about 2% to leave the respondents' well-being unaffected.
I confirm this finding qualitatively in Column (1) of Table (ref). Column (2) reports the numerical results from Table 1 of luttmer2005 (main respondent column), in which $\hat{\gamma}_2/\hat{\gamma}_1 = -2.23$. In column (1) I implement regression (ref) the publicly available NSFH data merged with the PUMA income variable constructed by Luttmer (the replication data construction is described in Appendix (ref)). I obtain similar results in both sign and magnitude, with $\hat{\gamma}_1 = 0.0877$ and $\hat{\gamma}_2 = -0.229$ for a ratio of $\hat{\gamma}_2/\hat{\gamma}_1 = -2.614$.\footnote{That I am not able to match the numerical results exactly is likely explained by the many choices involved in how exactly to define some of the control variables, or possible updates to the underlying NSFH data over the last two decades.}
If regression (ref) is correctly specified---that is $\mathbbm{E}[R_i|X_i,W_i]$ is indeed a linear function of $X_i$ and $W_i$---then Theorem (ref) implies via (ref) that the quantity $\tilde{\beta}_2(x,w)/\tilde{\beta}_1(x,w)$ is approximately constant over values $x$ of the treatment variables and $w$ of the control variables (and equal to $-2.614$), where recall that $\tilde{\beta}_j(x,w)$ is a weighted average of $\partial_{x_j} h(x,U_i)$ over individuals with $U_i$ such that their happiness is exactly at the threshold between two response categories when $X_i=x,W_i=w$. This is consistent for example with a structural function that takes the linear form $h(x,u) = x^T \beta+u$, in which case $\tilde{\beta}_2(x,w)/\tilde{\beta}_1(x,w) = \beta_2/\beta_1 = -2.614$. However, the magnitudes of $\beta_1$ and $\beta_2$ would not be identified separately, even with this strong functional form assumption about $h$. Appendix (ref) reports results in which $R_i$ and $W_i$ are constructed by averaging responses of the main respondent and those of their spouse, which are similar.
If $\mathbbm{E}[R_i|X_i,W_i]$ is not linear in fact in $X_i$ and $W_i$, then OLS estimates of Eq. (ref) are not guaranteed to be interpretable in terms of causal effects, even if the assumptions of Theorem (ref) do hold. Appendix (ref) discusses this issue generally, and in this section I discuss the robustness of the ratio $\gamma_2/\gamma_1 = -2.614$ to relaxing this functional form assumption.
Column (3) of Table (ref) employs a semi-parametric estimator following robinsonestimator that assumes the partially linear form $\mathbbm{E}[R_i|X_i,W_i] = f(X_{1i},X_{2i})+\lambda^T W_i$, in which $\lambda$ is estimated by residualizing $R_i$ and each component of $W_i$ with respect to $X_i$, before performing a bivariate kernel regression of $R_i - \hat{\lambda}^T W_i$ on $X_i$ to estimate $f(\cdot, \cdot)$. In this specification the coefficients $\gamma_j$ from Eq. (ref) are replaced with $$\gamma_j(x) := \partial_{x_j}\mathbbm{E}[R_{i}|X_i=x,W_i=w]$$ which does not depend on $w$ owing to the additively separable structure of $\mathbbm{E}[R_i|X_i,W_i]$. This implies that the estimand $\gamma_j(x)$ can be interpreted as proportional to an average causal effect $\tilde{\beta}_j(x)$ that also does not depend on $w$.
The first two rows of Column (3) report $\gamma_j(X_i)$ averaged across the empirical distribution of $X_i$. These estimates of $\mathbbm{E}[\gamma_j(X_i)]$ are numerically fairly similar to the $\hat{\gamma}_j$ reported by luttmer2005 from the OLS specification (ref). These average derivatives appear to mask only minor non-linearity in $\mathbbm{E}[R_i|X_i,W_i]$ with respect to $X_i$. Dividing the first two rows yields an estimate of $-0.202/0.122 \approx -1.66$ for $\frac{\mathbbm{E}[\gamma_2(X_i)]}{\mathbbm{E}[\gamma_1(X_i)]} = \frac{\mathbbm{E}[\tilde{\beta}_2(X_i)]}{\mathbbm{E}[\tilde{\beta}_1(X_i)]}$, but when computing the ratio of regression derivatives evaluated at $X_i$ for each observation, and then averaging across the sample, one instead obtains a value of $\hat{\mathbbm{E}}[\hat{\gamma}_2(X_i)/\hat{\gamma}_1(X_i)] = -1.581$. This average ratio is reported in the row labeled “Ratio” of Column (3) and estimates
If we assume a weakly separable causal model $h(x,u) = \mathtt{h}(g(x),u)$, then this value of $-1.581$ in turn represents an estimate of $\mathbbm{E}\left[\frac{\partial_{x_1}g(X_i)}{\partial_{x_2}g(X_i)}\right]$, the overall population mean of the marginal rate of substitution between own income and neighbors' income, which is in this model common among all individuals sharing a value of $X_i$.\footnote{If we instead make the assumptions of Appendix Proposition (ref), then $\mathbbm{E}\left[\frac{\gamma_2(X_i)}{\gamma_1(X_i)}\right]$ is equal to $\int dF_{X}(x) \cdot \mathbbm{E}\left[\left.\frac{\partial_{x_1}h(x,U_i)}{\partial_{x_2}h(x,U_i)}\right|H_i \in \tau_{V_i}, X_i=x\right]$, using as well that $\mathbbm{E}\left[\left.\frac{\partial_{x_1}h(x,U_i)}{\partial_{x_2}h(x,U_i)}\right|H_i \in \tau_{V_i}, X_i=x,W_i\right]$ does not depend on $W_i$ if $\mathbbm{E}[R_{i}|X_i=x,W_i=w]$ is separable between $x$ and $w$.} This value is substantially smaller than the $-2.614$ reported in column (1), which assumes linear conditional means.
Column (5) of Table (ref) shows the PUMA/own ratio to be similar when the controls $W_i$ are omitted, and fully non-parametric regression of $R_i$ on $X_{1i}$ and $X_{2i}$ becomes feasible. For comparison, column (4) implements OLS with no controls. Taking column (5) as our estimate of $\mathbbm{E}[\tilde{\beta}_2(X_i)/\tilde{\beta}_1(X_i)]$ would avoid the functional form restriction that $\mathbbm{E}[R_i|X_i, W_i]$ be linear in $W_i$, but at the expense of requiring Assumption EXOG to hold without the control variables $W_i$.\footnote{I report only the point estimate for the sample mean of $\hat{\gamma}_2(X_i)/\hat{\gamma}_1(X_i)$ in columns (3) and (5) of Table (ref), as a bootstrap computation of standard errors would be computationally intensive in the case of (3) given the number of control variables $W_i$. Standard errors are computed for the sample means of $\hat{\gamma}_2(X_i)$ and $\hat{\gamma}_1(X_i)$ separately in (3) and (5), but neither are clustered at the PUMA level as the npregress command in Stata does not accommodate cluster robust inference.} The gap in the PUMA/own ratio between linear and nonlinear models is much greater without controls, suggesting that the control variables eliminate much of the non-linearity with respect to $X_i$ in the conditional mean of $R_i$.
While we know by Corollary (ref) that the regression derivatives reported in Table (ref) average over respondents who are on the margin between two adjacent response categories, we also know from Theorem (ref) that we can isolate causal effects for respondents that are on a single such margin $r$ and $r+1$, for some $r \in \{1, 2, \dots 6\}$.
Table (ref) reports coefficients from a linear probability model that takes, for a given $r$, the conditional expectation function $\mathbbm{E}[\mathbbm{1}(R_i \le r)|X_i=x,W_i=w] = P(R_i \le r|X_i=x,W_i=w)$ to be linear in $X_i$ and $W_i$, with coefficients $(\gamma_{1r},\gamma_{2r}, \lambda_r)$ specific to that response category $r$, i.e. $\mathbbm{1}(R_i \le r) = \gamma_{1r} X_{1i}+\gamma_{2r} X_{2i} + \lambda_r^TW_i+\epsilon_{ri}$ with $\mathbbm{E}[\epsilon_{ri}|X_i,W_i]=0$.
Table (ref) reveals that the sign of $\gamma_{1r}$ is positive for all $r$ when it is statistically significant, the sign of $\gamma_{2r}$ is consistently negative when it is statistically significant, and the ratio $\gamma_{2r}/\gamma_{1r}$ is never differs from the “aggregate” value of -2.614 recovered by mean regression in a statistically significant way.
The information in Table (ref) is further visualized in Figure (ref). The top-left panel depicts the coefficients $\gamma_{1r}$ versus $r$. By Theorem (ref), we know that if the linear model correctly captures the conditional mean function of $R_i$ given $X_i$ and $W_i$, then each $\gamma_{1r}$ captures a positively weighted aggregation of the marginal effect of own income $\partial_{x_1} h(x,U_i)$ across individuals whose $U_i$ put them on their individual-specific threshold between response categories $r$ and $r+1$.
The clear hump-shaped pattern across values of $r$ could be explained by heterogeneity in the mean causal effect among the individuals at each of the thresholds, or by differences in the density of individuals at that threshold. Suppose that this effect were a constant $\partial_{x_1} h(x,U_i) = \beta_1$ for all $i$. Then, Theorem (ref) shows that $\gamma_{1r}$ would be equal to $\beta_1 \cdot \mathbbm{E}[f_H(\tau_{V_i}(r)|X_i,V_i,W_i)]$ for each $r$.\footnote{This uses that since $\gamma_{1r}$ does not depend on $x$ or $w$, we must have that $\mathbbm{E}[f_H(\tau_{V_i}(r)|x,V_i,w)|W_i=w]=\mathbbm{E}[f_H(\tau_{V_i}(r)|x_i,V_i,W_i)]$ for all $x$ and $w$.} The quantity $\mathbbm{E}[f_H(\tau_{V_i}(r)|X_i,V_i,W_i)]$ is unobservable, but note that for any $r \in \{1, 2, \dots 6\}$ the observable probability $P(R_i=r) = P(R_i \le r)-P(R_i \le r-1)$ identifies the quantity
where the approximation takes the density $f_H(\tau_{V_i}(r)|X_i,V_i,W_i)$ to be roughly constant on the interval $[\tau_{V_i}(r),\tau_{V_i}(r-1)]$. This will be a good approximation if that interval is small with high probability (i.e. in the limit of many categories), in which case $P(R_i=r)$ is roughly proportional to $\mathbbm{E}[f_H(\tau_{V_i}(r)|X_i,V_i,W_i)]$, if $\mathbbm{E}[\tau_{V_i}(r)-\tau_{V_i}(r-1)]$ does not vary much with $r$. The bottom-left panel of Figure (ref) depicts $P(R_i=r)$ and reveals that it does indeed capture the same basic pattern as $\gamma_{1r}$.
Similarly, the negative values $\gamma_{2r}$ depicted in the top-right panel of Figure (ref) capture a positive aggregation of the marginal effect of PUMA income on $H$ with the weights $\mathbbm{E}[f_H(\tau_{V_i}(r)|X_i,V_i,W_i)]$. Again, the pattern of $\gamma_{2r}$ mirrors that of $P(R_i=r)$, which is consistent with a model in which this effect $\partial_{x_2} h(x,U_i)$ is captured by a single number $\beta_2$ for all $i$. Finally, we see in the bottom-right panel of Figure (ref) the observation made after Table (ref), that the pattern cancels out and $\gamma_{2r}/\gamma_{1r}$ is roughly constant across $r$ (categories 1 and 2 are omitted due to being very imprecisely estimated). An F-test of equality of $\gamma_{2r}/\gamma_{1r}$ across all $r$ fails to reject (p-value: $0.98$).
Overall, the strong similarity in the shapes of the first three panels of Figure (ref) are suggestive that the differences in $\gamma_{1r}$ and $\gamma_{2r}$ are driven by the underlying latent density of happiness, than by heterogeneity in causal effects across the happiness distribution. This is consistent with a simple constant effects model in which $\gamma_{2r}/\gamma_{1r} = \beta_2/\beta_1$, or more generally by a weakly-separable model of the form $h(x,u) = \mathtt{h}(g(x),u)$.
Figure (ref) compares the gender balance and education of respondents that on the margin between categories $r$ and $r+1$, for each $r$, with that of the population as a whole. These comparisons are based on Proposition (ref) in Appendix (ref), which leverages additional assumptions to identify averages of an attribute $A_i$ among marginal respondents.
The upper panels of Figure (ref) report estimates of $\mathbbm{E}[A_i|H_i = \tau_{V_i}(r)]$, under an assumption that $\{X_{i} \perp\!\!\!\!\perp (A_i,U_i,V_i)\}|W_i$ and imposing the additional restriction that the sign of the effect of household income on happiness is the same for all units (not that this assumption is not imposed for the main results). The implementation further takes the conditional expectation of $A_i \cdot \mathbbm{1}(R_i\le r)$ to be linear in $x$ and $w$ and assumes a linear probability model for $\mathbbm{1}(R_i\le r)$ (see Appendix (ref) for details).
In particular, the top left panel displays 95% confidence intervals for $\mathbbm{E}[A_i|H_i = \tau_{V_i}(r)]$ versus $r \in \{2,3 \dots 6\}$ when $A_i$ is taken to be an indicator for the main respondent $i$ attending college.\footnote{Confidence intervals for $r=1$ are dropped in all panels of Figure (ref) for visibility, as the standard error is much larger than for other $r$.} The horizontal line (orange) depicts the overall sample mean (an estimate of $\mathbbm{E}[A_i]$). For none of the margins $r$ can we reject the null hypothesis that the average rate of college among marginal respondents for that category is the same as the overall population mean. A similar result appears in the top-right panel, in which this calculation is repeated with $A_i$ equal to $i$'s years of education. There is some weak evidence that individuals on the margin on categories four and five out of seven have fewer years of education than the average. This is consistent with the finding of CPBL that lower-education individuals are more likely to “bunch” at focal points in the response space $\mathcal{R}$, for example the midpoint (which is indeed 4 on the 1 to 7 scale).
The bottom panels of Figure (ref) exploit the identification of the relative odds for a binary $A_i$, comparing marginal respondents to the population as a whole: $$\frac{P(A_i=1|H_i = \tau_{V_i}(r),x,w)/P(A_i=0|H_i = \tau_{V_i}(r),x,w)}{P(A_i=1|x,w)/P(A_i=0|x,w)},$$ See Eq. (ref) in Appendix (ref). This result makes use of the weaker assumption in Proposition (ref) that only assumes that a binary $A_i$ would represent a valid control variable to add to $W_i$. In this case all that is required is to implement regressions of $\mathbbm{1}(R_i \le r)$ on $x$ and $w$ separately by subsample defined by $A_i$.
The bottom panels report 95% confidence intervals for this ratio of odds, with the horizontal line (orange) depicting unity (equal odds in both populations). The bottom-right panel sets $A_i$ to be an indicator for the main respondent being female, and compares the relative odds of being female among marginal respondents to the population overall. None are statistically different from unity. For clarity, the confidence interval for $r=3$, which is very large, is not shown.
Overall, the results of this section indicate there is some weak evidence that marginal respondents have somewhat less education than the overall population, for the central category in the response space. No differences are detected across gender. The rightmost confidence interval in each panel of Figure (ref), labeled “Avg”, replaces indicators for $R_i \le r$ with $R_i$ to approximate an “average” comparison considering all of the response categories at once. In all cases, we do not find any evidence that the marginal respondents overall differ from the infra-marginal respondents in education or gender.
The results of the proceeding sections suggest that the results of the OLS regression implemented by luttmer2005 can indeed be interpreted as being informative about the causal effects of own and PUMA income on subjective well-being. If one is willing to assume that these two treatments are as-good-as-randomly assigned the sense that $\{(X_{1j},X_{2j}) \perp\!\!\!\!\perp (U_i,V_i)\}|W_i$ (implying EXOG), then the coefficients $\gamma_1$ and $\gamma_2$ from Eq. (ref) have causal interpretations under weak and fully non-parametric assumptions about the latent heterogeneity underlying causal effects and response functions. Using OLS does require that $\mathbbm{E}[R_i|X_i,W_i]$ is indeed linear in $X_i$ and $W_i$, but this caveat is not a product of the ultimate outcome of interest $H_i$ being unobserved: rather, a correct specification of conditional mean functions is important for any selection-on-observables research design with control variables or setting in which multiple treatment variables are considered (see Appendix (ref) for details). Nonetheless, the quantitative estimates are similar if assuming a semi-parametric regression function that is partially linear in the controls, or a fully non-parametric regression that drops the controls altogether.
Although the magnitudes of $\gamma_1$ and $\gamma_2$ are not directly interpretable in terms of causal effects, the coefficient of $\gamma_2/\gamma_1$ is. This supports the main conclusion of luttmer2005 that relative-earnings considerations are indeed important to subjective well-being. A ratio of roughly -2 suggests that a 1% income increase to one's neighbors' average income would require a roughly 2% increase to one's own household income, in order to leave individuals equally happy overall. This interpretation in terms of a marginal rate of substitution is justified if one assumes a potential outcomes model that is weakly-separable in the treatments. The bottom-right panel of Figure (ref) finds that the ratio $\gamma_{2r}/\gamma_{1r}$ of category-specific regression coefficients is relatively constant over $r$, a key implication of a weakly-separable model. This supports the interpretation of $\gamma_2/\gamma_1$ as a marginal rate of substitution; another sufficient condition for this interpretation would be to assume utility to be quasi-linear in the log of own-income, as we saw in Section (ref).
Finally, I find that although the local regression derivatives of $R_i$ at a specific value of $X_i=x$ only capture causal effects among individuals that are indifferent between two response categories when $X_i=x$, these marginal respondents do not appear to be substantially different than infra-marginal respondents in terms of education or gender.
The analysis thus far has considered what is identified by examining how the conditional distribution of $R$ changes over infinitesimal differences in $X$. This section now considers taking discrete differences in treatment values (nesting the results thus far in the limit of small changes). I find that differences in the distribution of $R$ over discrete changes in $X$ can again be interpreted causally, and identify the sign of causal effects if those effects have the same sign across units. However, unlike the case with continuous treatments, magnitudes cannot be quantitatively compared between regressors absent further assumptions. Discrete treatment variables are prevalent in practice, so this highlights a limitation of, e.g. experiments with two treatment arms and subjective outcomes.
Consider any two fixed values $x$ and $x'$, and define $\Delta_i := h(x',U_i)-h(x,U_i)$ to be the “treatment effect” of moving from $X_i=x$ to $X_i=x'$ for unit $i$. Further, let $f_H(y|\Delta,x,v,w)$ denote the density of $H_i$ conditional on $\Delta_i=\Delta$, $X_i=x$,$V_i=v$ and $W_i=w$. As before, let $P(R_{i} \le r|x,w)$ denote a shorthand for $P(R_{i} \le r|X_i=x,W_i=w)$. The following expression shows what is identified from the conditional distribution of $R_i$ across this discrete change between values $x$ and $x'$:
Similar to Theorem (ref), Theorem (ref) shows that the change in $P(R_i \le r|W_i=w,X_i=x)$ over discrete changes in $x$ can be written as a positive linear combination of the causal effect of that variation in $X$ on $H$---a quantity proportional to a parameter of the form $\tilde{\Delta}$ introduced in Section (ref). The quantity $\bar{f}_H(\tau_{V_i}(r)|\Delta_i,x,V_i,w)$ is positive for each $i$ but unknown to the researcher, determined in part by individuals' reporting functions and the underlying distribution of $H_i$. Intuitively, respondents with treatment effect value $\Delta$ are “counted” in the above average if there exists a positive mass of such individuals with $(X_i,W_i)=(x,w)$ and happiness $H_i$ in the range $\tau_{V_i}(r)-\Delta$ to $\tau_{V_i}(r)$. Note that Theorem (ref) exhausts all implications of the observable data $(R_i,X_i)$ regarding variation in the potential outcome functions $h(x,u)$ with respect to $x$ (for a fixed value of the controls $W_i$).\footnote{Given any such fixed $w$, once $P(R_{i} \le r|X_i=x,W_i=w)$ is known for all $r$ for some fixed reference value $x$ of the explanatory variables, along with the distribution of $X_i|W_i=w$, the only remaining information available from the data takes the form of differences $P(R_{i} \le r|x',w)-P(R_{i} \le r|x,w)$ for various values of $x'$ and $r$.}\\
Theorem (ref) as a limiting case of Theorem (ref): A similar expression to that of Theorem (ref) shows up in the “bunching design", which leverages bunching at kinks in decision-makers' choice sets for identification of behavioral elasticities. Since the kink compares just two distinct slopes, an identification problem emerges for elasticity parameters blomquist_bunching_2019. An assumption sometimes used sidestep this issue is that the kink is “small” (e.g. saez2011,kleven2016, see goff2022 for a discussion). An analogous assumption in the context of Theorem (ref) would be that $\Delta_i$ is small with probably one so that for each $\Delta \in supp\{\Delta_i\}$, the density $f_H(h|\Delta,x,v,w)$ is approximately constant for all $h$ between $\tau_v(r)-\Delta$ and $\tau_v(r)$. Under this assumption, Theorem (ref) would simplify to:
Eq. ((ref)) exactly recovers the weighting over individuals achieved by Theorem (ref) using continuous variation in $x$. In particular, the quantity $\mathbbm{E}[\Delta_i|\tau_v(r), x, v,w]$ appears above with the same weight $-dF_{V|W}(v|w)\cdot f_H(\tau_v(r)|x,v,w)$ as $\mathbbm{E}[\partial_{x_j} h(x,U_i)|\tau_v(r),x,v,w]$ does in Eq. ((ref)). Unfortunately, the constant density assumption used to obtain (ref) is quite hard to justify except in the limit that $\Delta_i$ is very small with probability one.\footnote{If we consider the limit $x' \rightarrow x$ with the two differing only in component $j$, this approximation becomes exact and Eq. ((ref)) applied to $(P(R_{i} \le r|x',w)-P(R_{i} \le r|x,w))/(x_j'-x_j)$ reduces to Theorem (ref). See Lemma SMALL in goff2022.} Section (ref) thus explores this issue further when $\Delta_i$ is not small, in the context of mean regression.\\
Intuition for Theorem (ref): We can obtain some intuition for Theorem (ref) as depicted in Figure (ref). Suppose there are two response categories $\mathcal{R} = \{0,1\}$ with a common reporting function $r(h) = \mathbbm{1}(h \ge \tau)$. By iterating expectations over $\Delta_i$, we can consider a single value $\Delta$ of $\Delta_i$ at a time. Thus we aim to show that $\mathbbm{E}[R_{i}|x',\Delta]-\mathbbm{E}[R_{i}|x,\Delta] = \bar{f}_H(\tau|x,\Delta)\cdot \Delta$, using that $P(R_{i} \le 0|x',\Delta)=1-\mathbbm{E}[R_{i}|x',\Delta]$. In Figure (ref), I make the conditioning on $\Delta_i=\Delta$ implicit to simplify notation, taking an example in which $X_i$ is an indicator for marriage with $x'=1$, $x=0$.\\
Mean regression: As our main focus is regressions capturing the conditional mean of $R_i$ with $\mathcal{R}$ an integer response scale, let us as in Eq. ((ref)) aggregate Theorem (ref) across the response categories $r$ to obtain:\footnote{To obtain the notation of Eq. ((ref)) in the introduction from (ref), define $\bar{f}_H(\Delta,x,v,w) := \sum_{r=0}^{\bar{R}-1} \bar{f}_H(\tau_v(r)|\Delta,x,v,w)$.}
Recall from Theorem (ref) that derivatives of the conditional distribution of $R$ yield causal effects $\nabla_x h(x,U_i)$ with weights proportional to $\sum_r f_H(\tau_v(r)|x,v)$. By contrast, (ref) shows that discrete differences in $X$ recover treatment effects $\Delta_i = h(x', U_i)-h(x,U_i)$ with “weights” that themselves depend upon $\Delta_i$ through $\sum_r \bar{f}_H(\Delta_i,x,v,w)$. Since this quantity depends not only on the density of $H$ at response thresholds $\tau_v(r)$ but also the density at points within $\Delta$ of such thresholds through $\bar{f}$, the two weighting schemes do not lead to estimands that can obviously be directly compared.\\
Note: whether or not $\mathbbm{E}[R_i|x',w]-\mathbbm{E}[R_i|x,w]$ is positive or negative does not reflect the sign of the average treatment effect: $\mathbbm{E}[\Delta_i]$. Rather, it depends on how positive and negative treatment effects are aggregated over by the weights $\sum_r \bar{f}_H(\tau_v(r)|\Delta,x,v,w)$. If the CDF functions (or equivalently, quantile functions) of $h(x,U_i)$ and $h(x',U_i)$ cross, then there must be some individuals with $\Delta_i<0$ while others with $\Delta_i>0$.\footnote{Specifically, then $P(\Delta_i < 0) \ge \sup_t \left\{F_{h(x',U_i)}(t)-F_{h(x,U_i)}(t)\right\}$ and $P(\Delta_i > 0) \ge \sup_t \left\{F_{h(x,U_i)}(t)-F_{h(x',U_i)}(t)\right\}$; see e.g. fanpark}. This connects Theorem (ref) to the result of bondandlang, discussed further in Appendix (ref).
Given the foregoing analysis, Theorems (ref) and (ref) together imply that regression coefficients between discrete and continuous treatment variables can be meaningfully compared quantitatively in terms of causal effects in the limit that effects $\Delta_i$ for the discrete treatment are very small, if the conditional mean function is indeed linear.
More generally, a researcher who is interested in comparing a local regression derivative to the mean difference across two discrete groups can construct ratios like:
for some $x$,$x'$, and $x''$. For example, if $X = (income, marriage)$ with $x'=(y,married)$ and $x=(y,unmarried)$ for any income $y$ and $x''=(y,m)$ for $m \in \{married, unmarried\}$, then Eq. (ref) would yield a comparison of regression contrasts involving income to those involving marriage. If $\mathbbm{E}[R_i|X_i,W_i]$ were fully linear, then the numerator of (ref) would not depend on $y$ or $w$ and the denominator would not depend on $y$, $m$ or $w$, yielding a ratio of two linear regression coefficients.
Out goal now is to examine the causal interpretation of Eq. (ref) outside of the limit that $\Delta_i$ is very small. Combining Eq. ((ref)) with Corollary (ref), we know that the ratio in Eq. (ref) is equal to
To interpret this as informative about the relative magnitudes of $\Delta_i$ and $\partial_{x_1}h(x'',U_i)$, the relevant question is how similar the sum $\sum_{r} f_H(\tau_{V_i}(r)|x'',V_i,w)$ over densities at the thresholds is to the corresponding sum over mean densities: $\sum_r \bar{f}_H(\tau_{V_i}(r)|\Delta_i,x,V_i,w)$, at least on average. If these quantities tend to be close to one another in magnitude, then Eq. ((ref)) uncovers something close to the ratio of two convex averages of causal effects. If they differ by an unknown amount, then interpreting ((ref)) in terms of the relative magnitudes of causal effects is not possible.
Reasoning about the magnitudes involved in (ref) is challenging in full generality, but it is possible to derive analytical results to guide our intuition by assuming that there are “many” response categories in $\mathcal{R}$. Given the definition of $\bar{f}$, notice that $\sum_{r} f_H(\tau_v(r)|x'',v,w)$ and $\sum_r \bar{f}_H(\tau_v(r)|\Delta,x,v,w)$ are similar for a given $(\Delta,v,w)$ if
Observe that the two sides of ((ref)) can only differ because the summation occurs over $H_i$ evaluated at the discrete thresholds $\tau_v(r)$. If instead the sums over $r$ were replaced by integrals over all possible values of $H_i$, we would have $\int \left\{\frac{1}{\Delta} \int_{h-\Delta}^{h} f_H(y|\Delta,x,v,w)dy\right\}dh = \int f_H(h|x'',v,w)\cdot dh$, which holds trivially because both sides evaluate to unity for any $\Delta,v,x,w$ and $x''$.\footnote{This is immediate for the RHS, which integrates a density. To see it for the LHS, reverse the order of integrals to obtain $\int dy \cdot f_H(y|\Delta, x,v,w) \left\{\frac{1}{\Delta}\int_{y}^{y+\Delta} dh\right\} = 1$.} Thus it would seem that we have a second “limit” in which discrete and continuous regression differences can be compared: when there are many response categories. However, I show in Appendix (ref) that discrete sums over the thresholds do not exactly correspond to equal-weighted integrals over $h$ in the limit of a continuum of response categories. Rather, in this limit the integrals also involve the quantity $r'(h,v)$, which measures how responsive response function $v$ is at $h$. Nevertheless, the intuition provided by the above logic suggests that looking at the limit of many categories may provide a tractable means of evaluating the quality of Eq. ((ref)) as an approximation.
In Appendix (ref) I define a formal notion of the response categories being “dense” in the space of latent $H_i$, for each reporting function type $v$. This dense response limit allows us to conceptualize there as being an infinite number of response categories, while remaining contained between $0$ and a fixed $\bar{R}$. The dense response limit delivers a tractable approximation which may be reasonable to apply in instances in which the survey question offers many response categories between a lower and upper limit (e.g. integers from 0 to 100).
Proposition (ref) in Appendix (ref) shows how bounds on the ratio of total weights in Equation (ref) can be obtained in the dense response limit when each individual spaces out the thresholds $\tau_v(r)$ at roughly equal intervals---yielding reporting functions that are individually piecewise-linear. Intuitively, the assumption of linear reporting eliminates the effect of $r'(h,v)$, but only within the range of $h$ upon which each individuals' reporting function is increasing. Proposition (ref) gives two sets of bounds. First, a more general bound suggests that discrete contrasts will tend to overstate causal effects relative to regression derivatives, by a factor that is upper bounded by two. A second bound further assumes that the “sensitivity” of individual reporting functions is not too heterogeneous, and suggests that the inflation factor can also be bounded by the reciprocal of the fraction of the population that do not bunch at the endpoints of the response scale. This bound is close to unity when there are few such bunchers, which can be verified empirically.
To assess the performance of the theoretical bounds described above, Appendix (ref) simulates several data-generating-processes (DGPs) for $H_i$ and for the response functions $r(\cdot, V_i)$. The simulations generally provide an optimistic picture that the weights have similar overall magnitude in the numerator and denominator of Equation (ref), across a wide variety of DGPs. Thus $\left\{\mathbbm{E}[R_i|x',w]-\mathbbm{E}[R_i|x,w]\right\}/\partial_{x_j}\mathbbm{E}[R_i|x,w]$ can be interpreted as close to a ratio of weighted averages of causal effects in those DGPs considered. In general, results do not seem to differ substantially whether the number of response categories is small, or whether there are few or many different reporting functions present in the population. When treatment effects become very large relative to the dispersion of happiness in the population, non-linearity in the density of the conditional distribution of happiness becomes important and the sense in which comparisons of magnitude can become misleading is apparent in the simulations.
Appendix (ref) investigates the implications of Proposition (ref) for practical regression analysis with mixed discrete and continuous regressors, focusing both on the common approach of linear regression and non-parametric alternatives.
This paper has investigated the identification of causal effects when using subjective responses as an outcome variable. Such reports typically ask individuals to choose a response from an ordered set of categories, and how individuals use those categories can be expected to differ by individual $i$. Nevertheless, researchers may be willing to suppose that individual responses reflect the value of a well-defined latent variable $H_i$.
Without observing $H_i$ and without assuming it is possible to rank individuals by $H_i$ on the basis of their responses $R_i$, we have seen that the conditional distribution of $R_i$ given exogenous covariates $X_i$ can still be informative about the effects of $X$ on $H$. While this allows one to observe the sign of causal effects under the assumption that this sign is common across individuals, we've seen that different discrete conditional mean comparisons can impose different total weightings over the causal effects of individuals in the population. Simulation evidence as well as theoretical results suggest the impact of this problem for comparisons of magnitude is somewhat limited in practice, and the problem goes away entirely in the limit of continuous treatment variables. Nevertheless, the results suggest that care is warranted in comparing the magnitude of regression coefficients across explanatory variables, even when they are as good as randomly assigned.
The results of this paper suggest three practical implications for using regression analysis for causal inference with subjective ordinal outcomes. First, the critique of bondandlang that such responses are only ordinarily meaningful is most acute when comparing large, heterogeneous populations that differ among many dimensions. Isolating causal effects using exogenous variation in individual treatment variables is not subject to this critique in the sense that regression derivatives identify the sign of a convex average of causal effects, even though reporting functions are unknown to the researcher. Second, to make such regression derivatives quantitatively meaningful, researchers should focus on comparing across treatment variables when more than one is available. While non-parametric regression methods are preferred from the standpoint of identification, this is not specific to the analysis of subjective outcome variables. Finally, notwithstanding the above, researchers should exercise some caution when comparing the magnitudes of two discrete treatment effects or between a discrete treatment effect and the slope for a continuous treatment. The relative magnitudes of convex averages of causal effects can still be partially identified in such settings with further assumptions, though weakening these assumptions represents a possible avenue for future research.
\printbibliography