Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,654 characters · 14 sections · 73 citation commands
Loss aversion and the welfare ranking of policy interventions
\thispagestyle{empty}
Keywords: Welfare, Loss Aversion, Policy Evaluation, Stochastic Ordering, Directional Differentiability
JEL codes: C12, C14, I30
\setcounter{page}{1}
Policy interventions often generate heterogeneous effects, giving rise to gains and losses to different individuals and sectors of society. Classically, the welfare ranking of policy interventions, conducted under the Rawlsian principle of the “veil of ignorance”, has deemed such gains and losses irrelevant: all policies that produce the same marginal distribution of outcomes should be considered equivalent for the purpose of welfare analysis. Atkinson70,Roemer98, Sen00. However, more recent approaches have focused precisely on how different individuals are affected by a given policy HeckmanSmith98, CarneiroHansenHeckman01. Our paper relates to this latter approach and focuses on an important characteristic of the individual valuation of policy-induced effects: loss aversion.\footnote{Loss aversion is a well established empirical regularity, documented in a wide variety of contexts KahnemanTversky79, SamuelsonZeckhauser88, TverskyKahneman91, RabinThaler:01, Rick11}
There are two main reasons why loss aversion can be important for the welfare ranking of policy interventions. First, as shown in CarneiroHansenHeckman01, the individual gains and losses caused by a policy have important political economy consequences. Public support for that policy, and for the authorities that implement it, depends on the balance of gains and losses experienced and valued by different individuals in the electorate. In this context, there is mounting empirical evidence indicating that the electorate often exhibits loss aversion. This aversion to losses among constituents, in turn, drives the actions of policy makers, as documented in situations as diverse as government support to the steel industry in US trade policy, and President Trump's attempted repeal of the Affordable Care Act FreundOezden08, AlesinaPassarelli19.
Second, political economy aside, there are important situations where policy makers have strong normative reasons for incorporating loss-aversion in their own welfare ranking of public policies. This has been proposed in a range of fields. In a recent example, Eyal20 shows that the Hippocratic principle of “first, do no harm” has led US and EU policy makers to delay the development of Covid-19 vaccines by rejecting human challenge trials, that involve purposefully infecting a small number of vaccine trial volunteers. In this situation, policy makers placed more weight on the potential harm to a small number of volunteers, than on the vast potential benefits of accelerating the availability of a vaccine to large swathes of the population. Similarly, in the context of minimum wage legislation, Mankiw14 proposes that policy makers should be loss-averse in their approach to policy evaluation and adopt a “first, do no harm” principle: “As I see it, the minimum wage and the Affordable Care Act are cases in point. Noble as they are in aspiration, they fail the do-no harm test. An increase in the minimum wage would disrupt some deals that workers and employers have made voluntarily”. Along the same lines, in the context of development economics, scholars such as Easterly09 have proposed a “first, do no harm” approach to foreign aid interventions in developing countries.
In this paper, we extend the toolkit available for the evaluation of policy interventions in contexts where it is sensible to incorporate loss-aversion in the welfare ranking of policy interventions. We are not proposing that all policies should be ranked under loss aversion sensitive criteria. Instead, we provide a methodology for conducting such a ranking in cases where aversion to losses is likely to be important. As discussed above, this may be the case either because policy makers are genuinely loss-averse, or because they know that the individuals exposed to the policy are so, and this leads them to incorporate loss aversion in policy ranking. To this end, our paper develops new testable criteria and econometric methods to rank distributions of individual policy effects, from a welfare standpoint, incorporating loss aversion. We make two main contributions to the literature.
Our first contribution is to propose loss aversion-sensitive criteria for the welfare ranking of policies. We adopt the standard welfare function approach Atkinson70: alternative policies are compared based on a welfare ranking, where social welfare is an additively separable and symmetric function of individuals' outcomes. It is well established that, for non-decreasing utility functions, this is equivalent to first-order stochastic dominance (FOSD) over distributions of policy outcomes. Analogously, our ranking is based on social value functions, which are additively separable and symmetric functions of individual gains and losses. We show that the social value function ranking with non-decreasing and loss-averse value functions TverskyKahneman91 is equivalent to a new concept we call loss aversion-sensitive dominance (LASD) over distributions of policy-induced gains and losses. FOSD requires that the cumulative distribution function of the dominated distribution lies everywhere above the cumulative distribution of the dominant distribution. In contrast, under LASD, the dominated cumulative distribution function must lie sufficiently above the dominant distribution function for losses such that the probability of potential losses cannot be compensated by a higher probability for potential gains. This is a consequence of loss-aversion. Except for the special case of a status quo policy (i.e. a policy of no change) where FOSD and LASD coincide, generally, as we show, LASD can be used to compare policies that are indistinguishable for FOSD.\footnote{The literature on stochastic dominance is vast and spans economics and mathematics - we refer the reader to, e.g., ShakedShanthikumar94 and Levy16 for a review. When dominance curves cross, higher order or inverse stochastic dominance criteria have been proposed. The former involves conditions on higher (typically third and fourth) order derivatives of utility function (e.g. Fishburn80, Chew83 to which EeckhoudtSchlesinger06 provided interesting interpretation, whereas the latter is related to the rank-dependent theory originally proposed by Weymark81 and Yaari87, Yaari88, where social welfare functions are weighted averages of ordered outcomes with weights decreasing with the rank of the outcome (see AabergeHavnesMogstad18 for a recent refinement of this theory). }
The LASD criterion relies on gains and losses, which under standard identification conditions can be considered treatment effects. It is well known that the point identification of the distribution of treatment effects may require implausible theoretical restrictions such as rank invariance of potential outcomes HeckmanSmithClements97. We thus extend our LASD criteria to a partially-identified setting and establish a sufficient condition to rank alternative policies under partial identification of the distributions of their effects. We use Makarov bounds Makarov82, Rueschendorf82, FrankNelsenSchweizer87 to bound the distribution of treatment effects when the joint pre and post-policy outcome distribution is unknown. This provides a testable criterion that can be used in practice, since the marginal distribution functions from samples observed under various treatments can usually be identified and Makarov bounds only rely on marginal information for their identification.
Our second contribution is to develop statistical inference procedures to practically test the loss averse-sensitive dominance condition using sample data. We develop statistical tests for both point-identified and partially-identified distributions of outcomes. The test procedures are designed to assess, uniformly over the two outcome distributions, whether one treatment dominates another in terms of the LASD criterion. Specifically, we suggest Kolmogorov-Smirnov and Cram\'er-von Mises test statistics that are applied to nonparametric plug-in estimates of the LASD criterion mentioned above. Inference for these statistics uses specially tailored resampling procedures. We show that our procedures control the size of tests for all probability distributions that satisfy the null hypothesis. Our tests are related to the literature on inference for stochastic dominance represented by, e.g., LintonSongWhang10, LintonMaasoumiWhang05, BarrettDonald03 and references cited therein. LintonMaasoumiWhang05 is an important contribution because in addition to developing tests for stochastic dominance of arbitrary order, they propose a Prospect Theory stochastic dominance test. Their test is intended for inferring dominance among a different family of value functions than ours, namely, the focus is on risk loving for gains and risk aversion for losses (i.e. so called S-shapedness, KahnemanTversky79), but not on loss aversion. We contribute to the literature by developing tests for loss averse-sensitive dominance, which are an alternative to standard stochastic dominance tests. Our tests widen the variety of comparisons available to empirical researchers to other criteria that encode important qualitative features of agent preferences.
The LASD criterion results in a functional inequality that depends on marginal distribution functions, and we adapt existing techniques from the literature on testing functional inequalities to test for LASD. However, in comparison with stochastic dominance tests, verifying LASD with sample data presents technical challenges for both the point- and partially-identified cases. The criterion that implies LASD of one distribution over another is more complex than the standard FOSD criterion, and hence requires a significant extension of existing procedures to justify the use of inference about loss averse-sensitive dominance with a nonparametric plug-in estimator of the LASD criterion.
In particular, the problem with using existing stochastic dominance techniques is that the mapping of distribution functions to a testable criterion is nonlinear and ill-behaved. A dominance test inherently requires uniform comparisons be made, and tractable analysis of its distribution demands regularity, in the form of differentiability, of the map between the space of distribution functions and the space of criterion functions. However, the map from pairs of distribution functions to the LASD criterion function is not differentiable. Despite this complication, we show that supremum- or $L_2$-norm statistics applied to this function are just regular enough that, with some care, resampling can be used to conduct inference.
For practical implementation, we propose an inference procedure that combines standard resampling with an estimate of the way that test statistics depend on underlying data distributions, building on recent results from FangSantos19. We contribute to the literature on directionally differentiable test statistics with a new test for LASD. Recent contributions to this literature include, among others, HongLi18, ChetverikovSantosShaikh18, ChoWhite18, ChristensenConnault19, FangSantos19, CattaneoJanssonNagasawa17 and MastenPoirier17.
When distributions are only partially identified by bounds, the situation is more challenging. The current state of the literature on Makarov bounds focuses on pointwise inference for bound functions (see, e.g., FanPark10,FanPark12, FanGuerreZhu17, and FirpoRidder08,FirpoRidder19). However, the LASD criterion requires a uniform comparison of bound functions, and the map from distribution functions to Makarov bound functions is also not smooth. Fortunately the problem has a similar solution to the point-identified LASD test. The resulting $L_2$- and supremum-norm statistics allow us to conduct inference for LASD in the partially identified case using functions that bound the relevant CDFs. The details are included in the online Supplemental Appendix.
We illustrate the practical use of our proposed criteria and tests with a very simple empirical application using data from BitlerGelbachHoynes06. This aims at exemplifying the use of our approach, rather than developing a fully fledged empirical investigation. We show that, in the case of a policy with gainers and losers, the use of our loss aversion-sensitive evaluation criteria may lead to a ranking of policy interventions that differs from that obtained when their outcomes are compared using stochastic dominance.
The rest of the paper is organized as follows. Section (ref) presents the basic definitions and notation and defines loss aversion-sensitive dominance. Section (ref) develops testable criteria for loss aversion-sensitive dominance. Section (ref) proposes statistical inference methods for LASD using sample observations. Section (ref) illustrates our methodology using a very simple empirical application that uses data from the experimental evaluation of a well-known welfare policy reform in the US. Section (ref) concludes. Our first appendix includes auxiliary results and definitions; our second one collects proof of the results in the paper.
In this section, we propose a novel dominance relation for ordering policies under the assumption that social decision makers consider the distribution of individual gains and losses under different policy scenarios. We call this criterion Loss Aversion-Sensitive Dominance (LASD).
Consider a random variable $X$ with cumulative distribution function $F$. Let $\mathscr{F}$ be the set of cumulative distribution functions with bounded support $\mathcal{X}$. We maintain the assumption throughout that $F \in \mathscr{F}$. The bounded support assumption is made to avoid technical conditions on tails of distribution functions. The aim of this paper is to provide theoretical criteria and econometric methods to rank policy interventions under LASD. The decision maker's goal is to compare policies $A$ and $B$ using the distribution functions of $X_A$ and $X_B$, labeled $F_A$ and $F_B$. Because the random variables $X_A$ and $X_B$ represent gains and losses due to the enactment of policies $A$ and $B$, they represent a change between agents' pre-treatment and post-treatment outcomes. To this end, let $Z_0, Z_A$ and $Z_B$ represent an agent's potential outcome under the status quo, treatment $A$ or treatment $B$, so that we may write $X_A = Z_A - Z_0$ and $X_B = Z_B - Z_0$. We assume that the variables $(Z_0, Z_A, Z_B)$ have marginal distribution functions $(G_0, G_A, G_B)$.
It may be assumed that $(F_A, F_B)$ are identified, or that only $(G_0, G_A, G_B)$ are identified. The former case is related to a point-identified model, and the latter to a partially-identified one. Theoretical results for both situations are shown here. Inference for the first situation is considered in this paper, while partially-identified inference results under the second situation are considered in the online Supplemental Appendix.
For example, in the empirical illustration considered in Section (ref), we observe quarterly household income before the enactment of a new welfare program, which we assume to represent realizations of $Z_0$. Next, we observe income for some households under a continuation of the old program, identifying those as realizations of $Z_A$, and income for other households under a new welfare program, labeling those observations as $Z_B$ realizations. If policymakers are interested in how one household's earnings evolve over time either by staying with the old program or switching to the new program, then the levels of $Z_A$ and $Z_B$ are not of primary interest, rather the changes represented by $X_A$ and $X_B$ are (and given the longitudinal nature of our data, it is natural to assume that $X_A$ and $X_B$ are identified). Our tests, detailed in Section (ref), use the null hypothesis that a household exhibiting loss-aversion would prefer to switch to the new welfare program, and search for evidence to the contrary.
The decision maker has preferences over $X$ (not $Z$) that are represented via a continuous function.
The social value function defined above is the value assigned to the distribution of $X$ by a social planner that uses the value function $v$ to convert gains and losses into a measure of well-being GajdosWeymark12. This value function $v$ does not have to coincide with any individual's $v$ in the population: as mentioned in the Introduction, the social planner is averse to individual losses either because individuals are loss-averse themselves (political economy motivation), or because she holds normative views that imply her loss aversion towards $X$. In either case, $v$ will exhibit loss-aversion, i.e. there is asymmetry in the valuation of gains and losses, where losses are weighed more heavily than gains of equal magnitude. Furthermore, $v$ assigns negative value to losses and positive value to gains and is non-decreasing. These properties are formally listed in the next definition.\footnote{This standard interpretation of the social welfare function can be further extended. For example, individuals may be uncertain about their counterfactual outcome and form an expectation of $v(\cdot)$ given $z_0$. Then we write $W(F) = \int \int v(x, z_0) \dd F_{X|Z_0} (x|z_0) \dd F_{Z_0}(z_0)$. Denoting a new value function $v^*(z_0) = \int v(x, z_0) \dd F_{X|Z_0} (x|z_0)$, i.e. an expected value for a given $z_0$, $W(F)$ is as in Definition (ref).}$^{,}$\footnote{An interesting direction for future research is to axiomatize the class of social value functions and possibly develop measures of loss aversion based on this class. Some inspiration for axiomatization may come from the inequality and poverty measurement literature. For example, ratio scale invariance (i.e. proportional changes to the units in which gains and losses are measured do not matter), may be a powerful axiom in obtaining a specific functional form.}
The properties in Definition (ref) are typically assumed in Prospect Theory together with the additional requirement of S-shapedness of value function, which we do not consider (see, e.g., p. 279 of KahnemanTversky79). Assumptions 1 and 2 are standard monotone increasing conditions. Assumption 3 expresses the idea that “losses loom larger than corresponding gains” and is a widely accepted definition of loss aversion TverskyKahneman92. It is a stronger condition than the one considered by KahnemanTversky79.
The following form of $W(F)$ will be useful in subsequent definitions and results.
Assume that the decision maker's social value function $W$ depends on $v$ which satisfies Definition (ref), and she wishes to compare random variables $X_{A}$ and $X_{B}$ which represent gains and losses under two policies labeled $A$ and $B$. The decision maker prefers $X_A$ over $X_B$ if she evaluates $F_A$ as better than $F_B$ using her SVF --- specifically, $X_A$ is preferred to $X_B$ if and only if $W(F_A) \geq W(F_B)$, where $W$ is defined in Definition (ref). Please note that $X_A$ is preferred to $X_B$ for every $v$ that is described by Definition (ref). This is what makes dominance conditions robust criteria for comparing distributions. This idea is formalized below.
In the next section we relate this theoretical definition to a more concrete condition that depends on the cumulative distribution functions of the outcome distributions, $F_A$ and $F_B$.
In this section we formulate testable conditions for evaluating distributions of gains and losses in practice. We propose criteria that indicate whether one distribution of gains and losses dominates another in the sense described in Definition (ref).
Recall that $Z_0, Z_A$ and $Z_B$ represent an outcome before or after a policy takes effect, while $X_A$ and $X_B$ represent a change from a pre-policy state to an outcome under a policy. The challenge of comparing variables $X_A$ and $X_B$ is well known in the treatment effects literature: because $X_A$ and $X_B$ are defined by differences between the $Z_k$, $F_A$ and $F_B$ depend on the joint distribution of $(Z_0, Z_A, Z_B)$, which may not be observable without restrictions imposed by an economic model. In subsection (ref) we abstract from specific identification conditions and discusses LASD under the assumption that $F_A$ and $F_B$ are identified. In subsection (ref) we work with a partially identified case where only the marginal distribution functions $G_0$, $G_A$ and $G_B$ are identified and no restrictions are made to identify $F_A$ and $F_B$.
The LASD concept in Definition (ref) requires that one distribution is preferred to another over a class of social value functions and is difficult to test directly. The following result relates the LASD concept to a criterion which depends only on marginal distribution functions and orders $F_A$ and $F_B$ according to the class of SVFs allowed in Definition (ref). In this section we assume that $F_A, F_B \in \mathscr{F}$ are point identified. This may result from a variety of econometric restrictions that deliver identification and are the subject of a large literature.
Theorem (ref) provides two different conditions that can be used to verify whether one distribution of gains and losses dominates the other in the LASD sense.\footnote{LASD is a partial order. Over losses, (ref) is a partial order because FOSD is a partial order. For the tail condition (ref) checking transitivity we have $\rbr{1-F_{A}(x)}-F_{A}(-x) \geq \rbr{1-F_{B}(x)}-F_{B}(-x), \rbr{1-F_{B}(x)}-F_{B}(-x) \geq \rbr{1-F_{C}(x)}-F_{C}(-x)$, and $\rbr{1-F_{A}(x)}-F_{A}(-x) \geq \rbr{1-F_{C}(x)}-F_{C}(-x)$. If $F_{A}(-x)-F_{B}(-x)=0$ then $F_{A}(-x)=F_{B}(-x)$ and using it in (ref) gives anti-symmetry.} These criteria compare the outcome distributions by examining how the distribution functions $(F_A, F_B)$ assign probabilities to gains and losses of all possible magnitudes. The particular way that they make a comparison is related to the relative importance of gains and losses. Consider condition (ref). For the distribution of $X_B$ to be dominated, its distribution function must lie above the distribution of $X_A$ for losses. $X_B$ can be dominated by $X_A$ in the LASD sense even when gains under $X_A$ do not dominate $X_B$ for gains --- that is, when $F_{A}(x)-F_{B}(x) \geq 0$ for some $x \geq 0$ --- as long as this lack of dominance in gains is compensated by sufficient dominance of $X_A$ over $X_B$ in the losses region. This is a consequence of the asymmetric treatment of gains and losses. Conditions (ref) and (ref) jointly express the same idea, but they help to understand how gains and losses are treated asymmetrically in condition (ref). In the losses region, condition (ref) is a standard FOSD condition. This is a consequence of loss aversion; note that in the extreme case where only losses matter, we would have (ref). In the gains region, dominance has to be sufficiently large so that under $X_A$, the probability of gains minus the probability of losses (of magnitude $x$ or larger) is no smaller than the corresponding difference for $X_B$.\footnote{We leave 1s on both sides of inequality (ref) for this interpretation to be more evident.} Inequality (ref) combines the two inequalities represented by (ref) and (ref) into a single equation.
It is interesting to note that LASD has one property in common with FOSD, namely, a higher mean is a necessary condition for both types of dominance. This follows directly from Definitions (ref) and (ref) by using $v(x) = x$.
Note that FOSD cannot rank two distributions that have the same mean --- that is, if $F_A \succeq_{FOSD} F_B$ and $\ex{X_A} = \ex{X_B}$, then $F_A = F_B$. This is not the case for LASD, as the next example demonstrates. Therefore, for example, equation (ref) may still be used to differentiate between two distributions with the same average effect.
It is important to note that LASD is a concept that is specialized to the comparison of distributions that represent gains and losses. Standard FOSD is typically applied to the distribution of outcomes in levels without regard to whether the outcomes resulted from gains or losses of agents relative to a pre-policy state --- in our notation, $G_A$ and $G_B$ are typically compared with FOSD, instead of $F_A$ and $F_B$. FOSD applied to post-policy levels may or may not coincide with LASD applied to changes. This means that even when a strong condition such as FOSD holds for final outcomes, if one took into account how agents value gains and losses it may turn out that the dominant distribution is no longer a preferred outcome. One could apply the FOSD rule to compare distributions of income changes, which implies LASD applied to changes, because FOSD applies to a broader class of value functions. However, this type of comparison would ignore agents' loss aversion, the important qualitative feature that LASD accounts for. The following example shows that the analysis of outcomes in levels using FOSD need not correspond to any LASD ordering of outcomes in changes.
In the previous example, policy $B$ left pre-treatment outcomes unchanged, or in other words, maintained a status quo condition --- we had $X_B = Z_B - Z_0 \equiv 0$. Suppose generally that $X_B$ has a distribution that is degenerate at $0$. Then $F_B(x) = 0$ for all $x < 0$ and $F_B(x) = 1$ for all $x \geq 0$. We define this as a status quo policy distribution, labelled $F_{SQ}$. When comparison is between a distribution $F_A$ and $F_{SQ}$, LASD and standard FOSD are equivalent. The distribution that dominates $F_{SQ}$ is necessarily only gains.
In many situations of interest the cumulative distribution functions of gains and losses, $F_A$ and $F_B$, are not point identified without a model of the relationship between $X_A$ and $X_B$. However, the marginal distributions of outcomes in levels under different policies, represented by the variables $Z_0$, $Z_A$ and $Z_B$, may be identified. Without information on the dependence between potential outcomes, we can still make some more circumscribed statements with regard to dominance based on bounds for the distribution functions. This section studies the LASD dominance criterion to the case that distribution functions $F_A$ and $F_B$ are only partially identified.
A number of authors have considered functions that bound the distribution functions $F_A$ and $F_B$. Taking $X_A$ as an example, Makarov bounds Makarov82, Rueschendorf82, FrankNelsenSchweizer87 are two functions $L$ and $U$ that satisfy $L(x) \leq F_A(x) \leq U(x)$ for all $x \in \mathbb{R}$, depend only on the marginal distribution functions $G_0$ and $G_A$ and are pointwise sharp --- for any fixed $x$ there exist some $Z_0^*$ and $Z_A^*$ such that the resulting $X_A^* = Z^*_A - Z^*_0$ has a distribution function at $x$ that is equal one of $L(x)$ or $U(x)$. WilliamsonDowns90 provide convenient definitions for these bound functions. For any two distribution functions $G_1, G_2$, define
For convenience define the policy-specific bound functions for $F_k$, $k \in \{A, B\}$ and all $x \in \mathbb{R}$, which depend on the marginal CDFs $G_0$ and $G_k$, by
Using these definitions we obtain a sufficient and a necessary condition for LASD when only bound functions of the treatment effects distribution functions are observable. The next theorem formalizes the result.
Theorem (ref) is an analog of Theorem (ref) and shows what effect the loss of point identification has on the relationship between dominance and conditions on the CDFs. In particular, one loses a simple “if and only if” characterization that depends on CDFs. Instead, LASD implies a necessary condition using some bound functions, while a different sufficient condition using other bound functions implies LASD. Inference using the necessary condition shown in Theorem (ref) is discussed in the online Supplemental Appendix. We remark that there may exist other features of the joint data distribution that do not depend only on pointwise features of the CDFs of changes and would result in a necessary and sufficient condition for LASD under partial identification. That is an interesting open question but is beyond the scope of this paper.
When the comparison is with the status quo distribution, the partially identified conditions simplify. Corollary (ref) below shows what can be learned about LASD from bound functions in the partially identified case.
In this section we propose statistical inference methods for the loss aversion-sensitive dominance (LASD) criterion discussed in previous sections. We consider the null and alternative hypotheses
Under the null hypothesis (ref) policy $A$ dominates $B$ in the LASD sense, similar to much of the literature on stochastic dominance. It is a simplification of the hypotheses considered for several potential policies discussed in LintonMaasoumiWhang05, who test whether one policy is maximal, and the techniques developed below could be extended to compare several policies in the same way in a straightforward manner.\footnote{LintonMaasoumiWhang05 consider a test for Prospect Theory by testing whether the integral of one CDF dominates the other. This paper considers a different approach in which we impose loss aversion on the value function, and then derive testable conditions on the CDFs.} The null hypothesis above represents the assumption that policy $A$ is preferred by agents in the LASD sense. Rejection of the null implies that there is significant evidence for ambiguity in the ordering of the policies by LASD. Unfortunately, a drawback of the proposed procedure is that rejection of the null does not inform one about which sort of value function $v$ results in a rejection. Strong orderings of policies can result in more information, although they constrain $v$ by construction, and such exploration is left for future research.\footnote{There is also another strand of literature that develops methods to estimate the optimal treatment assignment policy that maximizes a social welfare function. Recent developments can be found in Manski04, Dehejia05, HiranoPorter09, Stoye09, BhattacharyaDupas12, Tetenov12, KitagawaTetenov18, KitagawaTetenov19, among others. These papers focus on the decision-theoretic properties and procedures that map empirical data into treatment choices. In this literature, our paper is most closely related to Kasy16, which focuses on welfare rankings of policies rather than optimal policy choice.}
We consider tests for this null hypothesis given sample data observed under two different identification assumptions. We start with the case where one can directly observe samples $\{X_{Ai}\}_{i=1}^{n_A}$ and $\{X_{Bi}\}_{i=1}^{n_B}$ which represent agents' gains and losses, or in other words, we simply assume that the distribution functions of $X_A$ and $X_B$ are point-identified and their distribution functions can be estimated using the empirical distribution functions from two samples. Next we extend these results to the partially-identified case where no assumption about the joint distribution of potential outcomes under either treatment is made. In this case, we assume that three samples are observable, $\{Z_{0i}\}_{i=1}^{n_0}$, $\{Z_{Ai}\}_{i=1}^{n_A}$ and $\{Z_{Bi}\}_{i=1}^{n_B}$, representing outcomes under a control or pre-policy state and outcomes under policies $A$ and $B$. Then tests are based on plug-in estimates for bounds for $X_A = Z_A - Z_0$ and $X_B = Z_B - Z_0$.
We consider distribution functions as members of the space of bounded functions on the support $\mathcal{X} \subseteq \mathbb{R}$, denoted $\ell^\infty(\mathcal{X})$, equipped with the supremum norm, defined for $g: \mathbb{R}^k \rightarrow \mathbb{R}^\ell$ by $\| g \|_\infty = \max_j \{ \sup_{x \in \mathbb{R}^k} |g_j(x)| \}$. For real numbers $x$ let $(x)^+ = \max\{0, x\}$. Given a sequence of bounded functions $\{g_n\}_n$ and limiting random element $g$ we write $g_n \leadsto g$ to denote weak convergence in $(\ell^\infty, \| \cdot \|_\infty)$ in the sense of Hoffman-J\o rgensen vanderVaartWellner96.
In this subsection we suppose that the pair of marginal distribution functions $F = (F_A, F_B)$ is identified. In the Online Supplemental Appendix C, we provide results extending the dominance tests to the case that distribution functions $F_A$ and $F_B$ are only partially identified.
To implement a test of the hypotheses (ref) we employ the results of Theorem (ref) to construct maps of $F$ into criterion functions that are used to detect deviations from the hypothesis $H_0$. Specifically, recalling that $(x)^+ = \max\{0, x\}$, for the point-identified case we examine maps $T_1: (\ell^\infty(\mathbb{R}))^2 \rightarrow \ell^\infty(\mathbb{R}_+)$ and $T_2: (\ell^\infty(\mathbb{R}))^2 \rightarrow (\ell^\infty(\mathbb{R}_+))^2$, defined for each $x \geq 0$ by
and
Functions $T_1(F)$ and $T_2(F)$ are designed so that large positive values will indicate a violation of the null. Taking $T_1$ as an example, Theorem (ref) states that $W(F_A) \geq W(F_B)$ if and only if $F_B(-x) - F_A(-x) \geq (F_A(x) - F_B(x))^+$ for all $x \geq 0$, so tests can be constructed by looking for $x$ where $T_1(F)(x)$ becomes significantly positive. We will refer to $T_j$ as maps from pairs of distribution functions to another function space, and also refer to them as functions.
The hypotheses ((ref)) can be rewritten in two equivalent forms, depending on whether one uses $T_1$ or $T_2$ to transform distribution functions: letting $\mathcal{X} \subseteq \mathbb{R}_+$ be an evaluation set, we have
and
In the second set of hypotheses $0_{2}$ is a two-dimensional vector of zeros and inequalities are taken coordinate-wise.
The next step in testing the hypotheses (ref) and (ref) is to estimate $T_1(F)$ and $T_2(F)$. Let $\mathbb{F}_n = (\mathbb{F}_{An}, \mathbb{F}_{Bn})$ denote the pair of marginal empirical distribution functions, that is, $\mathbb{F}_{kn}(x) = \frac{1}{n_k} \sum_{i=1}^{n_k} \mathbf{1}\{X_{ki} \leq x\}$ for $k \in \{A, B\}$. These are well-behaved estimators of the components of $F$. Letting $n = n_A + n_B$, standard empirical process theory shows that $\sqrt{n} (\mathbb{F}_n - F)$ converges weakly to a Gaussian process under weak assumptions vanderVaart98. In order to conduct inference for loss aversion-sensitive dominance, we use plug-in estimators $T_j(\mathbb{F}_n)$ for $j \in \{1, 2\}$. See Remark (ref) in Appendix (ref) for details on the computation of these functions.
In order to detect when $T_j(\mathbb{F}_n)$ is significantly positive, we consider statistics based on a one-sided supremum norm or a one-sided $L_2$ norm over $\mathcal{X}$. Kolmogorov-Smirnov (i.e., supremum norm) type statistics are
Meanwhile Cram\'er-von Mises (or $L_2$ norm) test statistics are defined by
Because all the CDFs used in these statistics belong to $\mathscr{F}$, distributions with bounded support, the integrands in the $L_2$ statistics are square-integrable.
We wish to establish the limiting distributions of $V_{jn}$ and $W_{jn}$, for $j\in\{1,2\}$, under the null hypothesis $H_0: F_A \succeq_{LASD} F_B$. Two challenges arise when considering these test statistics. First, the form of the null hypothesis as a functional inequality to be tested uniformly over $\mathcal{X}$ is a source of irregularity. Let the joint probability distribution of $(X_A, X_B)$ be denoted by $P$. Because the null hypothesis, $F_A \succeq_{LASD} F_B$, is a functional weak inequality the asymptotic distributions of the test statistics $V_j$ and $W_j$ may depend on features of $P$. This is referred to as non-uniformity in $P$ in LintonSongWhang10, AndrewsShi13, and requires attention when resampling.
Second, due to the pointwise maximum function in its definition, $T_1$ is too irregular as a map from the data to the space of bounded functions to establish a limiting distribution for the empirical process $\sqrt{n}(T_1(\mathbb{F}_n) - T_1(F))$ using conventional statistical techniques. In contrast, $T_2$ is a linear map of $F$, which implies that $\sqrt{n}(T_2(\mathbb{F}_n) - T_2(F))$ has a well-behaved limiting distribution in $(\ell^\infty(\mathbb{R}_+))^2$.\footnote{The issues of a general lack of differentiability of functions arrived at by marginal optimization and a solution for inference based on directly characterizing the behavior of test statistics applied to such functions are studied in more generality in FirpoGalvaoParker21. However, we highlight that the tests described here are extensions of the results of that paper and are specifically tailored to this application.}
Despite the above challenges, we show that $V_{jn}$ and $W_{jn}$ (for $j \in \{1, 2\}$) have well-behaved asymptotic distributions, and furthermore, that the limiting random variables satisfy $V_1 \sim V_2$ and $W_1 \sim W_2$. This is an important result because it is the foundation for applying bootstrap techniques for inference. Before stating the formal assumptions and asymptotic properties of the tests, we discuss the two difficulties mentioned above in more detail.
The limiting distributions of $V_{jn}$ and $W_{jn}$ statistics depend on features of $P$. Let $\mathcal{P}_0$ be the set of distributions $P$ such that $F_A \succeq_{LASD} F_B$. These are distributions with marginal distribution functions $F$ such that $T_j(F)(x) \leq 0$ for all $x \geq 0$. To discuss the relationship between these sets of distributions and test statistics, we relabel the two coordinates of the $T_2$ function as
and
When $P \in \mathcal{P}_0$, both $m_1(x) \leq 0$ and $m_2(x) \leq 0$ for all $x \geq 0$.
More detail is required about the behavior of the two coordinate functions to determine the limiting distributions of $V_{jn}$ and $W_{jn}$ statistics. For $L_2$-norm statistics $W_{1n}$ and $W_{2n}$, we define the following relevant subdomains of $\mathcal{X}$, which collect the arguments in the interior of $\mathcal{X}$ where $m_1$ or $m_2$ are equal to zero:
Denote $\mathcal{X}_0(P) \subseteq \mathcal{X}$ as the set of $x$ where $T_1(F)(x) = 0$ or at least one coordinate of $T_2(F)$ equals $0$ for probability distribution $P$. As will be seen below, $\mathcal{X}_0(P)$ is the same for both the $T_1$ and $T_2$ functions, and when it is non-empty, test statistics have a nondegenerate distribution. Following LintonSongWhang10, we call $\mathcal{X}_0(P)$ the contact set for the distribution $P$. Given the above definitions, under the null hypothesis we can write
On the other hand, the supremum-norm statistics $V_{1n}$ and $V_{2n}$ need a different family of sets, namely the sets of $\epsilon$-maximizers of $m_1$ and $m_2$. For any $\epsilon \geq 0$ and $k \in \{1, 2\}$, let
An important subset of $\mathcal{P}_0$ are those $P$ for which test statistics have nontrivial limiting distributions under the null hypothesis --- that is, not degenerate at 0, which occurs when there is some $x$ such that $T_j(F)(x) = 0$ (note that there are no $x$ such that $T_j(F)(x) > 0$ when $P \in \mathcal{P}_0$). Define $\mathcal{P}_{00} \subset \mathcal{P}_0$ to be the set of all $P$ such that $\mathcal{X}_0(P) \neq \varnothing$. If $P \in \mathcal{P}_0 \backslash \mathcal{P}_{00}$ then $\mathcal{X}_0(P) = \varnothing$ and because the distribution satisfies the null hypothesis, $F_A$ strictly dominates $F_B$ everywhere and the criterion functions $T_j$ are strictly negative over $\mathcal{X}$. When $P \in \mathcal{P}_0 \backslash \mathcal{P}_{00}$, test statistics have asymptotic distributions that are degenerate at zero because test statistics will detect that policy $A$ is strictly better that $B$ over all of $\mathcal{X}$. When $P \in \mathcal{P}_{00}$, $T_j(F)$ is zero over $\mathcal{X}_0(P)$ and test statistics have a nontrivial asymptotic distribution over $\mathcal{X}_0(P)$. Thus, when $F_A \succeq_{LASD} F_B$, the asymptotic behavior of test statistics depends on whether $P \in \mathcal{P}_{00}$ or $P \in \mathcal{P}_0 \backslash \mathcal{P}_{00}$. Note that when $P \in \mathcal{P}_{00}$, we have $\lim_{\epsilon \searrow 0} \mathcal{M}^k(\epsilon) = \mathcal{X}_0^k(P)$ (that is, nonstochastic convergence in the sense of Painlev\'e-Kuratowski, see, e.g., RockafellarWets98) for whichever coordinate function actually achieves the maximal value zero.
The second challenge for testing is related to the scaled difference $\sqrt{n}(T_1(\mathbb{F}_n) - T_1(F))$ as $n$ grows large. Hadamard differentiability is an analytic tool used to establish the asymptotic distribution of nonlinear maps of the empirical process. Definition (ref) in Appendix (ref) provides a precise statement of the concept. When a map is Hadamard differentiable --- for example $T_2$, which is linear as a map from $(\ell^\infty(\mathbb{R}))^2$ to $(\ell^\infty(\mathbb{R}_+))^2$ and is thus trivially differentiable --- the functional delta method can be applied to describe its asymptotic behavior as a transformed empirical process, and a chain rule makes the analysis of compositions of several Hadamard-differentiable maps tractable. Also, the Hadamard differentiability of a map implies resampling is consistent when this map is applied to the resampled empirical process vanderVaart98 --- so, for example, the distribution of resampled criterion processes $\sqrt{n}(T_2(\mathbb{F}_n^*) - T_2(\mathbb{F}_n))$ is a consistent estimate of the asymptotic distribution of $\sqrt{n}(T_2(\mathbb{F}_n) - T_2(F))$ in the space $\ell^\infty(\mathbb{R}_+)$. On the other hand, consider the $T_1$ map. The pointwise Hadamard directional derivative of $T_1(f)(x)$ at a given $x \geq 0$ in direction $h(x) = (h_A(x), h_B(x))$ is
This map, thought of as a map between function spaces, $(\ell^\infty(\mathbb{R}))^2$ and $\ell^\infty(\mathbb{R}_+)$, is not differentiable because the scaled differences $(T_1(f)(x) - T_1(f + th_t)(x)) / t$ converge to the above derivative at each point $x$, but may not converge uniformly in $\mathbb{R}_+$. Despite the lack of differentiability of the map $F \mapsto T_1(F)$, we show in Lemma (ref) in Appendix (ref) that the maps $F \mapsto V_1$ and $F \mapsto W_1$ are Hadamard directionally differentiable, which implies these maps are just regular enough that existing statistical methods can be applied to their analysis. Later in this section we apply the resampling technique recently developed in FangSantos19 along with this directional differentiability to describe hypothesis tests using $V_{1n}$ or $W_{1n}$.
Having discussed the difficulties in the relationship between distributions and test statistics, we turn to assumptions on the observations. In order to conduct inference using either $T_1(\mathbb{F}_n)$ or $T_2(\mathbb{F}_n)$ we make the following assumptions.
Under these assumptions we establish the asymptotic properties of the test statistics under the null and fixed alternatives. Under the above assumptions, there is a Gaussian process $\mathcal{G}_F$ such that $\sqrt{n}(\mathbb{F}_n - F) \leadsto \mathcal{G}_F$. We denote each coordinate process $\mathcal{G}_{F_A}$ and $\mathcal{G}_{F_B}$, and for convenience define two transformed processes: for each $x \geq 0$ let
These will be used in the theorem below.
Theorem (ref) derives the asymptotic properties of the proposed test statistics. Parts 1 and 2 establish the weak limits of $V_{jn}$ and $W_{jn}$ for $j\in\{1,2\}$ when the null hypothesis is true. Recall that when $P \in \mathcal{P}_{00}$, $\lim_{\epsilon \searrow 0} \mathcal{M}^k(\epsilon) = \mathcal{X}_0^k(P)$, which is why $\mathcal{M}^k(\epsilon)$ terms are absent in the first part of the theorem. Remarkably, the test statistics using $T_1$ and $T_2$ criterion processes have the same asymptotic behavior despite the different appearances of the underlying processes and the irregularity of $T_1$. Part 3 shows that the statistics are asymptotically degenerate at zero when the contact set is empty, that is, when $P$ lies on the interior of the null region. Part 4 shows that the test statistics diverge when data comes from any distribution that does not satisfy the null hypothesis.
The limiting distributions described in Part 1 of Theorem (ref) are not standard because the distributions of the test statistics depend on features of $P$ through the $\mathcal{X}_0(P)$ terms in each expression. Therefore, to make practical inference feasible, we suggest the use of resampling techniques below.
The proposed test statistics have complex limiting distributions. In this subsection, we present resampling procedures to estimate the limiting distributions of both $V_{jn}$ and $W_{jn}$ for $j\in\{1,2\}$ under the assumption that $P \in \mathcal{P}_{00}$. Naive use of bootstrap data generating processes in the place of the original empirical process suffers from distortions due to discontinuities in the directional derivatives of the maps that define the distributions of the test statistics. In finite samples the plug-in estimate will not find, for example, the region where $F_A(x) - F_B(x) = 0$, where the derivatives exhibit discontinuous behavior. Our procedure involves making estimates of the derivatives involved in the limiting distribution and a standard exchangeable bootstrap routine, as proposed in FangSantos19.\footnote{Given a set of weights $\{W_i\}_{i=1}^n$ that sum to one and are independent of $\{X_i\}_{i=1}^n$, the exchangeable bootstrap measure is a randomly-weighted measure that puts mass $W_i$ at observed sample point $X_i$ for each $i$. This encompasses, for example, the standard bootstrap, $m$-of-$n$ bootstrap and wild bootstrap. See Section 3.6.2 of vanderVaartWellner96 for more specific details.}
In order to estimate contact sets, define a sequence of constants $\{a_n\}$ such that $a_n \searrow 0$ and $\sqrt{n}a_n \rightarrow \infty$ and let $\hat{m}_{1n}(x) = \mathbb{F}_{An}(-x) - \mathbb{F}_{Bn}(-x)$ and $\hat{m}_{2n}(x) = \mathbb{F}_{An}(-x) - \mathbb{F}_{Bn}(-x) + \mathbb{F}_{An}(x) - \mathbb{F}_{Bn}(x)$. Then for $W_j$ statistics define estimated contact sets by
When both sets are empty, replace both estimates by $\mathcal{X}$, as suggested in LintonSongWhang10 to ensure nondegenerate bootstrap reference distributions. Meanwhile, for $V_j$ statistics define estimated $\epsilon$-maximizer sets. For a sequence of constants $\{b_n\}$ such that $b_n \searrow 0$ and $\sqrt{n} b_n \rightarrow \infty$, let
Although the null hypothesis may imply that the maximum $m_1(x)$ is zero, the above formulas use the maximum of the sample analog without setting its maximum equal to zero, which is important for ensuring non-empty set estimates. Using these estimates, the distributions of $V_1$ and $W_1$ can be estimated from sample data (recall that Part 2 of Theorem (ref) asserts that these are the same distributions as those of $V_2$ and $W_2$). We conducted simulation experiments to choose these parameters using a few simulated data-generating processes, which are briefly discussed in the appendix in the context of simulations that suggest that the resulting tests have correct size and good power. Scaling the estimated processes by their pointwise standard deviation functions when estimating contact sets as in LeeSongWhang18 might result in better performance when distribution functions are evaluated near their tails, but we leave that rather complex topic for future research.
Resampling routine to estimate the distributions of $V_{jn}$ and $W_{jn}$ for $j = 1, 2$:
Next repeat the following two steps for $r = 1, \ldots, R$:
Finally,
The formulas in part 3 of the steps above are obtained by inserting estimated contact sets and resampled empirical processes in the place of population-level quantities into the functions shown in part 1 of Theorem (ref).
The resampled statistics are calculated by imposing the null hypothesis and assuming that the region $\mathcal{X}_0^j(P)$ is the only part of the domain that provides a nondegenerate contribution to the asymptotic distribution of the statistic under the null. The two cases of each part in the maximum arise from trying to impose the null behavior on the resampled supremum norm statistics, even when it appears the null is violated based on the value of the sample statistic. A simple alternative way to conduct inference would be to assume the least-favorable null hypothesis that $F_A \equiv F_B$, and to resample using all of $\mathcal{X}$. However, this may result in tests with lower power LintonSongWhang10 --- power loss arises in situations where $\mathcal{X}_0(P) \subset \mathcal{X}$ (strictly), so that the $T_j$ process is only nondegenerate on a subset, while bootstrapped processes that assume $\mathcal{X}_0(P) = \mathcal{X}$ would look over all of $\mathcal{X}$ and result in a stochastically larger bootstrap distribution than the true distribution.
The next result shows that our tests based on the resampling schemes described above have accurate size under the null hypothesis. In order to metrize weak convergence we use test functions from the set $BL_1$, which denotes Lipschitz functions $\mathbb{R} \rightarrow \mathbb{R}$ that have constant 1 and are bounded by 1.
The result in above theorem is stated in terms of the limiting variables $V_1$ and $W_1$ and bootstrap analogs. $V_1$ and $W_1$, using the functional delta method, are Hadamard directional derivatives of a chain of maps from the marginal distribution functions $F$ to the real line, and the derivatives are most compactly expressed as the definitions in Theorem (ref).
The bootstrap variables combine conventional resampling with finite-sample estimates of the maps defined in Part 1 of Theorem (ref), which is a resampling approach proposed in FangSantos19. Their result is actually more general --- it states that with a more flexible estimator $V_n^*$, we would obtain bootstrap consistency for $P$ in the null and alternative regions. Because our focus is on testing $F_A \succeq_{LASD} F_B$, however, our resampling scheme, and Theorem (ref), are done under the imposition of the null hypothesis. The resampling consistency result in Theorem (ref) implies that our bootstrap tests have asymptotically correct size for all probability distributions in the null region, in the same sense as was stressed in LintonSongWhang10. A formal statement showing size control over all of $\mathcal{P}_0$ is given in Theorem (ref) in Appendix (ref). Along with Part 4 of Theorem (ref), Theorem (ref) additionally implies that our tests are consistent, that is, that their power to detect violations from the null represented by fixed alternative distributions tends to one. This is because the resampling scheme produces asymptotically bounded critical values, while the test statistics diverge under the alternative.
The behavior of bootstrap tests under the null and alternatives is most easily examined using distributions local to $P$. We consider sequences of distributions $P_n$ local to the null distribution $P$ such that for a mean-zero, square-integrable function $\eta$, $P_n$ have distribution functions $F_n$ (where $P$ has CDF $F$) that satisfy
The behavior of the underlying empirical process under local alternatives satisfies Assumption 5 of FangSantos19 in a straightforward way Wellner92.
In the Supplemental Appendix we provide Monte Carlo numerical evidence of the finite sample properties of both point- and partially-identified methods. The simulations show that tests have empirical size close to the nominal, and high power against selected alternatives.
In this section we briefly illustrate the use of our approach using household-level data from a well-known experimental evaluation of alternative welfare programs in the state of Connecticut, documented in BitlerGelbachHoynes06. Aid to Families with Dependent Children (AFDC) was one of the largest federal assistance programs in the United States between 1935 and 1996. It consisted of a means-tested income support scheme for low-income families with dependent children, administered at the state level, but funded at the federal level. Following criticism that this program discouraged labor market participation and perpetuated welfare dependency, the Clinton administration enacted the 1996 Personal Responsibility and Work Opportunity Reconciliation Act (PRWORA), requiring all US states to replace AFDC with a Temporary Assistance for Needy Families (TANF) program. TANF programs differed amongst US states and were all fundamentally different from AFDC: they included strict time limits for the receipt of benefits and, simultaneously, generous earnings disregard schemes to incentivise work.
Under the policy framework of TANF, the state of Connecticut launched its own program, called Jobs First (JF) in 1996: this included the strictest time limit and also the most generous earnings disregard of all the US states. Nonetheless, there was a transition period during which a policy experiment was conducted by the Manpower Demonstration and Research Corporation (MDRC). A random sample of approximately 5000 welfare applicants was randomly assigned to one of two groups: half of them were assigned to JF and faced its eligibility and program rules; the other half were randomly assigned to AFDC (the program that JF aimed to replace in the state of Connecticut), thereby facing AFDC eligibility and program rules.
The MDRC experimental data include rounded data on quarterly income for a pre-program assignment period and also for a post-program assignment period, thereby allowing one to quantify and compare the income gains and losses experienced by the households that were randomly assigned to JF and ADFC\footnote{BitlerGelbachHoynes06 conduct a test comparing features of households before random assignment and find that they do not differ significantly in terms of observable characteristics. We check additionally that the income distributions were the same before the experiment split households among the two policies. We use a conventional two-sided Cram\'er-von Mises test for the equality of distributions. The statistic was approximately $0.78$ and its p-value was $0.55$, implying that before the experiment, the distributions are indistinguishable.} BitlerGelbachHoynes06 use these experimental data to compare the distribution of income between the beneficiaries of AFDC and JF. They find that while JF made the majority of individuals better-off, it also made a significant number of worse-off, especially after the JF time limit kicks in and becomes binding.\footnote{BitlerGelbachHoynes06 focus on quantile treatment effects (QTEs). If QTEs were to be used as a measure of the impact on any individual household in a welfare comparison, it would require the assumption of rank invariance across potential outcome distributions, which would be quite strong. Note that BitlerGelbachHoynes06 do not make this assumption.} In our simple empirical illustration we draw on BitlerGelbachHoynes06 and consider “AFDC” and “JF” as our alternative policies (equivalent to policies $A$ and $B$ in the previous sections). We illustrate our methods by constructing a LASD partial order to support a policy choice between these two programs\footnote{Although not directly relevant for our empirical illustration, it can be mentioned that the debate on the replacement of AFDC by TANF combined political economy concerns and also normative considerations about the appropriateness of policy-makers causing income losses to parts of the population. AlesinaGlaeserSacerdote01 use the AFDC as an empirical proxy for the generosity of the welfare state in the US and show that changes to this program had the potential to sway the electorate. At the same time, normative arguments supporting policy-makers' loss-aversion have also been put forth in this context. Peter Edelman, then a senior advisor to President Clinton, resigned in protest against this policy change, calling the replacement of ADFC by TANF a "crucial moral litmus test", as it risked causing important income losses to some households.} and comparing it with the partial ordering that would emerge if loss aversion were not taken into consideration using conventional first order stochastic dominance (FOSD).\footnote{Because assignment is random, we assume that the distribution functions of gains and losses under each policy, $F_{JF}$ and $F_{AFDC}$, are point-identified by the differences in incomes before and after random assignment.} Along the lines of BitlerGelbachHoynes06 we make this comparison separately for the time period up until the JF time limit for the receipt of welfare benefits becomes binding and for the period after that.
To make welfare decisions in terms of gains and losses, we use data on household income changes, i.e. the difference between households' income after exposure to the program (JF or AFDC) and before exposure to that program. We make this analysis separately for the period before the JF time limit become binding (TL) and for after that. We thus call pre-TL observations those that were made after random assignment to either of the policies (JF or AFDC) but before the time limit; we call post-TL observations those made after the JF time limit. We summarize household income (for both policies and pre/post TL periods) by averaging income over all quarters in the relevant time span.\footnote {We explored alternative definitions of our outcome of interest such as using the final quarter within the time span; generally these led to the same results, so we will not show them for the purpose of this simple illustration. } Changes in household income due to the AFDC and JF policies were defined as the natural logarithm of the average household income in all post-policy quarters (either JF or AFDC) minus the natural log of the average pre-policy quarterly household income. Thus, our analysis applies LASD to these changes in two separate periods, the pre-TL period and the post-TL one.
The left-hand side of Table (ref) shows the results of formal tests of the hypothesis ((ref)) using $W_{2n}$ statistics (Cram\'er-von Mises statistics applied to the empirical $T_2$ process).\footnote{Results for the other test statistics are qualitatively the same. They are collected in an Online Supplemental Appendix.} For the pre-time limit period we cannot reject the hypothesis that $F_{JF} \succeq_{LASD} F_{AFDC}$.\footnote{For this time period we cannot even reject the null of equality in the distributions of changes in income between households assigned to JF and AFDC.} However, everything changes when we make this comparison taking into account the post-time limit period. As mentioned above, after the time limit becomes binding, BitlerGelbachHoynes06 show that a sizeable number of households in JF experience total income losses, as they stop receiving welfare transfers; this does not happen amongst households on ADFC, which does not have a time limit. In order to rank the distribution of income changes under JF and AFDC using LASD we test the hypothesis that $F_{JF} \succeq_{LASD} F_{AFDC}$. As shown in the left-hand side of Table 1, this hypothesis is rejected for every significance level, reflecting the greater weight placed on the income losses experienced by JF beneficiaries.
To investigate how this rejection occurs, Figure (ref) displays the CDFs of gains and losses under the AFDC and JF policies around the JF time limit, then the way that the two $T_2$ coordinate processes compare them --- when looking at the coordinates in equation (ref), large positive values correspond to a rejection of the hypothesis $F_{JF} \succeq_{LASD} F_{AFDC}$. The positive parts of the $m_1$ and $m_2$ functions illustrated in the middle and right-hand plots of Figure (ref) are squared and integrated over estimated contact sets to arrive at the test statistic in the lower left of Table (ref). It can be seen in the second and third panels that the presumable reason that the JF policy does not dominate the AFDC policy using LASD is because the distribution of small gains and losses is more appealing in the AFDC program and the relation between small gains and small losses is preferable to JF.
What difference would it make if loss-aversion had been left out of this welfare ordering of social policies? In order to address this question we compare the welfare ordering obtained in the previous section with that obtained by ordering JF and AFDC according to first order stochastic dominance (FOSD). Using FOSD, the only relevant comparison is between the post-policy household income under JF and AFDC (household income before exposure to these policies is not material). We thus define our outcome of interest in levels, i.e. the natural log of the average household income under JF or AFDC. As before, we do this analysis separately for the two relevant time periods: before the JF time limit becomes binding and after it does.
The right-hand side of Table 1 shows the result of our FOSD tests. Either way post-policy outcomes are measured, we cannot reject the null that $G_{JF} \succeq_{FOSD} G_{AFDC}$. When measurements are made before and after exposure to the policy this is unsurprising, as prior to the JF time limit becoming binding none of the policies produces large income losses. However, even after the JF time limit becomes binding, the FOSD test still does not allow us to reject $G_{JF} \succeq_{FOSD} G_{AFDC}$, while the LASD test would lead us to categorically reject the dominance of JF over AFDC. This simple empirical illustration shows that, in practice, the consideration of loss aversion can change the welfare ordering of social policies.
Figure (ref) shows an analogous investigation into the way analysis would typically be conducted using first order stochastic dominance to compare outcomes, using contact sets as in LintonSongWhang10. The contact set was estimated using $a_n = 4\log(\log(n))$, corresponding to the tuning parameter choice of that paper, for both the FOSD and LASD tests (the smaller sequence $c_n = \sqrt{\log(\log(n))}$ was used for estimating near-maximizing sets in LASD tests). The left-hand plot in the figure shows the empirical distribution functions of outcomes under each program after the JF time limit. That is, the functions are based on levels of income rather than changes in income. The scaled difference $\sqrt{n}(\mathbb{G}_{n,JF} - \mathbb{G}_{n,AFDC})$ is displayed as the criterion function in the right-hand side of the panel, and the square of the positive part of this function is integrated over an estimated contact set. As the lower-right test in Table (ref) indicates, these differences are sometimes mildly positive, so that $G_{AFDC}$ is occasionally below $G_{JF}$ (especially at lower income levels) but the difference is not large enough to indicate a rejection of the hypothesis that $G_{JF} \succeq_{FOSD} G_{AFDC}$, as indicated by the p-value of the test. Because outcomes are measured in levels, there is no way to measure whether they represent gains or losses for agents, and so a simple difference is used here instead of the comparison that accounts for loss aversion used with changes in income.
Public policies often result in gains for some individuals and losses for others. We define a social preference relation for distributions of gains and losses caused by a policy: loss aversion-sensitive dominance (LASD). We relate these social preferences to criteria that depend solely on distribution functions. The assumption of loss aversion can lead to a welfare ranking of policies that is different from the one that would be brought about if classic utility theory and first-order stochastic dominance were used. We then propose empirically testable conditions for LASD based on our CDF-based criterion functions. Because data may come as differences between underlying random variables, we propose a point-identified version of these conditions and also a partially identified analog. We develop inference methods to formally test LASD relations and derive the corresponding statistical properties. We show that resampling techniques, tailored to specific features of the criterion functions, can be used to conduct inference. Finally, : illustrate our LASD criterion and inference methods with a simple empirical application that uses data from a well known evaluation of a large income support policy in the US. This shows that the ranking of policy options depends crucially on whether changes or levels are used and whether or not one takes individual loss aversion into account.