Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,054 characters · 16 sections · 122 citation commands
Sharp Bounds in the Latent Index Selection Model
\abstract{ A fundamental question underlying the literature on partial identification is: what can we learn about parameters that are relevant for policy but not necessarily point-identified by the exogenous variation we observe? This paper provides an answer in terms of sharp, analytic characterizations and bounds for an important class of policy-relevant treatment effects, consisting of marginal treatment effects and linear functionals thereof, in the latent index selection model as formalized in Vytlacil vytlacil. The sharp bounds use the full content of identified marginal distributions, and analytic derivations rely on the theory of stochastic orders. The proposed methods also make it possible to sharply incorporate new auxiliary assumptions on distributions into the latent index selection framework. Empirically, I apply the methods to study the effects of Medicaid on emergency room utilization in the Oregon Health Insurance Experiment, showing that the predictions from extrapolations based on a distribution assumption (rank similarity) differ substantively and consistently from existing extrapolations based on a parametric mean assumption (linearity). This underscores the value of utilizing the model's full empirical content in combination with auxiliary assumptions. }
\singlespacing Keywords: instrumental variables, latent index selection model, marginal treatment effects, partial identification, sharp bounds, counterfactuals, extrapolation, second-order stochastic dominance, stochastic ordering, majorization, Oregon Health Insurance Experiment\\[0.1in] JEL Classification: C14, C26, C50
\onehalfspacing
The method of instrumental variables (IV) forms a cornerstone of modern econometrics. In the basic program evaluation problem with imperfect compliance, Imbens and Angrist imbensangrist show that IV yields an average causal effect of treatment on outcomes when the instrument shifts treatment monotonically; this is the influential Local Average Treatment Effect ($\text{LATE}$) on the subgroup of compliers, whose treatment decisions depend on the instrument. By definition, the group of compliers --- and thus the identified causal effect --- are inextricably tied to the instrument at hand.
As a consequence, the variation that identifies a quantity of interest is not always the variation that is observed. For example, typical quasi-experimental settings do not provide a choice of instrument --- the instrument may be too coarse, or it may not correspond to a counterfactual of interest. Even when such a choice is available, as in the design of experiments, it is often constrained in practice by other considerations such as statistical power or limited resources. Furthermore, assuming that a treatment effect of interest is equal to the one that is identified precludes selection on unobservables, which belies the purpose of the model. This motivates an alternative structural approach to identification based on a marginal treatment effect ($\text{MTE}$) function in an empirically equivalent latent index selection model [Heckman and Vytlacil hv1999, hv2005; Vytlacil vytlacil]. A fundamental question, raised recently within this paradigm by Mogstad, Torgovitsky, and Santos mogstad, is the following: what can we learn about parameters that are relevant for policy but not necessarily point-identified by the exogenous variation we observe? \footnote{ Of course, the question transcends this paradigm. For previous sharp bounds on average treatment effects under exclusion restrictions, see Manski manski1989, manski, manski94, manski03, Balke and Pearl balkepearl, and Heckman and Vytlacil hv2000; for additional sharp bounds on the potential outcome distribution under varying exclusion restrictions, see Kitagawa kitagawa09; for sharp bounds on effects on the treated, see Huber et al. huber. Other related work considers similar exclusion and monotonicity assumptions but derives bounds on different parameters or in different frameworks. For respective examples, Tian and Pearl tianpearl2000 derive bounds on probabilities of causation, and Zhang and Rubin zhangrubin2003 derive bounds on treatment effects in a model with truncation by “death,” where some potential outcomes are undefined in the traditional sense. } The question nests a variety of applications, including inference about counterfactuals, inference about internally valid parameters that may not be point-identified, and specification tests.
This paper provides an answer in terms of sharp, analytically-derived bounds for an important class of target parameters in the latent index selection model, as formulated in Vytlacil vytlacil. Thus the paper builds most directly on the recent yet already influential contribution of Mogstad et al. mogstad, who introduce a computational framework for counterfactual inference that utilizes --- and exhausts --- the identifying power of observed conditional means for a closely related class of target parameters. A starting point of the present contribution is that consistency with conditional means --- henceforth, mean consistency --- does not exhaust the identifying power of the fully independent latent index selection model, which is frequently applicable in settings with (quasi)random assignment. In the absence of further assumptions or structure, mean consistency only provides information about point-identified mean outcomes and the $\text{LATE}$ among compliers. In addition to means, the empirically equivalent $\text{LATE}$ model identifies distributions of potential outcomes among certain subpopulations [Imbens and Rubin imbensrubin, Abadie abadie]. Consistency with identified distributions --- henceforth, distribution consistency --- imposes additional restrictions on target parameters, yielding some information about parameters other than the observed means and the $\text{LATE}$, even in the basic model.
As a motivating example, suppose an experiment subsidized treatment relative to a status quo, and a policy-maker is now interested in counterfactual local average treatment effects under intermediate subsidies. In particular, consider a policy that is expected to reduce treatment takeup by, say, 10% among experimental compliers. Information about conditional means does not bound the local average treatment effect under the policy in the latent index selection model without further structure or assumptions. Nevertheless, the model imposes such bounds because it identifies the marginal distributions of potential outcomes among compliers. What matters then for the policy-relevant treatment effect is which 10% of compliers “drop out” of treatment relative to the experiment subsidy. In one extreme, the dropouts are those in the highest decile of the complier treated outcome distribution and in the lowest decile of the untreated outcome distribution; in the other extreme, the reverse is true. This pair of extreme scenarios yield the sharp bounds on the effect of the policy because the model does not impose any bounds on the joint distribution of complier outcomes beyond consistency with the identified marginals.
This basic intuition of “trimming” identified outcome distributions is related to Lee lee. In a model of sample selection where a randomized treatment shifts sample selection monotonically, he derives sharp bounds for the average treatment effect among the always-sampled by thus bounding their average treated outcome. One significant difference in the present model with imperfect compliance is that the trimming intuition is “two-sided” because neither average potential outcome need be point-identified among a given subgroup.
The intuition is also generalized to analytically obtain sharp bounds on a class of weighted average treatment effects --- technically, linear functionals of the $\text{MTE}$ function --- via the theory of stochastic orders. \footnote{ Given Vytlacil's vytlacil equivalence between monotonicity and a latent index selection model, this approach could also be useful for bounding a similar class of treatment effects in the sample selection framework of Lee lee when the selection equation is more explicitly modeled. } More specifically, the sharp set of possible marginal treatment effect functions is characterized in terms of a mean consistency and a second-order stochastic dominance condition relative to the distribution of countermonotonic treatment effects among compliers ((ref)). The importance of the countermonotonic effect (among compliers) is that it pairs the highest treated outcomes with the lowest untreated outcomes and vice versa, thus yielding the extremal treatment effect distribution consistent with the identified marginals from which all candidate $\text{MTE}$ functions can be generated. Sharp lower and upper bounds on weighted average treatment effects are obtained by rearranging an extremal marginal treatment effect function to be co- or counter-monotonic with the weighting function ((ref)). The first result relies on properties of the convex stochastic order, and the second also on properties of the supermodular stochastic order. \footnote{ For a textbook reference on stochastic orders, see Shaked and Shanthikumar shakedshanthikumar. }
These results also relate and contribute to a literature on the identification of the treatment effect distribution and its functionals in the case of perfect compliance [Heckman et al. hsc1997, Fan and Park fanpark2010, Firpo and Ridder firporidder2019], because in this case the quantile function of treatment effects constitutes a possible $\text{MTE}$ function. I elaborate on this connection in (ref) after a presentation of the main results. Furthermore, these results naturally suggest a complete graphical representation in terms of uniformly sharp bounds ((ref)), which relates the quality of counterfactual inference to a well-known measure of statistical dispersion (the Gini coefficient) and yields a specification test for further parametric assumptions. \footnote{ I focus on the recent linear parameterization of Brinch et al. brinch, which is an instance of a broader control function approach [Heckman and Robb heckmanrobb, Bj\"orklund and Moffitt bjorklundmoffitt; see also Kline and Walters klinewalters]. }
A central feature of the basic latent index selection model is its lack of assumptions on the relationship between selection and potential outcomes. Unsurprisingly, this can translate into weak inference over counterfactual parameters, especially as they become increasingly different from the ones point-identified by the observed data. For example, reconsider the problem of bounding the effects of intermediate subsidies. The data-generating processes that attain the non-parametric bounds exhibit two extreme properties, which may be implausible in many applications. First, potential outcomes among compliers are perfectly countermonotonic: those with the highest realized outcomes if treated would have obtained the lowest outcomes if not treated, and vice versa. Second, compliers (negatively) select perfectly into treatment based on the resulting extremal treatment effects. Additionally, the basic model does not provide information on counterfactual treatment effects beyond compliers because one of the underlying potential outcomes is always unobserved. Economic theory suggests a number of auxiliary assumptions a researcher may consider in order to sharpen the non-parametric bounds. In their recent contribution, Mogstad et al. mogstad incorporate a variety of conditional mean assumptions into their computational framework in order to tighten partial identification.
Another contribution of this paper is that the basic $\text{MTE}$ characterizations can be used to incorporate and exhaust the content of additional distribution assumptions, as evidenced by two simple extensions. \footnote{ In a previous draft [Marx marx2019], I also propose a method for sharply incorporating conditional mean assumptions into the fully independent latent index selection model, using the representation of (ref). } In the first extension, I aggregate over observed covariates and provide a sharp policy-relevant interpretation of the aggregated bounds. Covariate information improves identification using marginal distributions by restricting the set of feasible joint distributions of potential outcomes. In the second extension, I sharply incorporate the rank similarity assumption of Chernozhukov and Hansen chernozhukovhansen. Incorporating this assumption into the latent index framework confers several advantages. First, the latent index framework defines a variety of policy-relevant treatment effects, which are typically not point-identified under rank similarity alone, even when the unconditional potential outcome distributions are point-identified. Second, the added structure of the framework permits an alternative identification strategy --- intuitively, extrapolation from compliers --- that relaxes the need for continuity of outcomes in the existing method. This is important for the subsequent empirical application, in which some outcomes are discrete and all outcomes have a mass point at zero. Third, the added empirical content of the rank similarity assumption is again summarized compactly and graphically in terms of uniformly sharp bounds [as in (ref)], even though conditional marginal distributions may no longer be fully identified. Thus the framework provides simple ways to compare or combine further assumptions with rank similarity, as well as providing an alternative test of (unconditional) rank similarity, as in Dong and Shen dongshen and Frandsen and Lefgren frandsenlefgren2018.
Empirically, I apply the methods using publicly available data from the Oregon Health Insurance Experiment (OHIE) [Finkelstein finkelstein-data]. The OHIE, which extended Medicaid to a randomly selected group of uninsured, low-income adults in Oregon, is unique in the experimental evidence it provides on the impacts of insurance coverage on health outcomes. Yet many questions of policy interest --- for example, the effects of broader expansions, encapsulated in proposals of “Medicare for All” --- require extrapolation of point-identified treatment effects to different subpopulations. Specifically, I revisit the question of how insurance coverage affects emergency room (ER) utilization. Previous analyses using the OHIE data find a positive effect of health insurance on ER visits across a range of complier subpopulations [Taubman taubman], yet predict negative effects across various measures of ER utilization for expansions of Medicaid to the never-treated [Kowalski kowalski2016]; these predictions are based on the linear extrapolation method of Brinch et al. brinch. My main empirical contribution is to show that an extrapolation based on the distribution assumption of rank similarity predicts positive (in one case, nonnegative) average treatment effects among the never-treated, which for some outcomes are even larger than the average treatment effect among compliers. Thus distribution information and assumptions may provide substantively different predictions than the ones obtained from mean information and assumptions alone.
The paper proceeds as follows. (ref) reviews the latent index selection model and the identification problem. (ref) provides the main results and relates them in more detail to the previous literature. (ref) extends the results to auxiliary information and assumptions. (ref) provides an empirical application using the publicly available data from the Oregon Health Insurance Experiment. (ref) concludes.
In what follows, random variables are defined on a common probability space $(\Omega, \mathcal{F}, P)$ with (cumulative) distribution function $F$. For a given random variable $X$, let $Q_X: [0,1] \to \mathds{R}$ denote its quantile function, $Q_X (q) = \inf \{ x: F_X (x) \geq q \}$, and let $I_X: [0,1] \to \mathds{R}$ denote its integrated quantile function (IQF), $ I_X (q) = \int_0^q \, Q_X (s) \, ds$. A random variable $X$ is said to second-order stochastically dominate another random variable $X'$, written $X \succeq_{SSD} X'$, if $I_X (q) \geq I_{X'} (q)$ for all $q \in [0,1]$. \footnote{ This definition is equivalent to the perhaps better-known SSD characterization of Rothschild and Stiglitz rothschildstiglitz70 in terms of a pointwise ordering of integrated distribution functions, $\int_{-\infty}^x F_X (r) \, dr$. Equivalence follows from two facts: i) the integrated distribution and integrated quantile functions are convex conjugates of one another [Ogryczak and Rusczczy\'nski, or2002] and ii) convex conjugation is order-reversing. } Two random variables $X$ and $X'$ defined on the same probability space are comonotonic if $(X, X') \stackrel{d}{=} (Q_X (V), Q_{X'} (V))$ and countermonotonic if $(X, X') \stackrel{d}{=} (Q_X (V), Q_{X'} (1-V))$, where $(X,X')$ denotes the random vector, $\stackrel{d}{=}$ denotes equality in distribution, and $V \sim U[0,1]$. \footnote{ Equivalently, the two random variables are comonotonic if: \[ F_{(X,X')} (x,x') = \min \{ F_{X} (x), F_{X'} (x') \} \] and countermonotonic if: \[ F_{(X,X')} (x,x') = \max \{ F_X (x) + F_{X'} (x') - 1, 0 \} \] Note these are the upper and lower Fr\'echet-Hoeffding copula bounds, respectively. } In words, comonotonic random variables can be thought of as nondecreasing functions of a single random variable $V$ and thus exhibit perfect positive dependence; analogously, countermonotonic random variables exhibit perfect negative dependence. The second-order stochastic dominance relation and monotonic random variables will play an important role in subsequent characterizations.
Proceeding with the variables of the model, let $D$ denote the binary realized treatment decision and let $Z$ denote an exogenous correlate of the treatment decision, or instrument. For simplicity of notation and to crystallize the intuition of the partial identification approach, I henceforth assume that the instrument $Z$ is also binary. For example, this occurs in policy evaluation settings with noncompliance, e.g. (ref). In this case $Z$ is the treatment assignment, and I henceforth adopt this terminology. Let $D_z$ denote the potential treatment decision that would be observed if the treatment assignment were exogenously set to $z$. Note that the realized treatment decision coincides with the potential decision for the realized treatment assignment, $D = D_Z$. Let $\bar{p}_z = \mathds{E} [ D | Z = z]$ denote the propensity score (or choice probability) as a function of treatment assignment.
Let $Y$ denote the realized outcome, and let $Y_{d,z}$ denote the potential outcome that would be observed if the treatment decision and assignment were exogenously set to $(d,z)$. Henceforth I proceed under the exclusion restriction that the treatment assignment affects the outcome only through its effect on the treatment decision. Therefore suppress the assignment subscript $z$ and denote the potential outcomes by $Y_d$. As previously, note that the realized outcome coincides with the potential outcome for the realized treatment decision, $Y = Y_D$. Let $\Delta = Y_1 - Y_0$ denote the treatment effect. The treatment effect is the primary object of interest to the researcher.
An identification challenge exists because the researcher only observes the treatment assignment and the realized treatment decision and outcome.
In other words, for (almost) every individual, the researcher observes either the treated outcome or the untreated outcome. Furthermore, the outcomes and the treatment effect may be correlated with the treatment decision through unobservables $U$.
The latent index selection model [Heckman and Vytlacil hv1999,hv2005] non-parametrically identifies a variety of weighted average treatment effects with a sufficiently rich source of variation. As shown by Vytlacil vytlacil, its basic assumptions are equivalent to those of the LATE model [Imbens and Angrist imbensangrist]. However, the selection model's structural underpinnings make it particularly useful for conducting counterfactual policy inference.
(ref) constitutes the basic formulation of Vytlacil vytlacil in the case of a binary instrument, with the normalization that unobservables $U$ follow a standard uniform distribution. As discussed in Vytlacil vytlacil, this normalization is technically innocuous given the other assumptions; at the same time, it underlies the marginal treatment effect approach to identification introduced by Heckman and Vytlacil hv2005 and pursued in this paper. I now turn to this approach.
Many causal questions of policy relevance have answers in terms of (weighted) averages of the treatment effect $\Delta$ across some group of unobservables $\mathcal{U} \subseteq [0,1]$. For a given unobservable $\mathcal{U} = \{ u \}$, Heckman and Vytlacil hv2005 define the marginal treatment effect:
The $\text{MTE}$ function has choice-theoretic foundations and is of interest in its own right. For example, in the OHIE application of (ref), $\text{MTE}(u)$ is the expected change in spending (ER charges) from health insurance coverage for unobservable type $U=u$. In the health insurance literature, this has been interpreted as the moral hazard of type $U=u$ [Einav et al. einav; Kowalski kowalski2016]. To the extent that the measured outcome proxies for the outcomes that matter to the decision-maker, the $\text{MTE}$ may guide the treatment decision.
The $\text{MTE}$ function also serves as a building block for other treatment effects. In particular, a wide array of policy-relevant average treatment effects can be expressed as weighted averages, or linear functionals, of the $\text{MTE}$ function:
for some bounded weighting function $w : [0,1] \to \mathds{R}$ [Heckman and Vytlacil hv2005]. Since unobservables are ordered according to their propensity for treatment via (ref), a simple yet important family of groups $\mathcal{U}$ are the intervals $[u,u']$. The corresponding set of treatment effects are the family of local average treatment effects:
which are obtained from (ref) by the weighting function $w_{u,u'} (s) = \mathds{I} (u \leq s \leq u') / (u'-u)$. As discussed by Heckman and Vytlacil hv2005, definition (ref) generalizes the instrument-specific treatment effect introduced and identified by Imbens and Angrist imbensangrist:
to include effects that are not identified by the instrument at hand, but may nevertheless be of policy interest. Perhaps the most widely studied is the average treatment effect for the population:
Beyond the $\text{LATE}$ family, other treatment effects expressible as (ref) include the average treatment effect on the treated, $\mathcal{U} = \{ U: D = 1 \}$. \footnote{ The fact that $ D $ is measurable with respect to $U$ follows from the selection equation (ref). } All aforementioned effects would be identified via (ref) if the $\text{MTE}$ function were identified.
The $\text{MTE}$ function is not fully identified in the basic model with a discrete source of variation. Yet there is empirical content; for example, certain functionals of the $\text{MTE}$, e.g. (ref), are identified [Imbens and Angrist imbensangrist]. Let $\mathcal{M}^*$ denote the set of candidate $\text{MTE}$ functions that are consistent with the data ((ref)) and model ((ref)). This set is sharp with respect to these assumptions. In what follows I assume that the data-generating model is correctly specified, so that the $\text{MTE}$ function exists and belongs to the set $\mathcal{M}^*$. Given the candidate set $\mathcal{M}^*$, a treatment effect $\text{WTE} (w)$ of form (ref) belongs to the set:
Thus $\text{WTE} (w)$ is (perhaps infinitely) bounded above and below by:
If $\mathcal{M}^*$ is a convex set, then the set of possible $\text{WTE} (w)$ is an interval, characterized by these upper and lower bounds. In this case, the bounds (ref) are sharp for $\text{WTE} (w)$ with respect to the model and data. Deriving the sharp bounds on $\text{MTE}$ and the class of $\text{WTE}$ functionals in the latent index model are the main theoretical contributions of this paper, which I turn to after a preliminary definition and result.
To state and prove the main results, I first define conditional random variables based on a common partition of the population and then compile some related identification results. Namely, let:
denote a random variable that groups (or stratifies, in the language of Frangakis and Rubin frangakisrubin) the population based on the potential treatment decisions for each treatment assignment; the names $\{a,c,n,d\}$ stem from the always-takers, compliers, never-takers, and defiers nomenclature of Angrist, Imbens, and Rubin angristimbensrubin. (ref) rules out defiers, $\{G = d\} = \emptyset$, which are henceforth omitted. For each other group $g$, let $(U_{g}, Y_{0,g}, Y_{1,g}, \Delta_g) \sim ((U,Y_0, Y_1, \Delta) | G = g)$ denote a random vector of eponymous random variables conditional on group membership $g$; these are well-defined for groups where $P (G = g) > 0$. The following preliminary lemma compiles known results about identification within groups in terms of the conditional random variables.
Identification of any group-conditional treatment effect $\Delta_g$ is notably absent from (ref). Instead, define $\Delta_g^-$ to be the group-conditional treatment effect if outcomes $Y_{0,g}$ and $Y_{1,g}$ were countermonotonic: that is, higher instances of $Y_{1,g}$ were paired with lower instances of $Y_{0,g}$. In contrast to the distribution of the true group-conditional treatment effect $\Delta_g$, the distribution of the extremal, countermonotonic treatment effect $\Delta_g^-$ can be recovered from the marginal distributions of $Y_{0,g}$ and $Y_{1,g}$, perhaps most simply through their quantile functions:
The countermonotonic treatment effect plays an important role in the subsequent partial identification results.
This subsection provides the main partial identification results for the marginal treatment effect ($\text{MTE}$) function and the class of weighted average treatment effects. As is to be expected, meaningful partial identification is only possible in the basic latent index model for the interval of compliers, where there is some information about both treated and untreated outcomes. However, the extension in (ref) shows how these identification results can nevertheless yield meaningful and sharp extrapolation beyond compliers when additional assumptions are made.
The main result characterizes the set $\mathcal{M}^*$ of candidate $\text{MTE}$ functions that are consistent with the data and model. \footnote{ Elaborating on the parenthetical condition, the data can only constrain feasible $\text{MTE}$ functions up to sets of measure zero, and occasionally (e.g. (ref)) it will be helpful to think of $\mathcal{M}^*$ instead as a set of equivalence classes of integrable functions. In that case $\mathcal{M}^*$ is also a subset of the normed vector space of integrable functions $L^1 [0,1]$. }
(ref) splits the characterization of $\text{MTE}$ functions into two conditions. The first condition, as is well-known, requires consistency of the $\text{MTE}$ function with the identified $\text{LATE}_{IA}$. Adding the second condition requires further consistency of the $\text{MTE}$ function with the distribution of countermonotonic treatment effects identified on the complier interval.
Intuitively, necessity of the condition follows in two parts. First, by definition of the $\text{MTE}$ as a function of conditional expectations of treatment effects, the true group-conditional distribution of treatment effects $\Delta_g$ is a mean preserving spread \footnote{In the sense of Rothschild and Stiglitz rothschildstiglitz70.} of the $\text{MTE}$ among that group:
The first condition (ref) implies (ref), but the true group-conditional distribution of treatment effects $\Delta_g$ in (ref) is typically unidentified --- even among compliers --- because the model and data only identify (some) marginal distributions of $Y_{d,g}$. However, each true group-conditional treatment effect is stochastically ordered relative to the group-conditional countermonotonic treatment effect, $\Delta_g \succeq_{SSD} \Delta_g^-$. Among compliers, where the countermonotonic treatment effect is identified, combinining with (ref) implies the necessity of (ref).
Conversely, the countermonotonic treatment effect $\Delta_c^-$ is always consistent with the model and data, and any feasible $\text{MTE}$ function on the complier interval can be constructed as if from this extremal distribution, even if that distribution differs from the true distribution $\Delta_c$. Furthermore, there is no restriction (beyond integrability) on the $\text{MTE}$ function outside the complier interval because one of the potential outcomes is unobserved.
The next result identifies candidate $\text{MTE}$ functions that attain the sharp bounds for the class of treatment effects $\text{WTE} (w)$ whenever finite bounds exist. These bounding solutions are characterized by two features. First, they are extremal, in the sense that the distribution of values they take among compliers is equal to the extremal distribution of feasible treatment effects $\Delta_c^-$ identified by (ref). Second, they are monotonic with respect to the weighting function.
The basic intuition of (ref) is as follows. First, the integrand of $\text{WTE} (w)$ defined in (ref) is a supermodular function of $\text{MTE}$ and the weighting function $w$ (it is just the multiplication of the two). Therefore, the value of $\text{WTE} (w)$ is weakly increased (decreased) by rearranging values of $\text{MTE}$ to match with weights in ascending (descending) order. Furthermore, the sharp set of $\text{MTE}$ in (ref) is closed under such rearrangement operations within group $g$ because further information about treatment propensity is unobserved. This yields monotonicity in (ref). Second, the value of $\text{WTE} (w)$ is increasing (decreasing) in mean-preserving spreads of comonotonic (countermonotonic) $\text{MTE}$. Applying this intuition within the group of compliers --- where an extremal distribution of treatment effects is identified --- yields the distribution of the bounding solutions in (ref). Finally, finite bounds exist in the basic model only for weighting functions that are essentially zero outside the interval of compliers. This is why the conditions on extremizing candidate functions in (ref) restrict their scope to compliers.
An important special case of (ref) is the family of counterfactual $\text{LATE}$ parameters. Here one obtains the intuitive result that some parameters which are “almost” the same as $\text{LATE}_{IA}$ are “almost” identified. In particular, such parameters are sharply bounded by the means of the extremal distributions where the highest or lowest quantiles are trimmed off; as explained in the introduction, this trimming intuition is reminiscent of Lee lee in a related model with sample selection. The bounds can be conveniently expressed using the integrated quantile function. \footnote{ In turn, the integrated quantile function of an integrable random variable $X$ can be obtained by cumulatively integrating over its quantile function, which is computationally straightforward. If $X$ is continuous, then the function also has a representation in terms of a quantile-truncated expectation, $I_X (q) = q \mathds{E} [ X | X \leq Q_X (q)]$; e.g. see Jewitt jewitt. }
The result follows from (ref) given the $\text{LATE}$ weighting function $w_{u,u+\delta} (s) = \mathds{I} (u \leq s \leq u+\delta) / \delta$. In particular, this indicator function places equal, positive weight on the interval $[u, u+\delta] \subseteq [ \overline{p}_0, \overline{p}_1]$ and zero weight elsewhere. Therefore, by (ref), the upper (lower) bounds are attained by candidate $\text{MTE}$ functions that have the highest (lowest) values of the complier countermonotonic treatment effect $\Delta_c^-$ arranged on the interval $[u, u+\delta]$, similarly to the trimming intuition of Lee lee. Evaluating yields the bounds in (ref).
Since the value of the $\text{MTE}$ function at a point $u$ can be almost everywhere recovered as the limit of $\text{LATE} (u, u+\delta)$ where $\delta \to 0$, it may appear that (ref) also imposes finite pointwise bounds on the $\text{MTE}$ function at some points. It does not. Moreover, it cannot because a given point is infinitesimal, and the model and data do not impose any constraints on the $\text{MTE}$ function for sets of measure zero. However, the following subsection shows how the model imposes not just pointwise but (among compliers) uniformly sharp bounds on a closely related object, namely the counterfactual average outcome as a function of the latent index treatment threshold. In turn, these pointwise bounds provide a simple, graphical way to summarize the sharp bounds derived so far.
This subsection summarizes the previous bounds in terms of pointwise (a fortiori, uniformly) sharp bounds for a real-valued function on $\mathds{R}$. This yields a simple and new graphical representation of the latent index model's empirical content for the $\text{MTE}$ function and its linear functionals. To that end, define the counterfactual average outcome $\bar{y} (p)$ to be the average outcome that would occur if the threshold in the latent index selection equation (ref) were exogenously set to $p \in [0,1]$: \[ \bar{y} (p) = p \mathds{E} [ Y_1 | U \leq p] + (1-p) \mathds{E} [ Y_0 | U > p]. \] The counterfactual average function also has a representation in terms of the $\text{MTE}$ function:
This well-known representation underpins the Local Instrumental Variable (LIV) procedure for estimation of the $\text{MTE}$ function [Heckman and Vytlacil hv1999, hv2000].
The following result provides pointwise sharp bounds on counterfactual averages for the interval of compliers, where finite bounds exist.
This result follows from (ref) upon substituting $\text{LATE}(\overline{p}_0,p) = \frac{\bar{y} (p) - \bar{y}(\overline{p}_0)}{p - \overline{p}_0}$, $\text{LATE}_{IA} = \frac{\bar{y} (\overline{p}_1) - \bar{y}(\overline{p}_0)}{\overline{p}_1 - \overline{p}_0}$ and simplifying. Furthermore, the propensity scores $\overline{p}_z = \mathds{E} [ D | Z = z]$ and the mean outcomes $\bar{y} (\overline{p}_z) = \mathds{E} [ Y | Z = z]$ are identified, so that (ref) indeed bounds the counterfactual average among compliers.
The bounds of (ref) on the counterfactual average function encode the sharp set and bound results of (ref) through operations of function rearrangement. \footnote{ The notion and language of rearrangement alludes to the close connection between the stochastic convex order (at the heart of (ref)) and the function majorization order (implicit in (ref)). Namely, random variables are ranked in the convex order if and only if their quantile functions are ranked in the majorization order; Furthermore, the quantile function effectively serves as a rearrangement operator. The fundamental majorization result is due to Hardy, Littlewood, and P\'olya hardylittlewoodpolya. A textbook reference (with a focus on the discrete theory) is Marshall et al. marshallolkinarnold. } Say that functions $m, \tilde{m}:[0,1] \to \mathds{R}$ are rearrangements of one another if the distribution of values taken by the functions are the same, $\tilde{m} (U) \stackrel{d}{=} m (U)$. As a special case, define the increasing rearrangement among compliers of an integrable function $m$ by: \[ m^{\uparrow_c} (u) =
\] The transformation ${}^{\uparrow_c}$ rearranges the function's values on the complier interval in increasing order via the quantile function of $m (U_c)$. Let $\mathcal{A}^*$ denote the sharp set of candidate counterfactual average outcome functions; let $a$ denote a generic element and $a'$ its derivative (where it exists). The following result restates the empirical content of (ref) in terms of rearrangements and the counterfactual average function.
One implication of (ref) is that the bounds $\psi_l$ and $\psi_h$ in (ref) are uniformly sharp on the complier interval: there exists a data-generating process consistent with (ref) and (ref) for which the true counterfactual average outcome coincides with the upper (lower) bounding function on the entire complier interval $p \in [\overline{p}_0, \overline{p}_1]$. Namely, the lower (upper) bounds are uniformly attained by a data-generating process with i) perfect countermonotonicity of complier potential outcomes, captured by $\Delta_c^-$, and ii) perfect negative (positive) selection into treatment among compliers, captured by comonotonicity (countermonotonicity) of $\Delta_c^-$ and $U_c$. Thus the bounding functions are themselves segments of feasible average outcome functions. Yet (ref) also shows that satisfying the uniform bounds is not sufficient for an absolutely continuous function to be a candidate counterfactual average function; the bounds must additionally be satisfied by the increasing (in fact, any) rearrangement among compliers.
The information encoded in the uniform bounds is sufficient to recover the sharp bounds on functionals $\text{WTE} (w)$ of (ref). In particular, the derivative of the uniform lower (or upper) bound recovers the distribution of the countermonotonic treatment effect $\Delta_c^-$ via its quantile function: \[ (\psi_l)' (p) = Q_{\Delta_c^-} \left( \frac{p - \overline{p}_0}{\overline{p}_1 - \overline{p}_0} \right) \quad \text{for almost every $p \in [\overline{p}_0, \overline{p}_1]$.} \] Thus $(\psi_l)' (U_c) \stackrel{d}{=} \Delta_c^-$, and so the sharp bounds for $\text{WTE} (w)$ among compliers can be obtained by rearranging the extremal $\text{MTE}$ function $(\psi_l)' (u)$ to be co- and counter-monotonic with the weighting function $w(u)$. Stated in reverse, the increasing rearrangement (among compliers) of candidate $\text{MTE}$ functions that attain the sharp bounds coincides (among compliers $u \in [\overline{p}_0, \overline{p}_1]$) with the segment of the extremal $\text{MTE}$ function $(\psi_l)' (u)$. Analogous logic holds for decreasing rearrangements and the upper uniform bound. Thus the uniform bounds on the counterfactual average summarize the empirical content of (ref) and (ref) on treatment effects.
An appeal is that these bounds can be depicted graphically, as illustrated empirically in (ref). This graphical representation helps to visualize, summarize, and compare the extrapolation power of the model and its special cases; additionally, as will be discussed in remaining (sub)sections, it helps to assess the consistency of additional assumptions.
The area between the uniform bounds provides an intuitive and theoretically grounded summary statistic of the extrapolation power conferred by the model and data. Namely, define the area of uncertainty $R (p,p'): [0,1]^2 \to \mathds{R}_+ \cup \{ \infty \}$ as the total area between the uniform bounds on an interval $[p, p']$: \[ R (p, p') = \int_p^{p'} [ \psi_h (u) - \psi_l (u) ] \, du \] and define the average area of uncertainty for $p' > p$ as $\bar{R} (p,p') = R(p,p') / (p' - p)$. For any integrable random variable $X$, define its Gini mean difference $\Gamma_X$ and, if $\mathds{E} [ X ] \neq 0$, its Gini index $\gamma_X$ by: \footnote{ The IQF representation of the Gini index can be found in Muliere and Scarsini mulierescarsini. Note that if $X$ takes negative values, the Gini index may be greater than one. }
The next result ties the (average) area of uncertainty in the basic model to the normalized and unnormalized Gini coefficients.
In the case of a nonzero treatment effect, the average area of uncertainty among compliers is a product of three features of the data: the first stage $\overline{p}_1 - \overline{p}_0$, the complier treatment effect $\text{LATE}_{IA}$, and the amount of mean-normalized dispersion around the countermonotonic treatment effect, as captured by the Gini index $\gamma_{\Delta_c^-}$.
Despite the apparent simplicity of this expression, care is required in its interpretation. Consider the size of the first stage. In addition to its direct effect on the average area of uncertainty, the size of the first stage also has an indirect effect of affecting dispersion by reducing uncertainty about the relationship between unobservables and potential outcomes. Thus, interpolation improves as the size of the first stage decreases (ignoring any statistical complications), but at the expense of being confined to a smaller interval. In contrast, even though $\text{LATE}_{IA}$ appears in the second expression, it does not have an effect on the average area of uncertainty per se. For example, adding a constant to the distribution of treatment effects would change $\text{LATE}_{IA}$ without any effect on the average area of uncertainty. Intuitively, this is because the term also enters in the denominator of the Gini index. Finally, the main determinant of the average area of uncertainty is the amount of dispersion in the distribution of the countermonotonic treatment effect. This can be as low as zero if potential outcomes are constant, or arbitrarily large for sufficiently dispersed distributions.
The (average) area of uncertainty is infinite among non-compliers in the basic model, which intuitively reflects the fact that the basic model provides no ability to reason counterfactually about treatment effects beyond compliers. Nevertheless, the average area of uncertainty remains well-defined as assumptions are added to the model, and thus also provides a useful family of summary measures for studying and comparing models within the latent index selection framework.
In a recent but already influential contribution, Mogstad et al. mogstad introduce a computational framework for counterfactual inference in a mean-independent latent index selection model. Their approach is based on i) a set of IV-like estimands that summarize the observed data, and ii) a set of marginal outcome functions that can encode various kinds of additional assumptions a researcher might wish to impose. Remarkably, even with the flexible nature of their framework, they show that their bounds can exhaust the empirical content of means. Thus their framework recovers the mean consistency condition (ref) in (ref).
Yet mean consistency per se provides no counterfactual content in the latent index selection model --- in other words, it places no restrictions on treatment effects beyond $\text{LATE}_{IA}$. In contrast, the distribution consistency condition (ref) of (ref) shows that the (fully independent) latent index selection model has some counterfactual content. It imposes bounds on treatment effects among subgroups of compliers; that is, it permits partial interpolation. \footnote{ An application suggested by Kowalski kowalski2016 in the context of the Oregon Health Insurance Experiment (studied in (ref)) is that a policy-maker might consider providing lottery winners the opportunity to receive discounted rather than free insurance coverage. Then interpolated treatment effects from the experiment would be policy-relevant. } Furthermore, full independence of the instrument is a common empirical assumption, which is frequently motivated by (quasi)random assignment. Perhaps for that reason, despite being logically stronger, full independence is also a common assumption in the theoretical literature. \footnote{ E.g. Imbens and Angrist imbensangrist and Vytlacil vytlacil. Heckman and Vytlacil hv1999, hv2005 assume full pairwise independence $(Y_d, U) \perp Z$ for $d = 0,1$. }
Another relative contribution of the present paper is that the sharp bounds are derived explicitly, rather than being characterized implicitly as the optimal values to a pair of infinite-dimensional convex programs akin to (ref). In fact, despite their infinite-dimensional nature, the solutions underlying the sharp bounds can be described relatively simply. Such explicit sharp bounds help the researcher reason about where they might be willing to entertain more (or fewer) assumptions. For example, a particularly stark feature of the sharp bounds in the basic model is the possibility of perfect countermonotonicity in the treatment effect among compliers. In (ref) I show how the results presented so far can be extended to accommodate a common distributional assumption at the other extreme, namely rank similarity or invariance. Thus the results of this paper can also be used as a basis for incorporating and comparing additional assumptions.
Finally, the sharp bounds can be used to test the validity of additional assumptions, such as parameterizations of the $\text{MTE}$ function. A popular and natural candidate is the linear marginal outcome assumption of Brinch et al. brinch, which point identifies the $\text{MTE}$ function with just a binary instrument. In other words, linearity reduces the sharp set $\mathcal{M}^*$ to a singleton, and thus achieves point identification of all weighted average treatment effects $\text{WTE} (w)$. The candidate $\text{MTE}$ function can be integrated to obtain a counterfactual average function, for which the uniformly sharp bounds of (ref) provide a specification test in the latent index selection model. In fact, consistency with the uniform bounds is also sufficient in the case of a linear $\text{MTE}$ candidate function. This follows because such a function is monotone, and therefore equal to either its increasing or decreasing rearrangement among compliers. These observations are collected in the following result.
As will be shown in (ref), similar results can be obtained for the joint consistency of parameterizations with additional assumptions, for example on the joint distribution of potential outcomes.
The present work also relates to a literature beginning with Heckman et al. hsc1997 that studies the distribution of the treatment effect in models of perfect compliance. Most relatedly, Fan and Park fanpark2010 and Firpo and Ridder firporidder2019 study bounds on the distribution of the treatment effect and functionals thereof. Fan and Park fanpark2010 show that the treatment effect satisfies a second-order stochastic dominance conditions relative to the comonotonic and countermonotonic treatment effect:
In contrast to the $\text{MTE}$ function, however, the second-order stochastic relation (ref) does not characterize the set of feasible treatment effects. As shown by counter-example in (ref), the same remains true even when the second-order bounds are combined with the first-order Makarov makarov bounds, also introduced in the treatment effect framework by Fan and Park fanpark2010. Thus, the known bounds on the distribution of treatment effects are not sharp in the full sense, and characterizing the sharp set of the distribution of treatment effects remains an open problem.
It is interesting to contrast this conclusion with the results of the present paper, which include a derivation of the sharp set on marginal treatment effects in the latent index selection model. To simplify the comparison, I consider the special case of perfect compliance $D = Z$ and drop the redundant complier subscript. In that case, the quantile function of the treatment effect $Q_\Delta (\cdot)$ is a feasible $\text{MTE}$ function, which corresponds to perfect (hypothetical) negative selection on the treatment effect. Thus the necessary and sufficient conditions for a feasible $\text{MTE}$ function are necessary conditions for a feasible quantile function. The conditions are not sufficient.
Rather, the $\text{MTE}$ function is akin to a fractional quantile function, and the problem of deriving sharp bounds on functionals of the $\text{MTE}$ is a relaxation of the problem of deriving sharp bounds on functionals of the quantile function. \footnote{ The $\text{MTE}$ relaxation of the quantile problem bears a noteworthy resemblance to Kantorovich's linear programming relaxation of Monge's optimal transport problem. For more detail in the context of economics, see Galichon galichon. } Thus, if the sharp bounds on an $\text{MTE}$ functional are attained by feasible quantile functions, then the bounds are also sharp over the set of feasible quantile functions. By (ref), the relaxed problem has relatively useful features, such as an analytical characterization and convexity of its feasible set. This suggests a potentially fruitful method for deriving bounds on functionals of the quantile function, even without a characterization of the quantile function sharp set. Even when the bounds are not attained by quantile functions, solutions to the relaxed $\text{MTE}$ problem may be insightful, as illustrated by a simple example in (ref).
I now discuss and extend the previous results in the case where (ref) holds conditional on additional observed covariates, denoted by $X$. Covariate information is useful for at least two reasons. The first reason is that it permits subgroup analysis and comparison. Specifically, the latent index rule (ref) can be generalized to accommodate covariates:
where $U \sim U [0,1]$, $U \perp X$, and $\overline{p}_{z,x} = \mathds{E} [ D | Z=z, X=x]$. Then analysis in terms of a covariate-conditional marginal treatment effect function $\text{MTE} (u,x) = \mathds{E} [ \Delta | U = u, X=x]$ proceeds as previously. The second reason is that aggregation over covariates can be used to improve unconditional inference. This will be the focus of the remaining discussion.
In order to endow the improved bounds with policy relevance, consider a family of alternative policy interventions $\mathcal{W} \subseteq \mathds{R}$ satisfying the following assumptions.
Following a line of research beginning with Heckman and Vytlacil hv2001, hv2005, policy invariance assumes that the alternative policies affect only the propensity for treatment, and not the potential outcomes, unobservables, or covariates. \footnote{ Policy invariance in (ref) is the same as in Carneiro et al. carneiro2010, carneiro2011. It can be weakened to an assumption on the distribution of these random variables across policies $w$; for more detail, see Heckman and Vytlacil hv2005. } Policy monotonicity and instrument inclusion impose structure on how alternative policies affect treatment relative to one another and relative to the observed experiment; in this respect, the substantively novel structure relative to the existing literature is the monotonicity across covariates. \footnote{ Conditional on covariates, policy monotonicity and instrument inclusion recover the essence of the stochastic policy definition of the aforementioned papers [Heckman and Vytlacil hv2001,hv2005, Carneiro et al. carneiro2010, carneiro2011], in which each policy $*$ is associated with a stochastic threshold $P^*$ such that $D^* = \mathds{1} [ U \leq P^* ]$. The reason is that the stochastic policies are assumed to be exogenous, $(Y_0, Y_1, U) \perp P^*$, and so the expected outcome $\mathds{E} [ Y^* ]$ is equivalent to a mixture over uniform policy cutoffs in the population: \[ \mathds{E} [ Y^* ] = \mathds{E}_{P^*} [ P^* \mathds{E} [ Y_1 | U \leq P^* ] + (1-P^*) \mathds{E} [Y_0 | U > P^* ]] \] The conceptual distinction between the approaches is whether a policy is associated with a shift of cutoffs relative to the observed instrument $Z$, or with a (mixture of) uniform cutoff(s) in the population relative to the counterfactual instrument realization $z$; the exogeneity of the instrument and policy shifts under consideration makes the two computationally equivalent. } However, covariate monotonicity appears reasonable for the typical policies that also satisfy policy invariance, such as prices or subsidies for treatment takeup in the absence of general equilibrium effects. \footnote{ The direction of monotonicity across covariates can be inferred from the data. In the case where this direction varies significantly with covariates, a similar analysis could be undertaken by splitting the space of covariate values into subsets of positive and negative response to the policies. See, e.g., Semenova semenova2021 for an approach to flexibly accommodating such covariate-conditional monotonicity in the related sample selection framework of Lee lee. } Finally, the policy cutoff $W$ satisfies $\tilde{D}_w = \mathds{1} [ W \leq w]$ for all $w \in \mathcal{W}$ by definition; continuity of the cutoff is a technically convenient assumption that allows the cutoff to be redefined in terms of its unconditional quantile $W := F_W (W)$, so that $W \sim U[0,1]$ and $\overline{w}_z = \mathds{E} [ \overline{p}_{z,X}]$ for $z \in \{0,1\}$. This normalization is adopted henceforth.
Next, define the marginal policy relevant treatment effect function:
While this definition differs somewhat from the preceding literature, \footnote{ Namely, Carneiro et al. carneiro2010, carneiro2011 consider a sequence of stochastic exogenous policies $\{ P_w^* \}_{w \in \mathcal{W}}$ and define $ \text{MPRTE} (w) = \lim_{w \to 0} \, ( \mathds{E} [ Y_w^*] - \mathds{E} [Y_0^*]) / (\mathds{E} [ D_w^*] - \mathds{E} [ D_0^*])$, where $P_0^* = \overline{p}_Z$. The difference in definitions is consistent with the distinction discussed in \autoref*{fn:policies}. } it yields a close parallel between the $\text{MTE}$ and $\text{MPRTE}$ functions. Of course, the connection between marginal treatment effects and policy relevance is a primary motivation for the study of the $\text{MTE}$ function and its functionals in the first place. Yet, the set of feasible $\text{MPRTE}$ functions is additionally constrained relative to the feasible characterization of (ref) because the extremal treatment effects underlying the bounds can be “at most” perfectly countermonotonic conditional on covariates.
Elaborating, let $ \Delta_{g,x}^-$ denote a random variable distributed identically to the treatment effect restricted to and perfectly countermonotonic within the stratum $(g,x)$. \footnote{ As in (ref), such a random variable is most easily defined in terms of its quantile function. Letting $Y_{d,g,x} \sim (Y_d | G=g, X=x)$, \[ \Delta_{g,x}^- \stackrel{d}{=} Q_{Y_{1,g,x}} (V) - Q_{Y_{0,g,x}} (1-V) \quad \text{where $V \sim U[0,1]$.} \]
} Then define $\Delta_{g,\bar{x}}^-$, the average group-conditional countermonotonic treatment effect across $X$, by the distribution function:
Let $W_g$ denote the random variable of policy cutoffs constrained to $G=g$. It follows analogously to the reasoning of (ref) that the set $\mathcal{R}$ of feasible $\text{MPRTE}$ functions is constrained to the set of integrable functions $r: [0,1] \to \mathds{R}$ satisfying the stochastic relations: \[ \mathds{E} [ r (W_c) ] = \text{LATE}_{IA} \quad \text{and} \quad r (W_c) \succeq_{SSD} \Delta_{c,\bar{x}}^- \] Furthermore, since the assumptions on $\nu$ in (ref) are agnostic about relative selection across covariates, this conditional countermonotonicity is the only additional restriction imposed. Of course, further parametric restrictions on $\nu$ would tighten bounds. \footnote{ For example, one might assume additive or proportional shifts in treatment takeup for policies across covariates. This is similar in spirit to the examples studied in Carneiro et al. carneiro2010, carneiro2011, except that the restrictions are applied across (rather than conditional on) covariates. } Similarly, as explored in the next extension, bounds can be tightened through further assumptions on the joint distribution of potential outcomes.
The next extension illustrates how the preceding framework and results can be used to (sharply) accommodate additional assumptions on the joint distribution of outcomes. Such assumptions address a stark feature underlying the sharp bounds in the basic model: among compliers where the distribution of treated and untreated outcomes is observed, treatment effects could be perfectly countermonotonic, and among non-compliers, treatment effects are not restricted at all. Equally starkly, suppose that the joint distribution of potential outcomes --- and thus the distribution of treatment effects --- were known. Such knowledge could significantly improve the quality of inference over counterfactual weighted treatment effects $\text{WTE} (w)$. Yet this still might not point identify all such effects because of the remaining uncertainty about the relationship between treatment effects and the propensity for selection into treatment.
Rank similarity or invariance is a strong but often reasonable assumption on the joint distribution of potential outcomes that clearly illustrates these improvements and limitations. The identifying power of both assumptions is studied previously by Chernozhukov and Hansen chernozhukovhansen, who focus on identification of the population potential outcome distributions in a generalized setting with possibly non-binary treatments and a non-separable selection rule. Those weaker assumptions suffice for identifying population quantiles, but not for identifying (let alone defining) policy-relevant parameters such as those derived from the $\text{MTE}$ function. This paper's contribution is to sharply incorporate the rank similarity and invariance assumptions into the latent index selection framework, which allows i) partial identification of weighted treatment effects $\text{WTE} (w)$ beyond the average treatment effect, and ii) the relaxation of a previous continuity assumption that is violated in the empirical application of (ref). \footnote{ In other related work, Vuong and Xu vuongxu identify individual treatment effects (and thus also the average treatment effect) under rank invariance and the monotonicity assumption of Imbens and Angrist imbensangrist. The key to their main result on point identification is the existence of a deterministic mapping between potential outcomes, which relies on continuity of potential outcome distributions. The existence of mass points explains the difference in our results, including why the $\text{ATE}$ in my application ((ref)) is not point-identified. }
Rank similarity is theoretically weaker than rank invariance because it allows random slippage in ranks conditional on the selection unobservable $U$. However, the empirical content of the two assumptions is identical because there always exist extremal data-generating processes under rank similarity that are rank invariant. Specifically, define (on the original probability space) an “as if” comonotonic treatment effect by: \footnote{ Thus defined, the comonotonic treatment effect is closely related to the quantile treatment effect (QTE) $Q_{Y_1} (t) - Q_{Y_0} (t)$, $t \in [0,1]$, whose early formulations date back to Lehmann lehmann and Doksum doksum. Importantly for the derivations that follow, however, the comonotonic treatment effect is defined as a random variable on the original probability space (also contrasting with the previously defined countermonotonic treatment effect). } \[ \Delta^+ = Q_{Y_1} (V_0) - Q_{Y_0} (V_0) \] This comonotonic treatment effect is consistent with rank invariance and stochastically bounds the feasible $\text{MTE}$ functions under rank similarity, analogously to the countermonotonic treatment effect in (ref). To formally state and prove this result, extend the notation for group $g$-conditional random vectors to $V_d$ and $\Delta^+$. Then:
(ref) is reminiscent of (ref) but does not seek to characterize the sharp set of candidate $\text{MTE}$ functions under the additional rank similarity assumption. Indeed, if the distribution of each $\Delta_g^+$ is identified from the data, then such a characterization follows from arguments analogous to the proof of (ref) and the fact that rank similarity does not impose any further joint restrictions on unobservables and treatment effects within or across groups $g$.
Now consider identification of the distributions of group-conditional comonotonic treatment effects $\Delta_g^+$. Given the same data ((ref)), Chernozhukov and Hansen chernozhukovhansen have derived moment conditions that can point-identify the distributions of potential outcomes, and thus the distribution of $\Delta^+$, under rank similarity and an additional assumption that the potential outcomes are continuous. To move beyond compliers in the latent index selection model, \footnote{ The comonotonic effect among compliers $\Delta_c^+$ is identified even in the basic model ((ref)) from knowledge of the complier quantile functions. See Abadie et al. abadieangristimbens for a study of the Local QTE among compliers with an emphasis on incorporating covariates. } I introduce an alternative identification strategy that leverages the added structure of the selection rule (ref) to accommodate discrete responses and mass points. For example, this is useful for the empirical application of (ref) because most experiment participants never visit the emergency room and thus receive an outcome of zero. The alternative identification strategy, which inherently relies on the additional structure of the latent index selection model, \footnote{ Or the empirically equivalent monotonicity model of Imbens and Angrist imbensangrist. } is to extrapolate from compliers.
Extrapolation from compliers consists of three steps. First, knowledge of both potential outcome distributions among compliers imposes restrictions between treated and untreated quantile functions for the population, and thus for any other subpopulation. These restrictions can be summarized in terms of quantile-quantile ($QQ$) plots of potential outcomes, defined for the population as:
Among subgroups $g$, define an analogous plot $QQ_g$ by replacing population quantile functions with group-conditional quantile functions. The next result establishes the restriction imposed by the complier $QQ_c$ plot on the population under rank similarity.
To simplify the remaining exposition, I also make an additional support assumption to ensure that the population $QQ$ plot is completely identified from compliers, i.e. $Q Q_c = Q Q $.
The assumption does not presume or imply continuity of the potential outcomes, and it is made primarily for technical convenience. \footnote{ The validity of the support assumption is often an empirical question, which is answered affirmatively when complier potential outcomes take all possible values. For example, in the simplest non-trivial case of binary outcomes, it is easy to verify whether both possible realizations sometimes occur. }
In the second step, the restrictions from the complier $QQ_c$ plot are combined with knowledge of one potential outcome distribution among the always- and never-treated to (partially) recover the other, previously unidentified potential outcome distribution. That is, define the complier-extrapolated quantile bounds:
for all $q \in [0,1]$. Then:
Under rank similarity ((ref)a), the identified $QQ_c$ plot serves as a dictionary for recovering pointwise bounds on the missing quantile function, using the quantile function identified in the basic model. If the potential outcomes are continuously distributed, then the dictionary is one-to-one, and all group-conditional potential outcome quantile functions are point-identified.
Finally, in the third step, bounds on the group-conditional comonotonic treatment effect distributions are obtained under rank invariance, i.e. by matching quantiles. Fix $V \sim U[0,1]$ and consider random variables with the following distributions:
In each group, at least one of the quantile functions is identified in the basic latent index selection model, and uniform bounds on the other quantile function are obtained from (ref). The following first-order dominance relation is immediate by (ref) of (ref).
The bounds of (ref) are useful for several purposes. First, they undergird an extension of the sharp bounds on $\text{WTE} (w)$ of (ref) to the rank similar latent index selection model.
The bounds on functionals with nonnegative weights are attained at the extremal identified distributions of comonotonic treatment effects. An important feature of (ref) relative to (ref) is that it provides finite bounds on all finite functionals $\text{WTE} (w)$ with nonnegative weights, and not just those with nonzero weights among compliers. This arises because rank similarity suffices for true extrapolation from the compliers whose potential outcomes are collectively identified in the basic model.
The bounds of (ref) can also be combined with (ref) to obtain uniformly sharp bounds on counterfactual averages under rank similarity.
In turn, the uniformly sharp bounds impose a necessary condition on the counterfactual average function. \footnote{ It is worth noting that the uniform sharpness of the bounds of (ref) relies on the full support condition of (ref). In its absence, the bounding quantile definitions (ref) and (ref) must be generalized, and then the extremal comonotonic effects underlying the lower or upper bound, e.g. $\overline{\Delta}_a^+$ and $\underline{\Delta}_n^+$, may no longer be jointly attainable by a single rank similar data-generating process. Nevertheless, pointwise sharpness suffices for the necessary condition of interest. The issue also does not arise in (ref) because the underlying extremal comonotonic effects in that case, e.g. $\overline{\Delta}_a^+$ and $\overline{\Delta}_n^+$ for the upper bound, remain jointly attainable. } This yields a simple test of whether rank similarity and further assumptions within the latent index selection framework are jointly consistent. Thus, for example, I find suggestive evidence that rank similarity is not consistent with the linearity assumption of Brinch et al. brinch in the empirical application of (ref). In other words, the linear extrapolation could not be generated by a rank similar or invariant process, even though such a distribution assumption may be reasonable in the context of health expenditures.
While rank similarity may be inconsistent with ancillary assumptions, it is always consistent with the basic latent index selection model itself. Intuitively, this follows from the logic of extrapolating from compliers in (ref). Any marginal distributions of complier outcomes define a rank invariant joint distribution of complier outcomes through the operation of matching quantiles. Any such joint distribution of complier outcomes can be extended to the always- and never-treated because in each case one outcome is unobserved and therefore unconstrained. If (ref) is violated and the supports of the observed outcome distributions vary across the never/always-treated and compliers, then the joint distribution of population outcomes may simply not be sufficiently identified to allow meaningful extrapolation without further assumptions.
Rank similarity can also be imposed conditional on observed covariates $X$ (along with (ref)). It is useful to distinguish between two formulations.
The first version of conditional rank similarity ((ref)a) imposes the previous unconditional rank similarity ((ref)a) within each subpopulation $X=x$. As in the preceding extension ((ref)), the same analysis then proceeds within subpopulations. Aggregating across subpopulations, define (on the original probability space) the covariate-conditional comonotonic effect by: \[ \Delta_{\bar{x}}^+ = Q_{Y_{1,X}} (\tilde{V}_0) - Q_{Y_{0,X}} (\tilde{V}_0) \] Under (ref), the group- and covariate-conditional comonotonic effect $\Delta_{g,\bar{x}}^+ \sim (\Delta_{\bar{x}}^+ | G=g)$ constrains the set of feasible $\text{MPRTE}$ functions, like the group-conditional comonotonic effect $\Delta_g^+$ constrained the $\text{MTE}$ function in (ref). However, whereas covariate information sharpened bounds in the basic model, weak conditional rank similarity relaxes bounds relative to the unconditional rank similarity assumption because the covariate-conditional comonotonic effect is feasible in the basic model and therefore invoking (ref) implies:
where $g = \emptyset$ is used to denote the unconditional case.
The second version of conditional rank similarity ((ref)b) is theoretically stronger than unconditional rank similarity ((ref)a) because it conditions equality of rank distributions on the relatively finer realizations of $(X,U)$. However, because the treatment effect $\Delta^+$ satisfying unconditional rank invariance ((ref)b) is feasible and extremal in either case, the second version of conditional rank similarity confers no additional identifying power relative to unconditional rank similarity in practice. As the name suggests, the second version of conditional rank similarity is also stronger than the first. Relatedly, strong conditional rank similarity ((ref)b) is testable with information on covariates [Dong and Shen dongshen, Frandsen and Lefgren frandsenlefgren2018]. The present framework suggests two new, albeit related, testable implications.
The first implication (ref) requires that the conditional $QQ_{c,x}$ plots across realizations $x$ must lie on a single totally ordered curve in $\mathds{R}^2$, which is testable with identified covariate-conditional quantiles, e.g. among compliers. This is essentially the main testable implication of Dong and Shen dongshen applied to potential outcomes rather than the underlying ranks. \footnote{ The main testable implication of Dong and Shen dongshen is that the distribution of ranks is equal across treatment status conditional on covariates. This implies the $QQ$ plot relation, and the reverse implication is also true under their assumption that potential outcomes are continuous. An advantage to an outcome test is that the quantile-quantile plots are unique even when the rank variables are not, as is the case when potential outcomes are not continuous. } The second implication (ref) requires that (ref) holds with equality in distribution. This implies equality of the uniform bounds among compliers under weak conditional rank similarity and unconditional rank similarity. Thus (ref) can be visually assessed in the graphical plot suggested in (ref) and implemented empirically in (ref). I now turn to this empirical application.
This section applies the methods and framework of the previous sections using publicly available data from the Oregon Health Insurance Experiment (OHIE) [Finkelstein finkelstein-data]. In 2008, a group of uninsured, low-income adults in Oregon were offered the chance to apply for Medicaid via a random lottery. This experiment has provided a unique opportunity to identify causal effects of health insurance across a variety of outcomes among lottery compliers. At the same time, using the data to inform many questions of policy interest --- such as the potential effects of charging a price for similar coverage, or of proposed expansionary policies like “Medicare for All” --- requires interpolating or extrapolating beyond point-identified local average treatment effects. Thus the setting is appealing for employing the methods proposed in this paper.
Specifically, I contribute to the study of the effects of Medicaid on emergency room (ER) utilization. Emergency rooms are an expensive and, when used for non-emergencies, inefficient form of medical care. Furthermore, the effect of expanded medical insurance coverage on emergency room utilization is theoretically ambiguous: health insurance simultaneously decreases costs of both emergency and non-emergency care, as well as possibly affecting health outcomes directly. In summary, the impact of Medicaid expansion on emergency care is an important empirical policy question.
The effect of Medicaid on emergency room utilization has been previously studied in the randomized control setting of the OHIE by Taubman et al. taubman and Kowalski kowalski2016, kowalski2021. Taubman et al. taubman find positive causal effects of Medicaid enrollment on the number of emergency room visits. These effects are significant (on the order of 40% relative to a control group) across a wide range of visit types, conditions, and subgroups. Kowalski kowalski2016 extends this analysis by assessing the external validity of the $\text{LATE}$ and conducting a variety of policy-relevant extrapolation exercises for three ER utilization outcomes --- whether a participant visited the ER, the number of ER visits, and the number of total ER charges --- under the additional assumptions of marginal outcome monotonicity and linearity. Kowalski kowalski2021 further extends the analysis, with a focus on reconciling the results with those from a previous 2006 health reform in Massachusetts. The following paragraph discusses these previous findings in more detail, in order to compare to my own.
Kowalski kowalski2016 finds that, under marginal outcome monotonicity, the data is consistent with an externally valid LATE for all ER utilization outcomes. Under the stronger linearity assumption of Brinch et al. brinch, the $\text{MTE}$ function is point-identified and decreasing in the propensity for treatment, indicating a higher expected treatment effect for those more likely to select into insurance. For all ER utilization outcomes, $\text{MTE} (u)$ is positive for all always-takers and negative for all or most never-takers. Furthermore, the linear extrapolation predicts that further expanding Medicaid coverage to the never-treated (who were either ineligible or would have chosen not to enroll in the experimental Medicaid expansion) would generate negative local average treatment effects among that subpopulation. This is in contrast to the positive average effect identified among compliers (who became eligible and enrolled in Medicaid under the expansion), and the even larger positive treatment effects predicted among the always-treated (who were already eligible for Medicaid). \footnote{ Kowalski kowalski2021 uses this extrapolation to reconcile the OHIE results with the seemingly contradictory results from a 2006 Massachusetts health reform, which showed that ER utilization decreased or stayed the same (Chen et al. chen, Smulowitz et al. smulowitz, Kolstad and Kowalski kolstadkowalski, Miller miller). Specifically, Kowalski kowalski2021 argues that Massachusetts compliers are similar to Oregon never-takers, for whom the linear latent index model implies negative treatment effects. } Thus the linear extrapolation has significant policy implications. Finally, there is some evidence that the functional form is misspecified in the case of ER charges: linearity implies that marginal treatment outcomes are negative for 45% of the sample, which is impossible because ER charges are nonnegative by nature.
The present analysis builds on the preceding work as follows. \footnote{ The publicly available OHIE data consists of 24,646 observations corresponding to lottery entrants who lived in a postal code where residents almost exclusively used one of the twelve hospitals for which hospital usage data was collected. Following the analysis of Kowalski kowalski2016, I further restrict to the subsample of 19,643 entrants who were the only entrants in their households, in order to maintain internal validity. It should be noted that variables in the public-use data have been censored and truncated in order to limit the identification of participants; however, the analysis of Kowalski kowalski2016 suggests that this should not significantly affect the conclusions reached. Since I study the same ER outcomes for the same subsample as Kowalski kowalski2016, the reader is referred to her empirical analysis for a set of summary statistics and a detailed study of outcomes and observable characteristics across different treatment groups. Finally, to maintain the focus on identification, standard errors are omitted, and the following statements are based on point estimates and not statistical significance. } First, incorporating distributions rather than just means yields sharp interpolation bounds in the basic latent index selection model, such as for the $\text{LATE} (\overline{p}_0, \overline{p}_1 - 0.1 (\overline{p}_1 - \overline{p}_0))$ corresponding to a 10% reduction of compliers (bounds for all treatment effects referenced in the text are provided in (ref)). Such interpolations are policy-relevant; for example, as discussed by Kowalski kowalski2016, a policy maker may consider providing lottery winners the opportunity to receive discounted rather than free insurance coverage, inducing some experimental compliers to decline coverage.
For example, for the number of ER visits, the model implies that this aforementioned LATE with discounted coverage is bounded below by -0.66 and above by 1.11 additional ER visits due to insurance coverage. The width of the bounds reflects the fact that the model places little structure on i) how potential outcomes are jointly distributed among compliers, and ii) which compliers are most likely to decline coverage if it is discounted rather than free. Thus, removing 10% of compliers could have a variety of effects that remain consistent with the basic model. More positively, the interpolations are finite without further assumptions, and they typically improve with covariate information, which constrains feasible extremal distributions. For example, including covariate information tightens the bounds to [-0.53,0.98]. Thus, for example, we can conclude that while discounted coverage leading to a 10% reduction in compliers could reduce ER visits among the corresponding compliers, it would not do so by more than roughly half a visit on average.
In comparison, the existing mean-consistent method of Mogstad et al. mogstad only recovers the sharp bounds for the binary ER visit outcome, because binary outcome distributions are summarized by their means. For the other outcomes, some interpolation from means remains possible because potential outcomes are bounded below, $Y_d \geq 0$; otherwise there exist mean-consistent distributions with arbitrary dispersion, and so no finite interpolation is possible using existing methods without further assumptions. Even with nonnegativity, the mean-consistent bounds of [-1.30,1.60] on the aforementioned $\text{LATE}$ with discounted coverage are notably wider than those that use additional information from identified distributions. For example, they fail to rule out that the discounted coverage will reduce ER visits among the corresponding compliers by at least one on average, even though such inference is possible given the assumptions of the model.
As a second contribution, incorporating distributions allows one to incorporate further distribution assumptions, such as (conditional) rank similarity, into the latent index selection model. In the OHIE context, rank similarity is satisfied if those who have relatively higher ER utilization without insurance also have relatively higher ER utilization with insurance. In other words, receipt of insurance does not systematically change the ordering of utilization in the population. A consistent finding across ER utilization outcomes under (conditional) rank similarity is that extrapolations to the never-treated yield a nonnegative counterfactual average treatment effect, $\text{LATE}(\overline{p}_1, 1)$. \footnote{ Methodologically, the proposed identification method leverages the added structure of the latent index selection model to (partially) identify distributions even though two of the ER outcomes are discrete and all ER outcomes have a mass point at zero. Because the outcomes are thus discontinuous, the previous identification result of Chernozhukov and Hansen chernozhukovhansen does not apply. } This presents a stark contrast to the linear extrapolation, which predicts a negative average treatment effect among the never-treated. For the indicator and number of ER visit outcomes, unconditional rank similarity makes the even stronger prediction that the $\text{MTE}$ function is everywhere positive, because the complier quantile functions are consistently ranked on their domain. Thus partial extrapolations, such as the $\text{LATE} (\overline{p}_1, \overline{p}_1 + 0.1(\overline{p}_1 - \overline{p}_0))$ corresponding to a 10% expansion of compliers, are also bounded below by zero for the two discrete ER outcomes.
For the total ER charges outcome, the fact that complier quantile functions are not ranked implies that the realized treatment effects among some participants are negative; \footnote{ In other words, $P ( Y_1 < Y_0 | G = c) > 0$. For a detailed discussion of bounds on such “inequality probabilities” in the context of health economics, see Mullahy mullahy2018. } thus some counterfactual average effects could also be negative. Yet the extrapolated effect $\text{LATE}(\overline{p}_1,1)$ among the never-treated is even larger than the $\text{LATE}_{IA} = \text{LATE}(\overline{p}_0, \overline{p}_1)$ among compliers but smaller than the extrapolated effect $\text{LATE}(0, \overline{p}_0)$ among the always-treated. This is inconsistent with the weaker $\text{MTE}$ monotonicity assumption. The non-monotonicity occurs because (conditional) rank similarity predicts that experiment participants in the extreme upper tail of the total charge distribution have large positive treatment effects, which account for most of the mean effect, and these participants are more highly represented among the never-treated than among compliers. This finding remains consistent with the adverse selection on health observed by Kowalski kowalski2021 because i) adverse selection in health across treatment propensity is distinct from variation in marginal treatment effects (or moral hazard) across treatment propensity, and ii) the finding is driven by the extreme right tail of expenditures among the never-treated, and not a discrepant ranking of mean untreated outcomes. Therefore, while estimated averages of distributions with such high dispersion are noisy, the estimates still suggest a plausible extrapolation story that differs from the one obtained by functional form assumptions. Summarizing the main novel empirical takeaway, I find that full extrapolations to the never-treated under rank similarity yield a consistently nonnegative treatment effect across a variety of measures of utilization, in contrast to previous findings of a negative treatment effect under a linear extrapolation of the $\text{MTE}$ function.
At the same time, the aim of the preceding comparison is not to argue that rank similarity is necessarily a better assumption than outcome monotonicity or linearity in the OHIE context, but rather to highlight the third contribution: the framework and its incorporation of distribution information allows the researcher to compare the content of various mean and distribution assumptions and assess their joint consistency. As an additional such example, in (ref), the $\text{MTE}$ function obtained under the linearity assumption for total charges is too dispersed to have been generated by a (conditionally) rank similar data-generating process among compliers. Thus again, the linearity and rank similarity assumptions do not appear jointly consistent. More generally, this shows how distribution information alone could in theory reject functional-form assumptions that point-identify functional parameters of interest, as in (ref). Additionally, the uniform bounds in (ref) provide a specification test of unconditional rank similarity, as formalized in (ref). In this case, the visible difference in unconditional and conditional rank similar bounds for the number of ER visits provides some suggestive evidence against unconditional rank similarity for that outcome; however, such evidence is nonexistent for the binary visit outcome and seemingly more limited for the ER total charge outcome. The statistical formalization of these observations is left to future work.
Finally, the preceding insights --- as well as the features of data required for sharply bounding all weighted treatment effects $\text{WTE} (w)$ across the different auxiliary assumptions --- are compactly summarized in the uniform bounds of (ref). Thus the proposed methods provide an overarching framework for the partial identification of treatment effects in the latent index selection model.
This paper characterizes the set of feasible $\text{MTE}$ functions and derives sharp, analytic bounds on weighted average treatment effects in the latent index selection model [Heckman and Vytlacil hv1999, hv2005; Vytlacil vytlacil]. The latent index selection model is a popular instrumental-variable framework for identification, which in many empirical settings defines a larger class of parameters than it identifies. For example, in the empirical setting of the Oregon Health Insurance Experiment, the researcher or policymaker may be interested in the effects of charging a price for expanded insurance coverage, or of expanding coverage to an even larger population. A theme of the paper is that sharp characterizations of the latent index selection model's empirical content fully utilize identified marginal distributions. Furthermore, such distribution information is useful for encoding new auxiliary assumptions into the latent index framework, which may yield quite different predictions than assumptions that use means alone.
There are several interesting possibilities for future work. One is to (sharply) incorporate other auxiliary distribution assumptions into the model. Another is to study whether the analytical expressions for bounds can be used to simplify results about statistical inference. Finally, and most generally, it would be interesting to study how the analytical toolkit and results in this paper could be used to obtain sharp, analytical results in generalized models that relax the binary nature of the treatment decision or the structure of the latent index selection rule.