EconBase
← Back to paper

Supercompliers

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,322 characters · 12 sections · 106 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\pagestyle{plain}

\thispagestyle{empty}

\setcounter{footnote}{0}

titlepage\begin{center} \\ {Supercompliers\footnote[1]{We thank Josh Angrist, Eric Auerbach, David Card, Amy Finkelstein, Toru Kitagawa, David S. Lee, Doug Miller, Jon Roth, Steve Pischke, and participants in the Princeton University Industrial Relations Section Centennial Symposium and the Stata Virtual Symposium for helpful discussions. Fiona Qiu provided excellent research assistance. Disclaimers: Any opinions and conclusions expressed herein are those of the authors and do not reflect the views of the Joint Committee on Taxation or any Member of Congress. Eng performed this work prior to joining the Internal Revenue Service. All views and opinions expressed herein do not represent the Internal Revenue Service.}} \end{center} \begin{singlespace} \begin{center} {Matthew L. Comey, Joint Committee on Taxation}\\ {Amanda R. Eng, Internal Revenue Service}\\ {Pauline Leung, Cornell University}\\ {Zhuan Pei, Cornell University and IZA}\footnote[2]{Comey: \href{[email removed]}{[email removed]}; Eng: \href{[email removed]}{[email removed]}; Leung: \href{[email removed]}{[email removed]}; Pei: \href{[email removed]}{[email removed]}}\\ \end{center} \end{singlespace} \begin{center} {December 2024} \end{center} \begin{abstract} \begin{singlespace} In a binary-treatment instrumental variable framework, we define supercompliers as the subpopulation whose treatment take-up positively responds to eligibility and whose outcome positively responds to take-up. Supercompliers are the only subpopulation to benefit from treatment eligibility and, hence, are important for policy. We provide tools to characterize supercompliers under a set of jointly testable assumptions. Specifically, we require standard assumptions from the local average treatment effect literature plus an outcome monotonicity assumption. Estimation and inference can be conducted with instrumental variable regression. In two job-training experiments, we demonstrate our machinery’s utility, particularly in incorporating social welfare weights into marginal-value-of-public-funds analysis. Key Words: Local Average Treatment Effect; Principal Stratification; Compliers; Supercompliers; Marginal Value of Public Funds JEL Codes: C10; H00; J24. \end{singlespace} \end{abstract}

Introduction

Seminal studies by ImbensandAngrist1994, AngristandImbens1995, and Angristetal1996JASA establish the now well-known local average treatment effect (LATE) interpretation of an estimand from an instrumental variable regression. In the canonical setting with a binary instrument and binary treatment, the LATE is the average treatment effect for individuals who comply with treatment assignment (i.e., compliers). Unsurprisingly, methodological interest in describing compliers followed (ImbensandRubin1997; Abadie2003), and estimation of complier characteristics, popularized by AngristandPischke2009, is now widely implemented in empirical studies.\footnote{Recent examples include Borghansetal2014AEJPo, Dahletal2014QJE, Dobbieetal2018AER, FinkNoto2019, OussStevenson2021, and AganetalforthcomingQJE.} Describing the characteristics of compliers provides a direct answer to the question: who are induced to take up treatment when eligible? The exercise is informative of the external validity of the identified treatment effect. As AngristandPischke2009 note: “if the compliant subpopulation is similar to other populations of interest, the case for extrapolating estimated causal effects to these other populations is stronger."

In this paper, we extend the complier literature and devise tools to provide a direct answer to a different but related question: who benefit from gaining treatment eligibility? This population consists of those whose treatment take-up responds positively to treatment eligibility and whose outcome improves following treatment take-up. In other words, it is the subset of compliers for whom treatment improves the outcome. We term this subpopulation “supercompliers.”

Supercompliers are a key building block of a LATE. Under a set of testable assumptions and given a binary outcome, we show that the supercomplier population share is the intent-to-treat effect. And since a LATE is the ratio of intent-to-treat and first stage effects, it is equivalent to the supercomplier population share as a fraction of the complier population share. Put simply, the prevalence of supercompliers drives the LATE as we know it. Therefore, learning about supercompliers can be even more informative than learning about the broader complier subpopulation. In the case where supercomplier characteristics differ from complier characteristics, it would be more effective to target treatment eligibility to an external population similar to the supercompliers.

As with compliers, supercompliers cannot be directly observed. However, we show that their characteristics are identified with expressions analogous to those of complier characteristics. Furthermore, our identification result gives rise to estimators for supercomplier average and distributional characteristics that can be easily implemented with standard instrumental variable regressions. Additionally, in our exploration of various estimators, we show that the plug-in complier characteristics estimators used in the literature can be equivalently implemented using simple instrumental variable regression. Finally, we provide results for the identification and estimation of characteristics for compliers whose outcomes are unaffected by treatment.

Our key results follow from standard LATE assumptions along with an additional assumption that we call outcome monotonicity. Outcome monotonicity is the outcome counterpart of the treatment monotonicity assumption in the LATE framework. It states that receiving treatment either does not affect the outcome or changes it in a single direction. In other words, treatment does not improve the outcome for some while harming others.

While outcome monotonicity is not universally tenable, it is plausible in many situations. It is reasonable to assume, for example, that a deworming drug does not lower a child's weight (AndrewsandShapiro2021). More broadly, Institutional Review Board approval for experiments is only granted conditional on demonstrated minimal risks of harm to participants, supporting the assumption of nonnegative treatment effects. Other studies have likewise imposed outcome monotonicity restrictions (e.g., Manski1997, ManskiPepper2000, and Lee2009), highlighting their applicability across a range of settings.

In other situations, there is ambiguity as to whether outcome monotonicity holds. Hence, it is important that we be able to assess the validity of our identifying assumptions and understand the consequence of violation. To that end, we first extend the seminal work by Kitagawa2015 to characterize sharp testable implications under the standard LATE assumptions and outcome monotonicity, and we develop an easy-to-implement joint test. Then, to understand the consequence of outcome monotonicity violations, we show that the bias in our estimand is linear in the share of the complier subpopulation harmed by treatment take-up. This result implies that the bias is likely small when few individuals violate the outcome monotonicity assumption.

Much of our econometric discussion centers around a binary outcome. However, we do not view this focus as restrictive. First, binary outcomes such as employment, welfare receipt, and education credential attainment are quite common in economic research. Second, as argued by BalkeandPearl1997JASA who also focus on binary outcomes, we can easily transform a multi-valued or continuous outcome into a binary variable by categorizing its value as “high” or “low.” Finally, our framework is still interpretable even when the outcome is not binary. Regardless of the distribution of the outcome variable, our estimand identifies a weighted average of supercomplier characteristics, where the weight on each individual is proportional to her treatment effect.

We view analysis of supercomplier characteristics as complementary to subsample heterogeneity analysis of a treatment effect. While conventional heterogeneity analysis asks which subgroups benefit more, our supercomplier analysis collects all beneficiaries and profiles them. Characterizing supercompliers also mirrors the increasingly common practice of describing compliers, which researchers have accepted as a natural complement to subsample heterogeneity analysis of treatment take-up.

In addition to describing beneficiaries, the supercomplier framework can also be used to enrich marginal value of public funds (MVPF; HendrenSprungKeyser2020) analyses, which have recently become commonplace in empirical economic and policy studies. An input into the MVPF is the social willingness to pay (WTP) for a policy, which is often quantified by the LATE of the policy on some outcome of interest (e.g., earnings). Although social WTP (and therefore MVPF) should be calculated by taking a weighted sum of treatment effects, with weights reflecting the planner's redistributive preferences, it is not feasible in practice when only one LATE is reported. We show, however, that when social welfare weights depend linearly on observed characteristics, a weighted MVPF can be computed using only information on supercomplier characteristics and the reported LATE. A key advantage of this approach is that access to the microdata may not be necessary for recomputing an MVPF with different social welfare functions.

We illustrate our proposed method using data from two well-known randomized experiments on job training. The first is the National Job Corps Study, where eligible youths were randomly assigned access to an intensive residential education and job training program (Schochetetal2008). The second is the National Job Training Partnership Act (JTPA) Study, where economically disadvantaged adults were randomly assigned access to Title II-A programs under the JTPA (Bloom_etal1997). We find that, for labor market outcomes, the Job Corps supercompliers are more advantaged relative to the study population. In contrast, among the group that experienced the largest impact from JTPA---adult women---supercompliers are less relatively advantaged. For a social planner who favors redistribution and places greater welfare weight on the more disadvantaged, our results suggest that the MVPF is lower for Job Corps and higher for JTPA.

Related theoretical results were independently developed in exemplary research by Yu2024.\footnote{We became aware of this paper in November 2024. The first version of Yu2024 was circulated in November 2022 as a job market paper. We first posted ours on arXiv in December 2022.} Yu2024 focuses on the political persuasion literature and interprets outcome-response type profiling in that context. We discuss a range of statistical issues not considered in Yu2024, such as interpretation of estimands under stratified randomization, connections between various complier estimators in the literature, identification in the presence of a mediator, and estimation of supercomplier characteristics quantiles. The most important and substantive difference is our emphasis on the interpretability of supercomplier characteristics with a non-binary outcome and the nexus to MVPF analysis. This emphasis significantly broadens the applicability of our framework.

Identification and Estimation

Theoretical Framework and Identification

We begin with an extension of the Rubin1974 potential outcomes framework with a binary instrument $Z \in \{0,1\}$, binary treatment $D \in \{0,1\}$, and binary outcome $Y \in \{0,1\}$. Let $D_z$ represent the potential treatment status for an individual when she is assigned the instrument value $z$. Let $Y_{zd}$ represent a potential outcome for $Z=z$ and $D=d$. With these notations, where each individual's $D$ and $Y$ do not depend on the $Z$ and $D$ of other individuals, we implicitly assume the stable unit treatment value assumption (e.g., Cox1958).

Throughout this paper, we also maintain the other standard assumptions for the identification of a local average treatment effect (LATE) by Angristetal1996JASA.

assumption[IV] \leavevmode \begin{enumerate}[label={\arabic*.}, align=left] • Random Assignment: $(Y_{00}, Y_{01}, Y_{10}, Y_{11}, D_0, D_1) \perp \!\!\! \perp Z$ and $0 < \Pr(Z=1) < 1$. • Exclusion: $\Pr(Y_{1d} = Y_{0d}) = 1 \text{ for } d\in\{0,1\}.$ • Treatment Monotonicity: $\Pr(D_1 \geq D_0) = 1.$ • First Stage: $\Pr(D_1 = 1) > \Pr(D_0 = 1).$ \end{enumerate}

With Assumptions (ref).1 and (ref).2 along with binary $D$ and $Y$, we can partition the population into 16 unobserved subpopulations or groups based on potential treatment and potential outcome values. These 16 groups also appear in BalkeandPearl1993 (the working paper version of BalkeandPearl1997JASA) and ChenFlores2015, and we refer to this partition as the “extended principal stratification.” First, we categorize individuals by how treatment $D$ responds to assignment $Z$. This corresponds to the well-known principal strata (FrangakisandRubin2002) in the LATE context, comprising “always takers,” “never takers,” “compliers,” and “defiers.” Second, the exclusion assumption allows us to extend the principal stratification to incorporate how outcome $Y$ responds to treatment $D$ using the same taxonomy. Correspondingly, we define $Y_d \equiv Y_{zd}$ for $z \in \{0,1\}$ to simplify notation. We index each of the 16 groups $G$ by the pair $to$, where $t \in \{a,n,c,f\}$ indexes the four strata based on treatment response and $o \in \{a,n,c,f\}$ indexes the four strata based on outcome response. For each value of $G$, all individuals have the same vector $(D_0,D_1,Y_0,Y_1)$. We fully define the 16 groups in Table (ref). Supercompliers correspond to the subpopulation $G = cc$, where $D_1 > D_0$ and $Y_1 > Y_0$. Assumptions (ref).3 and (ref).4 place restrictions on the shares of these subpopulations. Most importantly, Assumption (ref).3 rules out all treatment defiers, i.e. $G \in \{fa, fn, fc, ff \}$ (for clarity, we use the word “treatment” as a qualifier when referring to the conventionally defined always takers, never takers, compliers, and defiers). Together, the two assumptions also require a nonzero share of treatment compliers. Next, we impose an additional assumption that further restricts subpopulation shares from Table (ref).

assumption[Outcome Monotonicity and Reduced Form] \leavevmode \begin{enumerate}[label={\arabic*.}, align=left] • Outcome Monotonicity: $\Pr(Y_1 \geq Y_0 ) = 1.$ • Reduced Form: $\Pr(G = cc) \equiv \Pr(Y_1 > Y_0, D_1 > D_0)> 0.$ \end{enumerate}

Assumption (ref) is the outcome analog of Assumptions (ref).3 and (ref).4. Assumption (ref).1 rules out the existence of all outcome defiers (i.e., $G \in \{af, nf, cf, ff\}$). Together, Assumption (ref).3 and Assumption (ref).1 rule out 7 of the 16 groups, and the 9 remaining groups are bolded in Table (ref) (in contrast, BalkeandPearl1997JASA do not rule out groups with monotonicity assumptions but obtain partial identification of the average treatment effect instead). While Assumption (ref).4 implies a nonzero population share of treatment compliers and consequently a nonzero first stage, Assumption (ref).2 requires a nonzero population share of supercompliers, which implies a nonzero intent-to-treat effect (or reduced form) as shown in Lemma (ref) below. (Along with Assumptions (ref).1-(ref).3, Assumption (ref).2 implies Assumption (ref).4.)

lemmaUnder Assumptions (ref) and (ref).1, the supercomplier share is identified by: \begin{equation*} \Pr(G=cc)=E[Y|Z=1]-E[Y|Z=0]. \end{equation*}

All proofs are in the Appendix.

Lemma (ref) states that the share of the supercompliers is identified by the reduced form. This corresponds naturally to the well-known result that the share of treatment compliers is identified by the first stage. Since the reduced form is the product of the local average treatment effect and the first stage, the LATE (given a binary outcome) is simply the share of supercompliers as a fraction of the treatment compliers. This result is intuitive: only the outcome compliers have a nonzero treatment effect (their treatment effect is equal to one), while all other admissible treatment compliers have a zero treatment effect. The more supercompliers there are, the higher the LATE.

In fact, the supercomplier group is the only subpopulation under Assumptions (ref) and (ref) whose outcome changes with the assignment $Z$. In other words, they are the only ones who benefit from being assigned to the treatment group. As such, learning about supercompliers should be of great importance to policy makers.

As with compliers, we cannot directly identify members of the supercomplier subpopulation. However, we can identify the distribution of their characteristics using observed data. Let $X$ be a variable determined prior to random assignment (i.e., prior to the realization of $Z$), so it is reasonable to assume that jointly with $(Y_1,Y_0,D_1,D_0)$, $X$ is independent of $Z$ (hereafter we use $X \perp \!\!\! \perp Z$ as a shorthand for this joint independence). For simplicity, we consider the case of a one-dimensional $X$, but many of our results generalize to a covariate vector of any dimension. Let $h$ be a function such that $E[|h(X)|]<\infty$.

propositionUnder Assumptions (ref) and (ref) and provided that $X \perp \!\!\! \perp Z$, the supercomplier average of $h(X)$ is identified by: \begin{equation} E[h(X) | G=cc] = \frac{1}{RF} E[\pi h(X)], \end{equation} where $RF \equiv E[Y | Z=1] - E[Y|Z=0]$, and $\pi \equiv \kappa - (\kappa_0 Y + \kappa_1(1-Y))$ with \begin{align*} \kappa &\equiv 1 - \frac{D(1-Z)}{\Pr(Z=0)} - \frac{(1-D)Z}{\Pr(Z=1)} \\ \kappa_0 &\equiv \frac{(1-D)(1-Z)}{\Pr(Z=0)} - \frac{(1-D)Z}{\Pr(Z=1)} \\ \kappa_1 &\equiv \frac{DZ}{\Pr(Z=1)} - \frac{D(1-Z)}{\Pr(Z=0)}. \end{align*} The supercomplier average can also be identified by a Wald-type estimand: \begin{equation} E[h(X) | G=cc] = \frac{E[h(X)Y | Z = 1] - E[h(X)Y | Z=0] }{ E[Y | Z=1] - E[Y|Z=0] }. \end{equation}

The weights $\kappa$, $\kappa_0$, and $\kappa_1$ are the unconditional counterpart of those defined by Abadie2003 in his study of compliers. Lemma (ref) in Appendix (ref) provides the analog of Theorem 3.1 from Abadie2003 under our Assumption (ref). While Abadie2003 assumes the LATE assumptions hold conditional on $X$, our Lemma (ref) shows that all three weights can still be used to identify the $X$ distribution of the compliers when the assumptions hold unconditionally. Lemma (ref) also shows that $\kappa_0$ ($\kappa_1$) can still be used to identify the $Y_0$ ($Y_1$) distribution of the compliers under our assumptions.

remark\normalfont ({\bf Compliers and Supercompliers}) There is a natural parallel between the identification results for supercomplier characteristics above and those for complier characteristics from the literature. While the $\kappa$ weights from Abadie2003 “find compliers” (AngristandPischke2009), our $\pi$ weights find supercompliers. Heuristically, the $\kappa$ weight starts from the population and subtracts the treatment always-takers and never-takers, while the $\pi$ weight starts from the treatment compliers and subtracts the $G=ca$ and $G=cn$ groups. There is a complier analog to equation ((ref)) as well: KlineWalters2016 and MarbachandHangartner2020 show that the complier characteristics can be identified by replacing $Y$ with $D$ in ((ref)).
remark\normalfont ({\bf Perfect Treatment Compliance}) Our estimands from Lemma (ref) and Proposition (ref) are applicable to the case where $D=Z$, i.e., the case of perfect compliance with treatment assignment.\footnote{This is the setting Kowalski2020 considers, although she adheres to the conventional nomenclature and studies “defiers” where “take-up” is defined using an outcome variable.} Here, the supercomplier share from Lemma (ref) is equal to the population average treatment effect. And the expressions for supercomplier characteristics identification under perfect treatment compliance are isomorphic (i.e., identical up to variable labels) to those for complier characteristics identification under treatment noncompliance.
remark\normalfont ({\bf Identification of the Distributions of Characteristics}) When $h$ is the identity function, Proposition (ref) says that we can identify the average characteristics of the supercompliers. We can also choose $h=1_{[X \leq x]}$ for all $x \in \mathbb{R}$ and identify the entire c.d.f. of $X$ among the supercompliers. More generally, when $X$ is $k$-dimensional, letting $h=1_{[X \leq x]}$ for $x \in \mathbb{R}^k$ identifies the joint distribution of the random vector $X$ among the supercompliers.
remark\normalfont ({\bf Share and Characteristics of Other Groups}) We can also identify the shares and characteristics of the other two groups within the unobserved treatment complier population: $G=ca$ and $G=cn$. See Appendix (ref) for details.
remark\normalfont ({\bf Adding a Mediator}) The tools we have developed can also shed light on how a binary mediator, $M$, operates on different populations (the causal chain is $Z \to D \to M \to Y$). If we extend the exclusion restriction and monotonicity assumption to cover $M$ (that is, $Y_{zdm}$=$Y_m$ and $M_1>M_0$), our supercomplier estimand simply identifies the characteristics of those with $D_1>D_0,M_1>M_0,Y_1>Y_0$ (the “superdupercompliers”). In addition, we can identify the shares and characteristics of those with $D_1>D_0,M_1=M_0=m$ and $D_1>D_0,M_1>M_0,Y_1=Y_0=y$ for $m,y=0,1$. See Appendix (ref) for details.

We can generalize Lemma (ref) and Proposition (ref) to accommodate non-binary $Y$, where supercompliers are still defined as those with $D_1>D_0$ and $Y_1>Y_0$ and referred to with the shorthand $G=cc$. The proposition below shows that regardless of the distribution of $Y$, the reduced-form estimand in Lemma (ref) identifies the supercomplier population share scaled by the average treatment effect among the supercompliers. Similarly, the Wald estimand in Proposition (ref) identifies a weighted average of supercomplier characteristics, where the weight of each individual is proportional to her treatment effect.

propositionUnder Assumptions (ref) and (ref), the reduced form identifies the supercomplier share scaled by the supercomplier average treatment effect: \begin{equation} \Pr(G=cc)E[Y_{1}-Y_{0}|G=cc]=E[Y|Z=1]-E[Y|Z=0]. \end{equation} With the additional assumption that $X \perp \!\!\! \perp Z$, the Wald estimand identifies supercomplier characteristics weighted by treatment effect: \begin{equation} \frac{E[h(X)(Y_{1}-Y_{0})|G=cc]}{E[Y_{1}-Y_{0}|G=cc]}=\frac{E[h(X)Y|Z=1]-E[h(X)Y|Z=0]}{E[Y|Z=1]-E[Y|Z=0]}. \end{equation}

As we demonstrate in Sections (ref) and (ref), Proposition (ref) substantially widens the applicability of our machinery to different types of outcomes and is essential for our MVPF analysis.

More broadly, while our discussion here has focused on randomized experiments, our framework can be applied to quasi-experimental research designs by making appropriate adjustments to the underlying assumptions. In a regression discontinuity design, for instance, we can extend our logic and identify supercompliers at the policy threshold under smoothness conditions. We now turn to estimation and inference.

Estimation and Inference

The Wald-type estimand in ((ref)) and ((ref)) can be implemented via a two-stage least squares (2SLS) regression. For example, when $h(X)$ is the identity function---corresponding to the mean characteristics of supercompliers---we can run this regression in Stata as

equation[equation omitted — 108 chars of source]

where {\tt XY} is a variable defined as the product of $X$ and $Y$.

However, equation ((ref)) does not estimate an exact sample analog of the Abadie-style equation ((ref)) from Proposition (ref). As we show in Appendix (ref), the estimand from equation ((ref)) also has a Wald-type representation similar to ((ref)):

equation[equation omitted — 161 chars of source]

where $\tau \equiv \Pr(Z=1)$ denotes the proportion of units assigned to treatment. Therefore, the sample analog of ((ref)) can also be implemented using 2SLS. But we need to first transform $Y$ by subtracting from it the proportion of observations in the control group, and then use this transformed outcome variable in the 2SLS regression (it turns out that we can ignore the sampling variation in estimating $\tau$ when conducting inference). As discussed in Appendix (ref), neither of the two estimators based on ((ref)) and ((ref)) has an asymptotic variance that dominates the other in all data generating processes (DGPs).

Estimators based on complier analogs of both ((ref)) and ((ref)) have been used in existing studies, for which $Y$ is replaced by $D$. AngristandPischke2009 implement the complier analog of ((ref)). Other empirical studies referenced in the introduction (e.g., FinkNoto2019) use a different estimator, which is equivalent to the complier analog of ((ref)) (see Appendix (ref) for details). However, these studies do not use 2SLS regression to estimate complier characteristics. Instead, implementation involves assembling separately estimated quantities. In addition, these studies either do not report standard errors on estimated complier characteristics or report bootstrapped standard errors. Our results here imply that inference results can be easily obtained for both estimators using existing Stata commands, including those that account for a weak instrument.\footnote{For estimating complier characteristics, a weak instrument has the standard meaning---$Z$ fails to generate sufficient variation in $D$. For supercompliers, we have a weak instrument problem if $Z$ fails to generate sufficient variation in $Y$.}

Recent work by AngristHullWalters2023 follows and extends the arguments by Abadie2002 and uses additional estimators to characterize compliers. In addition to the complier analog of ((ref)), which amounts to using $\kappa_1$ weights, they also consider weighting covariates by $\kappa_0$, equivalent to replacing $Y$ by $1-D$ in ((ref)). AngristHullWalters2023 also propose a “pooled" estimator, which is the average of the $\kappa_1$ and $\kappa_0$ weighted estimators and can be implemented via a stacked regression. While we can easily use the supercomplier analogs of these additional estimators, replacing $Y$ with $1-Y$ seems unnatural when $Y$ is non-binary. Therefore, we simply use the sample analog of the Wald estimand in ((ref)) and ((ref)) to estimate supercomplier characteristics, which is akin to the complier IV estimators used by Alsanetal2024 when the treatment variable is non-binary.

remark\normalfont ({\bf Characteristics Distribution and Quantile Estimation}) As we discussed in Remark (ref), we can identify the entire $X$ distribution among supercompliers. It is straightforward to estimate the c.d.f. of $X$: we can replace the dependent variable in ((ref)) with the product of the indicator function $1_{[X \leq x]}$ and $Y$ and estimate a series of 2SLS regressions by varying $x$. To estimate the supercomplier quantiles of $X$, we can minimize a weighted sum of the check function per BassettandKoenker1982. However, the same challenge encountered by Abadieetal2002 in complier quantile estimation is also present here---since the individual $\pi$ weight may be negative, the sample objective function is usually non-convex and is therefore difficult to minimize. Abadieetal2002 overcome this challenge by using weights conditional on $(D,Y,X)$. Directly applying this strategy to the supercomplier setting does not lead to nonnegative weights, but applying a modified strategy works, in which we use the $\pi$ weights only conditional on $(Y,X)$ or just $X$. See Appendix (ref) for details.
remark\normalfont ({\bf Conditional Independence and Stratified Randomization}) Our identification and estimation results can be naturally generalized to accommodate cases where independence (Assumption (ref).1) holds conditionally on covariate set $W$---we just need to add $W$ into the conditioning set in Lemma (ref) and Proposition (ref). A common situation that calls for conditional independence is stratified randomized experiments, in which researchers typically include stratum fixed effects in treatment effect regressions (BruhnandMcKenzie2009). It is then natural to also include the stratum fixed effects in the IV regression estimating supercomplier characteristics. In Appendix (ref), we follow Blandholetal2022 to show that the resulting population regression coefficient still identifies a non-negatively weighted average of supercomplier characteristics across strata.

The Outcome Monotonicity Assumption

While Assumption (ref) is standard in the RCT literature, Assumption (ref).1 (outcome monotonicity) warrants more discussion. First, we point out that previous studies have maintained similar assumptions. For example, in their influential studies of treatment effect partial identification, outcome monotonicity is what Manski1997 and ManskiPepper2000 refer to as “monotone treatment response” and what Lee2009 refers to simply as “monotonicity.” In their motivating examples, outcome monotonicity is taken to mean that the demand curve is weakly downward sloping (Manski1997), that education does not decrease wages (ManskiPepper2000), or that participating in the Job Corps training program does not lower employment (Lee2009).\footnote{ChenFlores2015 extend Lee2009 to bound treatment effects under imperfect compliance. Their monotonicity assumption is akin to Jobs Corps take-up not lowering employment for the treatment compliers, while Lee2009 assumes assignment to Jobs Corps does not lower employment. Under treatment monotonicity, these assumptions are equivalent.} Additionally, outcome monotonicity is implicitly assumed in the classic constant parameter endogenous treatment model of Heckman1978. As with Assumption (ref), the practical plausibility of Assumption (ref) depends on the context. For example, it is quite plausible for the relationship between training program participation and subsequent employment or between health insurance coverage and doctor visits to be weakly positive. However, there is more ambiguity concerning the relationship between, say, health insurance coverage and out-of-pocket medical spending. Health insurance may lead to significant savings during emergency room visits, but it may also incentivize healthcare utilization and lead to higher spending.

Given uncertainty in the plausibility of this and the other identifying assumptions, researchers would be prudent to test them. We state such a test below.

Assumption Testing

We build on and extend results from Kitagawa2015 and propose a sharp characterization of Assumptions (ref).1-(ref).3 and (ref).1.\footnote{We exclude Assumption (ref).4 (nonzero first-stage) and Assumption (ref).2 (nonzero reduced-form) from the joint test below for two reasons. First, this exclusion is consistent with Kitagawa2015 who only tests Assumptions (ref).1-(ref).3 and not Assumption (ref).4. Second, testing for nonzero first-stage and reduced-form is straightforward and a must-do in any empirical study.} The resulting test takes the form of a set of inequalities that must jointly hold.\footnote{Note that testing treatment and outcome monotonicity (Assumptions (ref).3 and (ref).1) amounts to testing for the existence of treatment and outcome defiers, respectively. While we do not consider it here, Kowalski2020 proposes a finite sample test of the existence of outcome defiers (Kowalski2020 simply refers to them as defiers) in a perfect compliance framework. Extending her results to accommodate incomplete take-up is an avenue for future research.} Consistent with Kitagawa2015, “sharp" here means that if the inequalities hold, then we can construct a data generating process which satisfies Assumptions (ref) and (ref) and rationalizes the observed data. Formally,

propositionGiven the potential outcomes model described in Section 2.1, (i) under Assumptions (ref).1-(ref).3 and (ref).1, the following inequalities hold \begin{align} \Pr(Y = 0, D=1 | Z = 1)-\Pr(Y = 0, D=1 | Z = 0) &\geq 0 \\ \Pr(Y = 1, D=0 | Z = 0)-\Pr(Y = 1, D=0 | Z = 1) &\geq 0 \\ \Pr(Y = 1 | Z = 1)-\Pr(Y = 1 | Z = 0) &\geq 0 ; \end{align} (ii) if inequalities ((ref))-((ref)) hold, there exists a joint distribution of $(Y_{11},Y_{10},Y_{01},Y_{00},D_1,D_0,Z)$ that satisfies Assumptions (ref).1-(ref).3 and (ref).1 and induces the observed distribution of $(Y, D, Z)$.

Our identification results for the size of the treatment complier subpopulations (Lemma (ref) and Proposition (ref) in the Appendix) reveal an intuitive interpretation of inequalities ((ref))-((ref)). The left-hand-sides identify the population shares of the $cn$, $ca$, and $cc$ groups, respectively, under our assumptions. Therefore, the testable implication of the identifying assumptions is simply that these quantities are weakly positive, which may not be the case if, for example, the disallowed subpopulation shares are positive.

Our test is related to a result presented in Machadoetal2019, who study the identification of the sign of the average treatment effect in a LATE setting. Their Theorem 3.2(ii) establishes a set of inequalities that hold if and only if the standard LATE assumptions are satisfied and the average treatment effect is nonnegative. In our binary outcome setting, the Machadoetal2019 inequalities turn out to be identical to those in Proposition (ref). While an assumption of a nonnegative average treatment effect is implied by and therefore weaker than our outcome monotonicity assumption, data cannot tell the two apart. Indeed, the proof of Proposition (ref) reveals that all the information available for evaluating outcome monotonicity is contained within the reduced form estimate (Inequality (ref)), which translates to a nonnegative average treatment effect.

Testing the inequalities one-by-one is straightforward: Each requires a one-sided test based on a treatment-control comparison. To test all three inequalities jointly, we propose running a “stacked” regression, where each stack corresponds to a single inequality. That is: first, create three copies of the data; second, for each individual $i$, define $Y_{i}^{1} = (1-Y_i)D_i$, $Y_{i}^{2} = Y_i(1-D_i)$, and $Y_{i}^{3} = Y_i$; and third, estimate the regression

equation[equation omitted — 161 chars of source]

where $s$ indexes each stack, $\phi^{s}$ is the stack specific constant, and standard errors are clustered at the individual level. The joint test amounts to testing whether $\theta \equiv \min(\theta_1,\theta_2,\theta_3)$ is nonnegative. Specifically, because the estimators $(\hat{\theta}_1,\hat{\theta}_2,\hat{\theta}_3)$ are asymptotically multivariate normal with a covariance matrix we can consistently estimate, we can simulate the asymptotic distribution of $\hat{\theta}$ under the null. We reject the null hypothesis and conclude violations of the identifying assumptions when $\hat{\theta}$ is less than the 5th percentile in that distribution.

remark\normalfont Inequality ((ref)) corresponds to testing outcome monotonicity (Assumption (ref).1). If we wish to test outcome monotonicity under the LATE Assumptions (Assumptions (ref).1-(ref).3)---as opposed to testing all these assumptions jointly---we can simply rely on inequality ((ref)) alone.
remark\normalfont We can generalize the test in Proposition (ref) by removing the binary restriction on $Y$. In Appendix (ref), we state a sharp joint test of our identifying assumptions, which allows $Y$ to have an arbitrary distribution. We implement this generalized test in our empirical applications where $Y$ represents earnings, an important outcome in MVPF analyses.
remark\normalfont We can incorporate covariates to increase power in the outcome monotonicity test. Specifically, instead of using inequality ((ref)) to test Assumption 2.1, we check whether the inequality holds when further conditioning on covariates. A simple way to implement this covariate-augmented test is to first discretize the covariate space (if necessary) and then jointly test the conditional analog of inequality ((ref)) via a stacked regression, in which each stack corresponds to a value of the covariate vector. In Appendix (ref), we provide a proof of concept by showing that this test indeed rejects outcome monotonicity in a well known example by Bitleretal2006,Bitleretal2017 where Assumption (ref).1 fails.

While the result from the exercise in Appendix (ref) is reassuring, we acknowledge that our testing procedures cannot detect all violations of the identifying assumptions. By following the same construction as in the proof of Kitagawa2015's Proposition 1.1(ii), we can show that for any observed $(Y,D,Z)$ that satisfies inequalities ((ref))-((ref)), there exists an underlying joint distribution $(Y_{11},Y_{10},Y_{01},Y_{00},D_1,D_0,Z)$ that violates some of the identifying assumptions. If we use inequality ((ref)) or its extension conditional on covariates as mentioned in Remark (ref) to only test outcome monotonicity, we will fail to detect a violation if for every covariate value, the $G=cf$ group (or the “compfier" group) share is nonzero but is less than the $G=cc$ (supercomplier) share. Finally, like any other statistical validity test, we may lack the precision to detect a violation in a given sample. For all these reasons, we investigate the consequence of outcome monotonicity violations next.

Relaxing Outcome Monotonicity

Proposition (ref) below shows what the supercomplier characteristics estimand identifies when we relax outcome monotonicity (Assumption (ref).1). It is analogous to Proposition 3 in Angristetal1996JASA on relaxing treatment monotonicity (our Assumption (ref).3).

propositionUnder Assumptions (ref) and (ref).2 and provided that $X \perp \!\!\! \perp Z$, \begin{align*} \frac{E[h(X)Y|Z=1]-E[h(X)Y|Z=0]}{E[Y|Z=1]-E[Y|Z=0]}=&E[h(X)|G=cc]+\&\xi\left\{ E[h(X)|G=cc]-E[h(X)|G=cf]\right\} , \end{align*} where we define \begin{equation*} \xi \equiv \frac{\Pr(G=cf) }{E[Y|Z=1]-E[Y|Z=0]}. \end{equation*}

Proposition (ref) says that when outcome monotonicity does not hold, the estimand for supercomplier characteristics will be biased unless the supercompliers and the compfiers have the same average characteristics. The bias increases linearly with the compfier share, which can be interpreted as the degree of an outcome monotonicity violation. The bias is small when the degree of outcome monotonicity violation is low.

In spite of our discussion here and in Section (ref), we acknowledge that needing outcome monotonicity is a drawback of our supercompliers framework relative to subsample heterogeneity analysis, which increasingly relies on machine learning techniques (e.g, AtheyandImbens2016, WagerandAthey2018; see Smith2022 for a recent review). But our approach also has several advantages. First, it allows imperfect compliance with treatment, from which machine learning methods often abstract away. Second, it avoids any approximation of the conditional expectation of potential outcomes or treatment effects with covariates. Consequently, our result does not depend on the choice of a particular machine learning algorithm or tuning parameters therein (see, for example, Knausetal2021EJ for an extensive exercise comparing techniques across many data generating processes). Finally, our approach is transparent to understand, easy to implement, and cheap to compute with existing statistical software commands.

Relationship to Welfare Analysis

In this section, we discuss how supercomplier characteristics can be used to enrich welfare analyses. Specifically, documenting the differences between supercompliers and compliers informs how social welfare would change if a social planner put different “weights” on individuals of different circumstances. More concretely, we describe how social weights are (and are not) used in the widely adopted framework of HendrenSprungKeyser2020, and elaborate on how to appropriately weight treatment effects in this framework using supercomplier characteristics.

HendrenSprungKeyser2020 advocate for the systematic reporting of the marginal value of public funds, or MVPF, to facilitate comparisons of public policies. The starting point for the MVPF is that the government is interested in the impact of a policy on social welfare, defined as the weighted sum of individual utilities: \[ W=\sum_{i}\eta_{i}U_{i}. \] $U_{i}$ is expressed as a money-metric, and $\eta_{i}$ is the social welfare impact of transferring \$1 to individual $i$.\footnote{Any utility function can be normalized to a money-metric by dividing by the marginal utility of income.} A small policy change, denoted by $dp$, impacts social welfare by:

equation[equation omitted — 116 chars of source]

where $WTP_{i}$ is interpreted as individual $i$'s “willingness to pay” (WTP) for the policy. Equation ((ref)) encapsulates the benefits of the policy change. The cost of the policy equals its impact on the government's budget, denoted by $G$. HendrenSprungKeyser2020 define the MVPF formally as \[ MVPF=\frac{\sum_{i}WTP_{i}}{G}, \]or, the ratio of the (unweighted) societal willingness to pay to the government net costs. Note that while the benefits of a policy change should incorporate social weights, the formal definition of the MVPF leaves the weights out, presumably because it is difficult to practically implement. HendrenSprungKeyser2020 suggest, therefore, to use MVPFs to compare policies that benefit similar beneficiaries, so as to “reduce the role of social preferences in driving conclusions.” As shown below, we can use supercomplier characteristics to incorporate weights into the MVPF to reflect a social planner's redistributional preferences, so long as the weights are functions of the characteristics reported.

The MVPF of a policy is calculated by plugging in estimates of its effect on the government budget (denominator) and, in many cases, the willingness to pay (numerator).n a typical application, the policy's effect on beneficiaries' (and possibly their children's) earnings are translated into an estimate of changes in tax revenue collected. The denominator may also incorporate spillover effects on government spending, such as changes in healthcare utilization among Medicare and Medicaid beneficiaries, changes in public school enrollment, or changes in individual contact with the criminal justice system. The willingness to pay is the dollar value that is transferred to individuals (before behavioral responses, per the envelope theorem) for policies that include a transfer or tax. However, if a policy affects later-life outcomes (e.g., additional education or improved health), causal estimates of these effects could better capture social benefits and therefore enter the numerator.

The use of causal estimates as “plug-in” components of an MVPF calculation means that it corresponds to the subpopulation for which the causal estimates are identified. In particular, when using estimates from a randomized controlled trial with imperfect compliance as inputs, the resulting MVPF is the value of the policy for the (treatment) complier group.

A social planner may desire to place different weights on individuals' willingness to pay, even within the complier group. Suppose that the WTP is measured by the effect of the policy on an outcome $Y$, so that $WTP_{i}=Y_{1i}-Y_{0i}$. When $Y$ is binary, it is easy to see that the weighted WTP for the complier group is

equation[equation omitted — 122 chars of source]

In other words, the weighted WTP is simply the (unweighted) local average treatment effect on the outcome of interest ($\text{LATE}_Y$), multiplied by the average social welfare weight of the supercompliers. We show in Appendix (ref) that this logic extends to the case where $Y$ is not binary. An important implication of this result is that if the social weights $\eta_{i}$ are a linear function of observable characteristics $X_{i}$, one can calculate the appropriately weighted WTP (and corresponding MVPF) from reported supercomplier characteristics. Linearity need not be restrictive in practice. In our empirical applications below, for example, baseline family income was recorded in predefined bins. By including an indicator variable for each bin (i.e., a fully saturated set of dummies), we effectively allow the welfare weights to vary flexibly across income categories, and the specification remains linear.

While it is possible to simply estimate the weighted willingness to pay directly by estimating the treatment effect on weighted outcomes, the formulation in ((ref)) shows that it can be done without microdata. If supercomplier characteristics are documented along with the LATE, a reader can construct a customized weighted WTP and MVPF themselves, to the extent that the desired social weights are a linear function of those reported characteristics. A major advantage of this approach is that it allows for readers to compute the MVPF using different social weights, which may not necessarily be the same as the weights chosen by the researchers who estimate the relevant treatment effect.

Empirical Applications

To demonstrate the utility our tools for characterizing supercompliers, we consider experimental evaluations of the Job Corps and Job Training Partnership Act programs.

Job Corps

Job Corps is a federally sponsored job training program for disadvantaged youths in the United States. The program offers a suite of services, including academic education, vocational training, and job search assistance. In the 1990s, Department of Labor sponsored the National Job Corps Study, where eligible Job Corps applicants between ages 16 and 24 were randomly assigned access to Job Corps. Researchers have used this experiment to study the effects of Job Corps on various outcomes (JCImpacts2001).

Consistent with Job Corps's mission, there were large impacts on educational attainment during the 48-month follow-up period in the public use data (Schochet2008Data). Members of the treatment group were 22.3 percentage points more likely to obtain a vocational certificate (37.5 percent compared to 15.2 percent; $t$-stat: 27.1; $N$: 11,151). Among students without a high school credential at random assignment, members of the treatment group were 15.0 percentage points more likely to obtain a GED (41.6 vs. 26.6 percent; $t$-stat: 14.3; $N$: 8,579). Job Corps also improved labor market outcomes, though the effects are more moderate. Specifically, the intent-to-treat effects on employment and earnings in the 16th quarter after random assignment---the final quarter of study by JCImpacts2001---were 2.4 percentage points (71.1 vs. 68.7 percent; $t$-stat: 2.59; $N$: 10,872) and \$266, or 11 percent of the control mean, ($t$-stat: 4.76; $N$: 10,872), respectively.

In Table (ref), we present supercomplier and complier characteristics with respect to educational attainment, as well as differences in characteristics between supercompliers and both compliers and the full experimental population (standard errors for the differences are estimated via stacked regressions). In Panel A, we examine characteristics for those attaining a GED, restricting to the sample without high school credentials prior to random assignment. In Panel B, we examine characteristics for those attaining a vocational credential. In both panels, we define (treatment) compliers as those who ultimately enrolled in Job Corps. They made up 73 percent and 72 percent of the respective samples. Panel A shows that supercompliers with respect to GED attainment were more likely to be female and white, and also to have had employment prior to random assignment, all relative to both the full population and compliers. Although not statistically significant at the 5 percent level, supercompliers also appeared to be slightly more advantaged with respect to family income. Similar patterns hold for supercompliers with respect to vocational certificate attainment (Panel B), and here we also find significant differences in arrest history, where supercompliers were less likely to have had a prior arrest.

In Table (ref), we present supercomplier characteristics with respect to a labor market outcome. Because of the relatively weak intent-to-treat effect on employment reported above, we focus on earnings in the 16th quarter after random assignment. The patterns in Table (ref) share many similarities with Table (ref): supercompliers were more likely to be white, older, less likely to have been arrested, and more likely to have had an employment history. But the magnitude differences are much more pronounced for two of these characteristics. Compared to Table (ref), supercompliers were substantially more likely to be white and age 20 or older. While not statistically significant at the 5 percent level, point estimates suggest that supercompliers were also more likely to come from higher-income families. Finally, for all outcomes in Tables (ref) and (ref), our tests based on Propositions (ref) and (ref) cannot reject our identifying assumptions at the 5 percent level, indicating their plausibility.

Overall, our labor market findings combined with our results on educational attainment suggest that participants who benefited most from Job Corps tended to be those who were less marginalized upon application: they were less likely to have had an arrest history, more likely to have been white, more likely to have had an employment history, and tended to be from families with higher income. We return to this point and its normative implications at the end of Section (ref).

Job Training Partnership Act

The National Job Training Partnership Act (JTPA) Study was a randomized controlled trial evaluating job training programs in 16 sites across the US in the late 1980s. The JTPA programs primarily target economically disadvantaged adults and out-of-school youths. Eligible applicants to each site's JTPA program were randomly assigned to a control group or a treatment group, which was offered one of three service types: classroom training, subsidized on-the-job training, or other services. Following Bloom_etal1997, we focus on earnings in the 30-month follow-up period as the main outcome of interest. HendrenSprungKeyser2020 also rely on this outcome for the MVPF calculation of JTPA.

While Bloom_etal1997 estimated program effects for four distinct groups---adult men, adult women, male youths, and female youths---we restrict our analysis to adult women.\footnote{The treatment take-up effects were around 60 percent for all four groups.} Among the four groups, adult women saw the largest impact of the intervention with an intent-to-treat effect of \$1,190, or 10 percent of the control mean ($t$-stat: 3.63; $N$: 6,102). Adult men saw an intent-to-treat effect of \$1,051, or 6 percent of the control mean ($t$-stat: 2.01; $N$: 5,102). The intent-to-treat effects for the youth groups were insignificant at the 5 percent level. Since the intent-to-treat effect is the supercomplier “first stage,” it is only strong for adult women by the conventional rule of thumb of $F$-stat $>10$.\footnote{Recent studies by Leeetal2022 and Angristandkolesar2024 reexamine the reliability of inference procedures for just identified single-instrumental-variable models. We first note that the “first-stage” $F$-stats (square of the $t$-stats of the intent-to-treat effects) for the two panels in Table (ref) are sufficiently strong (above the threshold of 142.6 according to Leeetal2022) to dispel inference concerns. In addition to the statistics reported in Tables (ref) and (ref), we have also examined the Anderson-Rubin (AR) confidence intervals for supercomplier characteristics: the 95 percent AR confidence intervals are 9 to 13 percent (17 to 25 percent) wider for the characteristics reported in Table (ref) (Table (ref)) than their conventional counterparts. However, we can show that the statistical significance (at the 5 percent level) of the differences for white and older youths in Table (ref) is unchanged with AR intervals by applying a conservative test that results from the Bonferroni inequality. Angristandkolesar2024 advocate for screening for the sign of the “first stage” coefficient, and all of our intent-to-treat effects have the expected signs.}

Following the same format as Tables (ref) and (ref), Table (ref) reports the characteristics of the adult women in the study. We find that compliers were similar to the experimental population, but that supercompliers were somewhat different. With the caveat that standard errors are large and that none of the differences are statistically significant at the 5 percent level, we see that supercompliers were less attached to the labor force at baseline: they had lower prior earnings and weeks worked, were more likely to be on Aid to Families with Dependent Children (AFDC, i.e., cash assistance), and had lower family income.\footnote{Note that estimated characteristics shares of supercompliers are not guaranteed to lie between zero and one (in fact, neither are complier characteristics estimates). It is reassuring that none of our point estimates are negative or larger than one in Tables (ref)-(ref).} As with Job Corps, the test based on Proposition (ref) also cannot reject our identifying assumptions in the JTPA data for the earnings outcome.

We now turn to the normative implications of these results. Following the framework from Section (ref), we estimate the JTPA program's socially weighted MVPF. HendrenSprungKeyser2020 estimate that the unweighted MVPF of the JTPA program for adults (pooling men and women) is 1.38. This is calculated by assuming that participants value the program by their post-tax earnings impacts, net of AFDC benefits. Since the earnings impacts for women appear to be driven by less advantaged participants, and to the extent that one places a larger social weight on these participants, the weighted MVPF will be higher.\footnote{While the WTP calculated by HendrenSprungKeyser2020 contains earnings and AFDC impacts for both men and women, we only weight the earnings impacts for women. As mentioned above, the reduced form earnings effect for men is not strong enough to compute supercomplier weights. The same is true of AFDC impacts for both men and women. Furthermore, JTPA impacts on AFDC contribute minimally to the MVPF calculation: without accounting for it, the unweighted MVPF is 1.35 (vs. 1.38).}

To compute the weighted WTP (the numerator of the weighted MVPF), we first estimate the average social weight of the supercompliers. We assume a textbook social welfare function $\Psi(u)=\frac{u^{1-\phi}}{1-\phi}$ Salanie2011, where $\phi$ denotes the extent to which the social planner desires redistribution and $u$ is income. A value of $\phi=0$ means that a dollar transfer has the same impact on social welfare regardless of income, while larger values of $\phi$ imply greater preference for redistribution. For a given $\phi$, the estimation of the average supercomplier social weight proceeds in three steps:

enumerate• Calculate the marginal social welfare of transferring \$1 to someone at the midpoint of each of the five income bins. That is, calculate $u^{-\phi}$ by setting $u$ equal to each midpoint. We truncate the top income bin (greater than \$12K) to (\$12K,\$15K) for the purpose of calculating marginal utilities.\footnote{This truncation artificially lowers the assumed income for individuals in the top bin, thereby inflating their marginal utility, resulting in a higher social weight for this income bin, and biasing against our results.} • Calculate the social weights by normalizing the marginal welfare so that they average to one across the five income bins. For example, when $\phi=0.5$, the weights are 1.83, 1.05, 0.82, 0.69, and 0.61, respectively. • Multiply each weight with the corresponding supercomplier share from Table (ref) and sum the products to arrive the average supercomplier social weight. When $\phi=0.5$, the average weight estimate is 1.29.

After these three steps, we multiply the average supercomplier weight with the reported LATE as per equation ((ref)), and plug the result into the companion program by HendrenSprungKeyser2020 to arrive at the weighted MVPF. We find that when $\phi=0.5$, the weighted MVPF is 1.63, while $\phi=1$ implies a weighted MVPF of 1.97.

As discussed above, this MVPF exercise depends on the normalization of the weights. Had we included additional higher income bins, supercompliers would receive an even greater weight, driving up the MVPF. The bottom line is that, whatever the choice of social welfare weights, it is possible to compute the weighted MVPF using only the reported supercomplier characteristics, provided that the weights are a linear function of those characteristics.

Finally, we note the contrast between the results for the JTPA and Job Corps studies. The earnings supercompliers (among women) in the JTPA study are less advantaged compared to the experimental population while the opposite is true in the Job Corps study. Although we do not have the data to estimate a weighted MVPF for Job Corps that directly corresponds to HendrenSprungKeyser2020, our results based on publicly available survey data indicate that incorporating social welfare weights would decrease the MVPF for Job Corps youths by up to 43 percent depending on the value of $\phi$.\footnote{HendrenSprungKeyser2020 estimate an MVPF of 0.18 for Job Corps based on 20-year follow-up impacts from Schochet2018, which come from tax data that are not publicly available. The impacts on earnings from the tax data are more muted even in the short run, and one explanation by Schochetetal2008 is that survey data includes informal earnings not captured by tax data.} In contrast, the MVPF with social welfare weights increases by up to 43 percent for the JTPA adults as shown above.

Conclusion

In this paper, we develop methods to characterize the supercomplier subpopulation in a canonical instrumental variable framework with a binary instrument and treatment. We show that under the standard LATE assumptions plus outcome monotonicity, we can identify the characteristics of supercompliers. Because the plausibility of our identifying assumptions may depend on context, we develop sharp joint tests of their validity. In the presence of outcome monotonicity violations, we show that the average characteristics we identify are a linear combination of the characteristics of those benefiting from and those harmed by treatment receipt, implying that the bias will be small when few individuals are harmed by treatment receipt. Our identification results lead to natural estimators, which can be easily implemented using instrumental variable regressions via existing Stata commands. Finally, we show that supercomplier characteristics can be used to incorporate social weights in a welfare analysis of the marginal value of public funds.

We illustrate the utility of our tools using data from two job-training experiments, the National Job Corps Study and the National Job Training Partnership Act Study. We find that participants whose labor market outcomes were improved by Job Corps had relatively higher baseline incomes, while the adult women who benefited from JTPA had relatively lower baseline incomes. To the extent that a social planner values redistribution, our results imply a lower MVPF for Job Corps and a higher MVPF for JTPA.

singlespace
table[table omitted — 1,971 chars of source]
table[table omitted — 648 chars of source]
table[table omitted — 827 chars of source]
table[table omitted — 681 chars of source]

\pagenumbering{arabic}

\setcounter{equation}{0}

\setcounter{footnote}{0}

\setcounter{figure}{0} \setcounter{table}{0}