EconBase
← Back to paper

Evaluating Counterfactual Policies Using Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

198,792 characters · 13 sections · 92 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Evaluating Counterfactual Policies Using Instruments

abstractWe study settings in which a researcher has an instrumental variable (IV) and seeks to evaluate the effects of a counterfactual policy that alters treatment assignment, such as a directive encouraging randomly assigned judges to release more defendants. We develop a general and computationally tractable framework for computing sharp bounds on the effects of such policies. Our approach does not require the often tenuous IV monotonicity assumption. Moreover, for an important class of policy exercises, we show that IV monotonicity---while crucial for a causal interpretation of two-stage least squares---does not tighten the bounds on the counterfactual policy impact. We analyze the identifying power of alternative restrictions, including the policy invariance assumption used in the marginal treatment effect literature, and develop a relaxation of this assumption. We illustrate our framework using applications to quasi-random assignment of bail judges in New York City and prosecutors in Massachusetts.

Introduction

The most common approach among practitioners for leveraging an \ac{IV} to tease out causal effects from observational data is to use \ac{TSLS} and interpret the resulting estimand as a \ac{LATE}---a (weighted) average of treatment effects for individuals whose treatment status changes in response to the \ac{IV} ImAn94. This interpretation relies on the well-known \ac{IV} monotonicity condition.

Yet in many applications of \ac{IV} designs, both statistical evidence and institutional details point to failure of IV monotonicity. For example, consider the popular leniency \ac{IV} design, which harnesses as-good-as-random assignment of judges or other decision-makers who differ in their leniency---the propensity to grant treatment---to study the effects of treatments such as incarceration, pretrial detention, or bankruptcy on various outcomes.\footnote{The leniency \ac{IV} design was pioneered by kling06. It has been used to leverage (quasi-)random assignment of a variety of decision-makers, including judges AiDo15,DoSo15,dgy18, patent examiners SaWi19, child welfare investigators doyle2007child, and doctors cgy22. See Table 1 in fll23 for additional references.} IV monotonicity requires that every defendant released by a given judge would also be released by all judges who are more lenient on average. Effectively, the judges need to agree on the ranking of the defendants in terms of their risk; they only disagree on the release cutoff. This is a stringent requirement, as pointed out in the original ImAn94 analysis (Example 2, p. 472). Institutional knowledge suggests that decision-makers may have different rankings owing to heterogeneity in skills cgy22 or preferences mueller-smith15. Moreover, IV monotonicity is frequently rejected by statistical tests fll23, as well as by direct evidence from judicial panels sigstad_monotonicity_2023. To address these concerns, a growing literature has proposed weaker or alternative conditions that still allow for a \ac{LATE}-type interpretation of the \ac{TSLS} estimand dechaisemartin17,strlb17,fll23, mtw21. Yet, the plausibility of these alternative assumptions is debated MoTo24.

Often, however, researchers may not ultimately be interested in the LATE per se, but rather in evaluating counterfactual policies.\footnote{A recent survey of the empirical literature by lss25 concludes that researchers are rarely interested in the LATE per se.} Researchers conducting a leniency IV, for instance, may be interested in policies that nudge the decision-makers to be more (or less) lenient, by, say, imposing release quotas adh22, providing algorithmic recommendations ady23, or imposing a presumption of non-prosecution for low-level offenses adh23. In the limit, universal release programs may treat everyone with certain observable characteristics albright22.

In this paper, we develop a general framework for directly evaluating counterfactual policies that alter treatment assignment. We make the goal of policy evaluation explicit by indexing potential treatments as $D(z, a)$, where $z$ is the value of the instrument (say, the identity of the judge in the leniency IV example), and $a$ is an indicator for whether the policy (say, a release quota) is implemented. We only observe data under the status quo ($a=0$), with the observed treatment given by $D(Z, 0)$, and the observed outcome given by $Y(D(Z,0))$, where $Y(d)$ is the potential outcome under treatment $d$. The parameter of interest is $\theta = E[Y(D(Z,1))]$, the average outcome (say, recidivism) under the counterfactual treatment assignment, $D(Z,1)$. Since $\theta$ is generally not point-identified, our goal is to derive bounds for it, i.e., an identified set of values consistent with the observable data.

Within this framework, we develop three sets of results. First, we derive tractable bounds for $\theta$ without imposing IV monotonicity. Second, we show that in a range of relevant cases, imposing IV monotonicity does not help tighten the identified set. Thus, while some form of monotonicity is essential for a \ac{LATE} interpretation, our results show that IV monotonicity is neither necessary nor sufficient to learn about counterfactuals. Third, we consider other assumptions that do help tighten the bounds, and illustrate their power in two applications.

Specifically, our first main result provides a computationally tractable characterization of the identified set for $\theta$ without imposing IV monotonicity. In related \ac{IV} settings kitagawa21, bai_identifying_2024, the identified set is commonly computed by enumerating all possible “response types”, defined by the values of $D(\cdot, \cdot)$. This proves computationally infeasible in our setting, because the number of response types grows exponentially in the number of judges $K$ if one does not impose \ac{IV} monotonicity. We show that one can bypass response type enumeration and compute the identified set as the solution to a linear program that scales linearly in $K$. The key observation is that it suffices to consider sets of marginals for $(Y(\cdot), D(z, \cdot))$, involving the decision of one judge $z$ at a time. This is because the observable data does not depend on the joint distribution of $D(z, a)$ across judges $z$, and without \ac{IV} monotonicity, this joint distribution is ex ante unrestricted. Our result builds on insights in RiRo14, who derived bounds on the average treatment effect in IV settings with binary outcomes involving $O(K^2)$ restrictions.

Our second set of results gives sufficient conditions under which the identified set for $\theta$ does not depend on whether one imposes IV monotonicity. Specifically, we show that this is the case when either (a) one of the potential outcomes is known, or (b) the policy counterfactual is a “sufficiently strong” encouragement (where the notion of “sufficiently strong” is formalized in (ref) below). Condition (a) is satisfied in many criminal justice settings, where, for example, one cannot commit pretrial misconduct unless one is released. Condition (b) is trivially satisfied for universal release programs, like that studied in albright22, and may also plausibly be satisfied by other counterfactual policies that strongly encourage judges to release more defendants. An important practical takeaway is that when these sufficient conditions are satisfied, the debate over IV monotonicity is somewhat of a red herring: instead of worrying about IV monotonicity, researchers interested in policy counterfactuals should turn attention to evaluating the validity of other assumptions that may in fact tighten the identified set.

Our third set of results considers several such assumptions. We begin by evaluating the policy invariance assumption of HeVy05. While \ac{IV} monotonicity effectively imposes that judges agree on their rankings of defendants under the status quo; policy invariance additionally imposes that this ranking is the same under the counterfactual. We show that policy invariance can indeed be helpful in tightening the identified set in some settings where IV monotonicity alone is not. Of course, since policy invariance is stronger than IV monotonicity, it may often be implausible in practice. We therefore develop a relaxation of policy invariance that may be more plausible yet still informative. In particular, while policy invariance imposes perfect agreement among judges in the ranking of defendants, our relaxation only imposes that judges do not disagree too often. We show how bounds on disagreement rates can be calibrated using settings where multiple decision-makers rule on the same cases, such as sigstad_monotonicity_2023's data on panels of judges. We further show that these disagreement bounds can be tractably incorporated into our linear program for calculating the identified set. We also discuss a variety of other economically-motivated restrictions that can be incorporated to further tighten the identified set. For example, in the pretrial release setting, judges are legally instructed to release defendants unless they are at high risk of committing pretrial misconduct. A natural assumption is then that the defendants released by judges under the status quo are lower-risk than the defendants they would marginally release in response to a policy encouragement.

Two applications illustrate the usefulness of our results. The first studies bail judges in New York City, who decide whether to release defendants awaiting trial, using aggregated data from adh22. We evaluate two counterfactual policies in this context. First, we consider a policy that releases all defendants ($D(Z,1)=1$), allowing us to evaluate what would happen if the universal release policy in Kentucky studied by albright22 were applied in this context. This policy is also relevant for evaluating the “disparate impact” of pretrial release decisions by race, as in adh22, whose disparate impact parameter depends only on the (race-specific) average outcomes under universal release and point-identified summary statistics. We obtain informative bounds for the universal release counterfactual, imposing only \ac{IV} validity (i.e., random IV assignment and exclusion). Our bounds suggest, for example, that the defendants marginally released by this policy will have higher misconduct rates than those released under the status quo, which is intuitive given the judges' instructions to only detain high-risk defendants. Our bounds for the disparate impact parameter are only moderately wider than the range of estimates obtained by different parametric assumptions in adh22, and our confidence interval excludes zero, indicating statistically significant disparate impact by race.

We also use this data to evaluate a quota policy that requires the bottom 90% of judges in terms of leniency to increase their release rates to match the top 10%. For this policy, only imposing IV validity yields trivial bounds that do not restrict the misconduct rates of the marginally-released defendants. This occurs because without any restrictions on agreement rates between judges, the judges below the 90th percentile of leniency may choose to marginally release only IV “never-takers” who were not released by any judge under the status quo, and the data thus contain no information about their misconduct potential. To tighten the bounds, we rule out such perfect discordance between judges above and below the 90th percentile by calibrating a bound on the average disagreement rate between these two groups of judges using sigstad_monotonicity_2023's data on panels of judges who rule on the same cases. This calibration yields informative bounds on the counterfactual.

The second application uses data from adh23, who study the effect of non-prosecution of misdemeanor cases by assistant district attorneys (ADAs) in Massachusetts on subsequent criminal justice involvement of the defendants. ADH23 report TSLS estimates. Since the fll23 test rejects IV monotonicity for several of the courts in their data, ADH23 appeal to fll23's average monotonicity condition for a weighted-LATE interpretation of TSLS\@. ADH23 also write that “if all arraigning ADAs acted like the most lenient ADAs in our sample\ldots [we] would likely see a reduction in criminal justice involvement”. Based on this discussion, we evaluate two counterfactual policies to make ADAs act more similarly to the most lenient ones (defined as the top 10%). First, we consider a simple re-allocation of cases to the most lenient ADAs. The impact of such a policy is point-identified, and our results suggest it would decrease crime, in line with the discussion in ADH23. Second, we consider a quota that requires the bottom 90% of ADAs to match the non-prosecution rate of the top 10%. For this policy, stringent bounds on agreement rates between ADAs are needed to conclude that the policy would decrease crime; the bounds appear unreasonably high given rejection of IV monotonicity under the status quo. We can conclude the policy would decrease crime under more plausible restrictions on disagreement, however, if we also restrict heterogeneity of treatment effects among people on whom the ADAs disagree. Thus, the conclusion that increasing non-prosecution likely leads to a reduction in crime, as argued in ADH23, can be obtained if one is comfortable imposing either a specific policy implementation (such as reallocation) or restrictions on the degree of treatment effect heterogeneity.

Our paper is motivated in part by the observation in previous work that the connection between \ac{LATE} and economic policies of interest is not entirely clear HeVy05,HeVy07i,manski1996jhr,manskiJPEMicro. \Textcite{HeVy99,HeVy05} introduced the \ac{MTE} framework for analyzing the impacts of policies that change treatment assignment. However, the canonical \ac{MTE} framework requires policy invariance. Our framework allows for evaluation of a larger class of counterfactual policies by using the generalized potential treatment function $D(z, a)$ and allowing for a wide range of restrictions on it, including relaxations of policy invariance.\footnote{In a complementary line of work, ura_policy_2025 consider a multi-index generalization of the MTE framework that allows for relaxations of IV monotonicity. Although the framework allows for partial identification of a variety of treatment parameters (such as ATE or ATT), analyzing impacts of general counterfactual policies appears challenging without further assumptions. \Textcite{sigstad_marginal_2024} studies conditions under which standard MTE methods correctly estimate parameters such as ATE or LATE even when \ac{IV} monotonicity is violated.} We build on IcTa00, which considers policy evaluation under the assumption that $D(z,0) =D(z',1)$ for a known pair $(z,z')$, so that the treatment assignment for $z'$ under the counterfactual matches that of $z$ under the status quo. Such agreement would arise under policy invariance if $z$'s release rate under the status quo matches that of $z'$ under the counterfactual, $E[D(z,0)]=E[D(z',1)]$. Our framework nests this perfect agreement assumption as a special case, but also allows for much more general restrictions on the potential treatments $D(z,a)$, including non-zero disagreement bounds.

Our paper also relates to a large literature that has considered bounds on parameters other than LATE in \ac{IV} settings BaPe97,manski_nonparametric_1990, swanson_partial_2018, bai_inference_2025, chen_bounds_2015, RiRo14. These papers have primarily focused on parameters such as the \ac{ATE}, or average effects for principal strata, in contrast to our focus on specific counterfactual policies.

Finally, several papers have provided conditions under which \ac{IV} monotonicity does not help to tighten the identified set for the \ac{ATE} or the marginal distributions of potential outcomes BaPe97,kitagawa21,HeVy01bounds,bai_identifying_2024. By contrast, kamat_identifying_2019 shows that IV monotonicity can matter for partially identifying the copula of the potential outcomes. We complement this literature by providing conditions under which IV monotonicity does not help to tighten the bounds on the impact of a counterfactual policy, which need not correspond to the \ac{ATE}.

The next section sets up the model and defines the identified set for $\theta$, the average outcome under the counterfactual. (ref) derives a tractable characterization of the identified set without imposing IV monotonicity. (ref) gives sufficient conditions under which imposing IV monotonicity does not help tighten the identified set, while (ref) gives examples of restrictions that do. The identifying power of these restrictions is illustrated with two applications in (ref). The appendices collect proofs and additional results.

Setup

Observable data and potential outcomes

We observe data on the triple $(Y, D, Z)$, where $Y$ is an outcome of interest, $D\in \{0,1\}$ is a binary treatment indicator, and $Z\in\mathcal{Z} = \{1, \dotsc, K\}$ is a multivalued instrument that influences the treatment. We wish to use the data to learn about the effect of a counterfactual policy, which we denote by $a=1$. We let $a=0$ denote the status quo observed in the data. To properly define the effect of this policy, we use the potential outcome notation, with $Y(d)$ denoting the potential outcome under treatment status $d\in\{0,1\}$, and $D(z, a)$ denoting the potential treatment under instrument value $z\in\mathcal{Z}$ and policy $a\in\{0,1\}$.\footnote{Similar notation appears in IcTa00.} We assume that the support of the joint distribution of $(Y(0), Y(1))$ is contained in a compact set $\mathcal{Y}^{2} \subseteq \mathbb{R}^{2}$.\footnote{For simplicity, we focus on the case where $Y$ is bounded, which ensures that the bounds on the average policy effect are finite. Our results readily extend to the case with unbounded outcomes if one imposes additional restrictions on the outcome, such as those considered in (ref), that guarantee that the average policy effect is finite. Even if the bounds are infinite, one could modify our framework to obtain bounds on quantiles, rather than the average policy effect.} The observed treatment and outcome are thus given by $D=D(Z,0)$ and $Y=Y(D)=Y(D(Z,0))$, while the potential treatments $D(z,1)$ are not observed. The policy effect relative to the status quo is given by $\theta-E[Y]$, where $\theta:=E[Y(D(Z,1))]$ is the average counterfactual outcome under the new policy.\footnote{For ease of exposition, we assume that the marginal distribution of $Z$ is invariant to the policy---e.g., imposing a quota on release rates does not affect the assignment of judges. It is straightforward to accommodate known changes in the marginal distribution of $Z$ by weighting the average counterfactual outcome by the density of the counterfactual $Z$ distribution relative to its status quo distribution.} Since identification of $E[Y]$ is trivial, we focus on $\theta$ as the parameter of interest. We further note that in many settings, it may be reasonable to impose that the counterfactual policy has a monotone effect on the treatment in the sense that $D(z,1) \geq D(z,0)$ (which we formalize in (ref) below), in which case:

equation[equation omitted — 120 chars of source]

The left-hand side is the treatment effect for the policy compliers who are induced to adopt the treatment by the counterfactual policy (e.g., the marginally-released defendants). Thus, when the “first stage” impact of the counterfactual policy is identified, the average effect for policy compliers is also a simple point-identified transformation of $\theta$. We follow the usual convention in identification analysis of abstracting from observable covariates; (ref) details how we incorporate covariates in our empirical illustrations.

To fix ideas, as a running example, consider a judge \ac{IV} design, where $D$ is an indicator for release of a defendant, $Z$ is the identity of a judge who is randomly assigned to the case, and the policy $a=1$ is an encouragement or a quota to release more defendants. $Y$ may correspond to various outcomes of interest. In many applications of leniency IV designs, the outcome is binary, such as an indicator for pretrial misconduct adh22, recidivism or criminal complaints adh23, mortality norris_effect_2024, employment dgy18, earning above the poverty rate kling06, or high-school graduation AiDo15. However, many of our results also allow for continuous outcomes, such as future academic performance of the defendant's child npw21 or average earnings kling06.

Since the potential outcomes are indexed by the treatment only, our notation implicitly embeds two exclusion restrictions: both the instrument ($z$) and the counterfactual policy ($a$) affect the outcome only by impacting the treatment. In the judge IV running example, the first exclusion restriction amounts to assuming that the judges affect the defendants' outcomes only through their release decisions. This would be violated if judges can make other decisions, such as set the terms of probation or fees owed, that may directly affect outcomes mueller-smith15. The exclusion restriction with respect to the counterfactual policy likewise imposes that, say, a quota on how many defendants are released affects defendants' outcomes only by changing which defendants are released. This could be violated if releasing more defendants reduces deterrence, thereby increasing crime directly; more generally, the exclusion restriction rules out general equilibrium effects of the policy.

Additionally, we assume throughout that the instrument is as-good-as-randomly assigned. In the judge \ac{IV} running example, this holds if judges are as-good-as-randomly assigned to defendants, which is the case in many of the empirical papers cited above, at least once we condition on court-by-time fixed effects. Together with the exclusion restrictions described above, random assignment implies that the instrument is valid in that it is independent of the potential treatments and potential outcomes,

equation[equation omitted — 142 chars of source]

The identification power of the instrument comes from its influence on the treatment.

The identified set for the average counterfactual outcome

To define the identified set for $\theta$, let $P$ denote the distribution of the observable data $(Y, D, Z)=(Y(D(Z,0)), D(Z,0), Z)$, and let $P^{*}$ denote the distribution of the model primitives, that is, the joint distribution of potential outcomes, potential treatments, and the instrument, $(Y(\cdot), D(\cdot, \cdot), Z)$. We say that $P^*$ generates $P$ if the implied distribution of $(Y, D, Z)$ under $P^*$ matches the observed data---i.e., if under $P^*$, the distribution of the triple $(Y(D(Z,0)), D(Z,0), Z)$ is $P$.

We encode any ex ante restrictions on the model primitives by restricting the family $\mathcal{P}^{*}$ of possible distributions for $P^{*}$. At minimum, since we maintain IV validity throughout, we must have that $\mathcal{P}^* \subseteq \mathcal{P}^{*}_{\textnormal{valid}}$, where $\mathcal{P}^{*}_{\textnormal{valid}}$ is the set of distributions satisfying the IV validity condition (ref). However, the family $\mathcal{P}^{*}$ may be smaller if the researcher imposes other restrictions on the model that we discuss in more detail below, such as IV monotonicity. Given $\mathcal{P}^{*}$, the identified set for the model primitives is given by the set of distributions in $\mathcal{P}^{*}$ that could have generated the observed data:

equation*[equation* omitted — 114 chars of source]

The identified set for the average counterfactual outcome, $\theta = E[Y(D(Z,1))]$, corresponds to the set of counterfactual means consistent with these primitives:

equation*[equation* omitted — 124 chars of source]

where we make explicit that the expectation is taken with respect to the distribution $P^{*}$. We suppress this dependence whenever it doesn't cause confusion.

For our analysis, it will be useful to distinguish between assumptions that restrict the joint distribution of decisions across decision-makers---i.e., restrict the dependence between $D(z, \cdot)$ and $D(z', \cdot)$ for $z \neq z'$---versus assumptions that only restrict the set of marginals $\{(Y(\cdot), D(z,0), D(z,1))\}_z$ involving one $z$ at a time. We refer to the former as cross-judge restrictions, and the latter as marginal restrictions, as formalized by the next assumption.

assumption[Marginal restrictions] The marginal distributions $\{(Y(0), Y(1), D(z, 0), \allowbreak D(z, 1)) \}_{z\in\mathcal{Z}}$ are contained in some collection of distributions $\mathcal{R}^{*}$, specified by the researcher.

Next, we list examples of both marginal and cross-judge restrictions on the potential treatments that we will consider in our analysis. (ref) then explore the identifying power of these restrictions (and variants thereof).

Examples of restrictions on potential treatments

Knowledge of policy implementation details will typically allow us to impose natural restrictions on the marginal distribution of $(D(z,1), D(z,0))$ for each $z$, which can be encoded by an appropriate choice of $\mathcal{R}^*$ in (ref). In the context of our running example, for many counterfactual policies, the marginal release rates will either be known or consistently estimable, allowing us to impose that $E[D(z,1)] = \alpha_{1, z}$ for all $z$. For instance, under a quota policy that requires judges to release at least a fraction $q$ of the defendants, we may set $\alpha_{1, z}= \max\{E[D\mid Z=z], q \}$ for each $z$, so that each judge who is currently below the quota increases their release rate to match it.

For quota policies or other policies that take the form of an encouragement, it is natural to also impose the following marginal restriction:

assumption[Policy monotonicity] $D(z,1)\geq D(z,0)$ for all $z$.

If a policy is a directive to treat everyone, such as a universal release program, then policy monotonicity holds trivially since $D(z,1)=1$ for all $z$. In the context of a quota policy, (ref) simply imposes that any defendant released without the quota in place would also be released when the quota is in place, which seems quite reasonable. Likewise, an algorithm that flags low-risk defendants as candidates for release may be expected to have a monotonic effect on release rates.\footnote{In some contexts, the algorithm may recommend release for some defendants, and recommend detention for others. In this case, we'd expect the algorithm to weakly increase release rates among those recommended to be released, and weakly decrease them among those recommended for detention. If the algorithmic recommendation is a function of observables, one could impose the appropriately signed version of monotonicity in each subpopulation.} We expect that for many counterfactual policies of interest policy monotonicity will be reasonable, and we will therefore maintain this assumption for many of our results.

A different type of monotonicity restriction that is commonly imposed is the ImAn94 IV monotonicity assumption:

assumption[IV monotonicity] For any pair $z, z'$, either $D(z,0) \geq D(z',0)$ (a.s.) or $D(z',0) \geq D(z,0)$ (a.s.).

This assumption allows one to interpret the \ac{TSLS} estimand as a \acf{LATE}, a weighted average of treatment effects for individuals whose treatment depends on the instrument under the status quo, i.e., those for whom $D(z,0)$ varies with $z$. In contrast to policy monotonicity, IV monotonicity involves restrictions on the joint behavior of judges, so it is a cross-judge rather than a marginal restriction. \Textcite{vytlacil02} shows that (ref) is equivalent to the existence of a latent index $U$ (not depending on $z$) and thresholds $\alpha_{z}$ such that

equation[equation omitted — 96 chars of source]

Intuitively, this representation says that all judges have the same ranking of defendants under the status quo ($U$), but potentially disagree on the threshold to use to determine whether a defendant is released ($\alpha_{z}$). As described in the introduction, (ref) may often be questionable in empirical contexts---judges may differ in their rankings of defendants owing to idiosyncratic perceptions of risk, different preferences over crime types, or differences in skill---and in fact is often rejected using statistical tests fll23, as well as direct measurement when we observe multiple judges making a decision about the same case sigstad_monotonicity_2023.

In addition to IV monotonicity, which imposes a common ranking $U$ under the status quo, we will also consider the even stronger cross-judge assumption that this common ranking does not change under the counterfactual---HeVy07i,HeVy05 call this condition policy invariance.

assumption[Policy invariance] $D(z, a)=\1{\alpha_{z} + \beta_z \cdot a \geq U}$, with $\beta_z \geq 0$.

Under (ref), the only effect of the policy is that it increases the thresholds judges use for release by a judge-specific amount $\beta_{z}$. The policy invariance assumption underlies the use of the marginal treatment effect (MTE) framework HeVy99,HeVy05 for policy analysis. In particular, the representation in (ref) can be used to define the MTE curve, $MTE(u) = E[Y(1)-Y(0) \mid U=u]$. Under (ref), we can then write the change in the average outcome from implementing the counterfactual policy, $\theta - E[Y]$, simply as a functional of the MTE curve. But without (ref), when $U$ may not correspond to judges' rankings under the counterfactual, the MTE curve may be insufficient for counterfactual policy analysis.

Note that (ref) implies both (ref) and (ref), and thus will be questionable whenever (ref) is questionable. Another way to see the difference between (ref) is that if we had data under both the status quo and the counterfactual, we could consider two possible instruments, $Z$ and $A$, where the latter is an indicator for whether the counterfactual policy is implemented. (ref) corresponds to the IV monotonicity assumption where the binary policy variable $A$ is the instrument. (ref) corresponds to IV monotonicity for the multivalued instrument $Z$, while (ref) corresponds to IV monotonicity for the two-dimensional instrument $\tilde{Z} = (Z, A)$. However, as argued in HeUrVy06 and mtw21, with multiple instruments, IV monotonicity is often implausibly strong, as it imposes strong restrictions on treatment choice.

Nevertheless, it is useful conceptually to consider what we can learn under (ref), since as we will see below, for some policy exercises (ref) will be useful while (ref) will not. This suggests that relaxations of (ref) may be useful in practice, a topic we turn to in (ref) below.

The identified set without monotonicity or policy invariance

In this section, we derive a tractable characterization of the identified set $\Theta_I$, without imposing IV monotonicity or policy invariance. Previous work that has considered partial identification in related IV settings has typically characterized the identified set by defining response types (a.k.a.\ principal strata) based on the values of $D(\cdot, \cdot)$ kitagawa21, bai_identifying_2024. The observable data on $(Y, D, Z)$ is then a mixture of the distributions of the potential outcomes across the response types, and one can characterize the identified set for the primitives $\mathcal{P}_I^*(P;\mathcal{P}^{*})$ by searching over mixtures of types that match the observed data. The identified set for $\theta$ can then be calculated by computing the minimum and maximum value of $E_{P^*}[Y(D(Z,1))]$ among $P^* \in \mathcal{P}_I^*(P;\mathcal{P}^{*})$.

The challenge with this approach is that if one does not impose instrument monotonicity, then the number of response types grows exponentially in the number of judges $K$. Specifically, if we impose no restrictions on $D(\cdot, \cdot)$, then there are four possible values of $(D(z,0), D(z,1))$ for each value of $z$, and hence $4^K$ possible response types. Imposing policy monotonicity ((ref)) rules out response types with $(D(z,0), D(z,1)) = (1, 0)$ for some $z$, which still leaves $3^K$ possible response types. In either case, the number of response types becomes extremely large with even a moderate number of judges: when $K=30$, for example, we have $3^K \approx 2 \cdot 10^{14}$ and $4^K \approx 10^{18}$, which makes characterization of the identified set by type enumeration computationally infeasible.

Fortunately, we will show that it is not necessary to enumerate response types to derive the identified set for the counterfactual policy of interest. The key observation is that we need not search over possible distributions $P^*$ for the model primitives; it is sufficient to restrict our attention to the collection of marginal distributions $\{(Y(0), Y(1), D(z, \cdot))\}_{z}$ that involve the potential outcomes and the potential treatments for one judge at a time. That is, we need not explicitly consider the dependence between $D(z,\cdot)$ and $D(z',\cdot)$ for $z \neq z'$. To see why we need not consider the joint behavior of the judges, observe that the counterfactual outcome can be written

equation*[equation* omitted — 101 chars of source]

which depends on $P^*$ only through the collection of marginals for $\{(Y(0), Y(1), D(z,1)) \}_z$ and the marginal distribution of $Z$. Likewise, the probability distribution of the observable data takes the form

equation*[equation* omitted — 84 chars of source]

which depends on $P^*$ only through the collection of marginals for $\{(Y(0), Y(1), D(z,0))\}_{z}$ and the marginal distribution of $Z$. It follows that both the observable data distribution $P$ and the average outcome under the counterfactual depend on $P^*$ only through the marginal distributions $\{(Y(0), Y(1), D(z, \cdot))\}_{z}$ and the marginal distribution of $Z$---they do not depend on the joint distribution of $D(z, \cdot)$ and $D(z', \cdot)$ for $z \neq z'$. Moreover, if we are not imposing any cross-judge restrictions, then our constraints on the model also only restrict the marginals of $\{(Y(0), Y(1), D(z, \cdot))\}_{z}$. This suggests that to compute the identified set for $\theta$, we can simply optimize over sets of marginals for $\{(Y(0), Y(1), D(z, \cdot)) \}_z$ that are consistent with the observable data.

(ref) and (ref) below formalize these observations. For ease of exposition, we state the results for the special case with discrete outcomes, and defer the general statement, which involves more notation, to (ref) in (ref). Suppose that the support $\mathcal{Y}$ of the outcomes is discrete, so that the collection of marginal distributions $\{(Y(0), Y(1), D(z, \cdot)) \}_{z}$ can be characterized by the collection of marginal probability mass functions $\{\pi_{z}(\cdot) \}_{z}$, where $\pi_{z}(y_{0}, y_{1}, d_{0}, d_{1})=P^{*}(Y(0)=y_{0}, Y(1)=y_{1}, D(z,0)=d_{0}, D(z,1)=d_{1})$. Our first result gives a simple way of verifying whether a conjectured collection of marginals $\{\pi_{z}(\cdot)\}_{z}$ is consistent with the data in the sense that there exists a distribution $P^{*}$ that generates both the data distribution $P$ and the conjectured marginals.

propSuppose that $\mathcal{Y}$ is finite. There exists a joint distribution $P^* \in \mathcal{P}^*_{I}(P;\mathcal{P}^{*}_{\textnormal{valid}})$ with marginals $\{\pi_{z}(\cdot) \}_{z \in \mathcal{Z}}$ if and only if $\{\pi_{z}(\cdot) \}_{z \in \mathcal{Z}}$ satisfy the following three conditions: \begin{enumerate} • They match the observable data: for every $y \in \mathcal{Y}$, $z \in \mathcal{Z}$ \begin{align*} \sum_{y_0 \in \mathcal{Y}, d_1 \in \{0,1\}} \pi_z (y_0, y, 1, d_1)&=P(Y=y, D=1 \mid Z=z), \\ \sum_{y_1 \in \mathcal{Y}, d_1 \in \{0,1\}} \pi_z(y, y_1, 0, d_1)&=P(Y=y, D=0 \mid Z=z). \end{align*} • They imply the same distribution for $(Y(0), Y(1))$: for all $y_0,y_1 \in \mathcal{Y}$ and for any $z \in \mathcal{Z}$ \begin{equation*} \sum_{(d_0, d_1) \in \{0,1\}^2} \pi_{z}(y_0, y_1, d_0, d_1) = \sum_{(d_0, d_1) \in \{0,1\}^2} \pi_{1}(y_0, y_1, d_0, d_1). \end{equation*} • They are valid probability mass functions: for all $(y_0, y_1, d_0, d_1) \in \mathcal{Y}^2 \times \{0,1\}^2$, and $z \in \mathcal{Z}$, $ \pi_z(y_0, y_1, d_0, d_1) \geq 0$ with $\sum_{(y_0, y_1, d_0, d_1) \in \mathcal{Y}^2 \times \{0,1\}^2} \pi_z(y_0, y_1, d_0, d_1) = 1$. \end{enumerate}

(ref) implies that to compute the identified set for $\theta$, instead of searching over all distributions $P^{*}$, it suffices to search over the marginal probability mass functions that satisfy the intuitive conditions (ref)--(ref). As the following \namecref{cor:optmize_over_pi} shows, this holds so long as the restrictions on $\mathcal{P}^{*}$ only take the form of marginal restrictions, as in (ref).

corollarySuppose that $\mathcal{Y}$ is finite and that $P^{*}$ satisfies (ref) and (ref) for some convex $\mathcal{R}^{*}$, but no further restrictions are placed on $\mathcal{P}^{*}$, that is $\mathcal{P}^{*}=\mathcal{P}^{*}_{\textnormal{valid}}\cap \{P^{*}\colon \{\pi_{z}(\cdot)\}_{z \in \mathcal{Z}}\in\mathcal{R}^{*}\}$. Then ${\Theta}_{I}(P;\mathcal{P}^{*})$ is given by an interval, with the upper endpoint given by the optimization \begin{equation} \sup_{\{\pi_{z}(\cdot) \}_{z \in \mathcal{Z}} \in \mathcal{R}^*} \sum_{z \in \mathcal{Z}} P(Z=z) \sum_{(y_0, y_1, d_0, d_1) \in \mathcal{Y}^2 \times \{0,1\}^2} \left(d_1 y_1 + (1-d_1)y_0 \right) \cdot \pi_{z}(y_0, y_1, d_0, d_1) \end{equation} subject to the constraints (ref)--(ref) in (ref); the lower endpoint is given by an analogous minimization.

\paragraph{Computation via linear programming}

(ref) implies that if $\mathcal{R}^*$ restricts the marginals $\{\pi_z\}_{z \in \mathcal{Z}}$ linearly, the identified set can be computed by solving a linear program that scales linearly with the number of instruments $K$.

To illustrate, suppose that the only restriction imposed by $\mathcal{R}^{*}$ is that it restricts the fraction of defendants released under the counterfactual, $E[D(z,1)]=\alpha_{1, z}$. For a given collection of marginal probability mass functions $\{\pi_{z}(\cdot)\}_{z}$, let $\pi$ denote the vector of length $4\abs{\mathcal{Y}}^{2}K$ that stacks all values of the marginals. Then the value of the objective function in (ref) can be written $\sum_{z, y_{0}, y_1, d_0, d_1} P(Z=z) (d_1 y_1 + (1-d_1) y_0) \pi_{z}(y_0, y_1, d_0, d_1)= \omega' \pi$, where the weighting vector $\omega$ stacks the weights $P(Z=z) (d_1 y_1 + (1-d_1) y_0)$. Furthermore, the constraints (ref)--(ref) in (ref) can be written as $2\abs{\mathcal{Y}}K+\abs{\mathcal{Y}}^{2}(K-1)+4\abs{\mathcal{Y}}^{2}K+1=O(\abs{\mathcal{Y}}^{2}K)$ linear constraints in $\pi$. Finally, the $K$ constraints $E[D(z,1)]=\alpha_{1, z}$ can be written as $\sum_{y_{0}, y_{1}, d_{0}}\pi_{z}(y_{0}, y_{1}, d_{0},1)=\alpha_{1, z}$, which are also linear in $\pi$. It follows that the optimization problem in (ref) can be written as a linear program in $O(K\abs{\mathcal{Y}}^{2})$ variables, subject to $O(K\abs{\mathcal{Y}}^{2})$ constraints. A related result appeared in RiRo14, who show that for binary outcomes, bounds on $E[Y(d)]$ when $\mathcal{R}^{*}$ is unrestricted can be obtained as a solution to $O(K^2)$ inequalities. Concurrent work by song_categorical_2025 shows that bounds on the marginal distribution of $Y(\cdot)$ when $\mathcal{R}^*$ is unrestricted can be characterized by $O(K)$ inequalities. Their result, however, does not directly cover policy counterfactuals like $\theta = E[Y(D(Z,1))]$, which is a functional of the joint distribution of $Y(\cdot)$ and $D(\cdot, \cdot)$.

The key feature of this program is that its dimension scales linearly with $K$, rather than exponentially, which makes it computationally fast even when $K$ is large. The linear scaling is preserved if we restrict $\mathcal{R}^{*}$ in other ways, so long as the restrictions are linear. For example, imposing policy monotonicity ((ref)) amounts to adding the $K$ constraints that $\sum_{y_{0}, y_{1}}\pi_z(y_0,y_1, 1, 0) = 0$ for all $z=1, \dotsc, K$.

remark[Discretizing $Y$] If $Y$ is continuous, one can obtain conservative bounds by considering a discretized version of $Y$. Given an initial grid of $Q$ points, let $Y^{ub}$ be $Y$ rounded up to the nearest grid point. Since by construction $Y^{ub} \geq Y$, it follows that $E[Y^{ub}(D(Z,1))] \geq E[Y(D(Z,1))] = \theta$. Hence, computing the upper bound of the identified set for $E[Y^{ub}(D(Z,1))]$ yields a potentially non-sharp upper bound on $\theta$. Likewise, one can obtain a conservative lower bound by computing the lower bound of the identified set after rounding down $Y$ to the nearest grid point.
remark[Comparing multiple policies] For simplicity, we have focused on evaluating a single counterfactual policy ($a=1$). In some settings, we may be interested in comparing the impacts of two counterfactual policies (e.g. a quota vs. automatic release), denoted $a=1$ and $a=2$. It is straightforward to extend the approach outlined above to this case by setting $\pi_z$ to correspond to the joint marginal distribution of $(Y(\cdot), D(z,0), D(z,1), D(z,2))$. The objective $E[ Y(D(Z,2)) - Y(D(Z,1))]$ is linear in $\pi_z$, and the dimension scales linearly in $K$ as before.
remark[Estimation and inference] For estimation and inference in our empirical illustrations, we exploit the fact that the data only enters the linear program through the $2\abs{\mathcal{Y}}K$ probabilities $P(Y=y,D=d\mid Z=z)$, which appear on the right-hand side of the data compatibility constraint (ref) and in the constraints on $E[D(z,1)]$ under the quota policy we consider. We form plug-in estimates of the identified set by replacing these probabilities with their sample analogs. For inference, we use a projection approach, whereby we first form a confidence band for these data probabilities, and then optimize over both $\pi$ and data-probabilities within the confidence band. An advantage of this approach is that it can accommodate many-decision-maker asymptotics, where $K$ grows with the sample size. We provide details in (ref).

Does IV monotonicity help tighten the identified set?

The previous section characterized the identified set for the counterfactual outcome $\theta = E[Y(D(Z,1))]$ without imposing IV monotonicity. This section evaluates the extent to which imposing IV monotonicity helps tighten the identified set for $\theta$. We present sufficient conditions under which imposing IV monotonicity alone does not help tighten the identified set. If these conditions hold, then the debate over the validity of IV monotonicity is somewhat of a red herring from the perspective of learning about counterfactuals. A researcher interested in counterfactual policy evaluation should therefore focus on alternative assumptions---such as policy invariance (and relaxations thereof)---that may help tighten the identified set.

\paragraph{Simple example for intuition.} We begin with a simple example to illustrate why imposing IV monotonicity need not help tighten the identified set. Suppose that $a=1$ corresponds to a universal release policy, so that $D(Z,1)=1$ with probability 1. Suppose further that $Y$ is binary and that $Y(0)$ is known to equal zero with probability one. This restriction arises frequently in the criminal justice setting, where, for example, one cannot fail to appear in court if detained while awaiting trial.\footnote{There are other leniency designs in which one of the potential outcomes is known to equal zero. For example, baron_discrimination_2024 study child welfare investigations and define the outcome $Y$ to be 1 if there is evidence of future misconduct in the child's original home, and thus by construction $Y=0$ if a child is placed in foster care (i.e. $Y(1)=0$). Likewise, cgy22 study whether doctors diagnose patients with pneumonia, and define $Y$ to be 1 if the patient subsequently shows signs of undiagnosed pneumonia. Thus, in their setting $Y(1)=0$ by construction.} The parameter of interest then reduces to $\theta = E[Y(1)]$.

In this setting, it is straightforward to show that without IV monotonicity, the sharp identified set that we characterized in (ref) corresponds to an intersection of judge-specific intervals, as suggested in manski_nonparametric_1990 (Lemma 3.1 of bai_identifying_2024 makes a similar observation). In particular, bounds for $\theta$ using only data on a given judge $z$ are given by the interval

equation[equation omitted — 106 chars of source]

The lower bound corresponds to the value of $E[Y(1)]$ if everybody not released by $z$ has $Y(1)=0$; the upper bound to the value if everybody not released by $z$ has $Y(1)=1$. The identified set for $\theta$ then corresponds to the intersection of the judge-specific intervals, $\Theta_I = \bigcap_z \mathcal{I}_z$.

On the other hand, under IV monotonicity, it is straightforward to show that the identified set corresponds to the judge-specific interval for the most lenient judge under the status quo, i.e. $\mathcal{I}_{z_{\max}}$, where $z_{\max} = \operatorname*{argmax}_z E[D(z,0)]$ is the identity of the most lenient judge under the status quo. Specifically, note that IV monotonicity implies that any defendant not released by judge $z_{\max}$ is also not released by any other judge, so that $D(z_{\max},0) =0$ implies $D(z,0) = 0$ for all $z$. It follows that the distribution of the observable data does not depend at all on the value of $Y(1)$ for defendants not released by judge $z_{\max}$, i.e., any binary distribution for $Y(1) \mid D(z_{\max},0)= 0$ is compatible with the observable data. Note, however, that by iterated expectations, $E[Y(1)]$ is a weighted average of $Y(1)$ for the defendants released and not released by judge $z_{\max}$, i.e. $E[Y(1)] = P^{*}(D(z_{\max},0)=1) P^{*}(Y(1) =1 \mid D(z_{\max},0)=1) + P^{*}(D(z_{\max},0)=0) P^{*}(Y(1)=1 \mid D(z_{\max},0) = 0)$. The first term is point-identified from the defendants that judge $z_{\max}$ releases. The identified set therefore corresponds to the interval obtained by placing trivial bounds on the outcome for defendants not released by $z_{\max}$, $P^{*}(Y(1)=1 \mid D(z_{\max},0) = 0) \in [0,1]$, which yields the interval $\mathcal{I}_{z_{\max}}$.

We thus see that without imposing IV monotonicity, we obtain the identified set $\bigcap_z \mathcal{I}_z$, whereas if we do impose monotonicity, we obtain the identified set $\mathcal{I}_{z_{\max}}$. Since adding an assumption cannot widen the identified set, and since $\bigcap_z \mathcal{I}_z$ is contained in $\mathcal{I}_{z_{\max}}$, it follows that when IV monotonicity holds, the identified set is the same whether we impose IV monotonicity or not. On the other hand, if IV monotonicity does not hold and is rejected by the data, it may be the case that the set $\bigcap_z \mathcal{I}_z$ is a strict subset of $\mathcal{I}_{z_{\max}}$: the sharp identified set from (ref) above without imposing monotonicity is then tighter than the naïve identified set $\mathcal{I}_{z_{\max}}$ that assumes IV monotonicity holds in the data (if IV monotonicity is rejected, then the identified set under IV monotonicity is formally empty).

The intuition for why IV monotonicity does not help in this setting is that under IV monotonicity, we learn nothing about the defendants not released by the most lenient judge, $z_{\max}$. By contrast, if IV monotonicity is violated, then some defendants not released by $z_{\max}$ may be released by another judge $z'$, and thus the data provides some information about their value of $Y(1)$. Hence, IV monotonicity is in this sense the least favorable configuration of judge release decisions, and thus imposing IV monotonicity does not help us tighten the identified set.

\paragraph{More general sufficient conditions.} Our simple example above had two salient features: (i) one of the potential outcomes was known ($Y(0)=0$), and (ii) the policy encouragement was very strong ($D(z,1)=1$). We next present a generalization showing that either of these features on its own is sufficient for IV monotonicity not to tighten the identified set. We first show that if $Y(0)=0$, then the identified set for any counterfactual policy satisfying policy monotonicity ($D(z,1) \geq D(z,0))$ does not depend on whether we impose IV monotonicity, regardless of whether the policy encouragement is strong or weak. Second, we show that if both potential outcomes are non-trivial, then IV monotonicity does not help to tighten the identified set provided that one assumes the policy encouragement is “sufficiently strong”, as formalized in condition (iii) below.

We first consider the setting where one of the potential outcomes is known. For simplicity of notation, we establish our result in the context of a counterfactual policy that imposes an average quota policy that restricts the average value of $D(Z,1)$. Fix $\alpha \in [0,1]$. Let $\mathcal{P}_{\alpha}^*$ be the set of distributions over the random vector $(Y(0), Y(1), D(\cdot, \cdot), Z)$ satisfying the \ac{IV} validity condition (ref), policy monotonicity ((ref)), and also

enumerate$\alpha$-average quota policy: $P^*(D(Z,1)=1) = \alpha$, • Known outcome under $D=0$: $P^*(Y(0) = 0)=1$.

Let $\mathcal{P}^*_{\alpha, Mon}$ be the subset of distributions in $\mathcal{P}_{\alpha}^*$ that further satisfy IV monotonicity ((ref)).

prop[No identifying power of monotonicity with known $Y(0)$] If the identified set $\Theta_I(P; \mathcal{P}^*_{\alpha, Mon})$ is non-empty, then (ref) has no identifying power in the sense that $\Theta_I(P; \mathcal{P}^*_{\alpha, Mon})=\Theta_I(P; \mathcal{P}_{\alpha}^*)$.

Although condition (i) focuses on a policy that imposes an average quota, it is possible to extend the proposition to the case in which there is a unit-specific quota $\alpha_{z,1}$, at the cost of additional notation.

We next consider the setting with a sufficiently strong policy encouragement, in the sense that the policy satisfies the following condition:

enumerate• Sufficiently strong encouragement: $P^*(D(z,1) =1 \mid D(z_{\max},0) =1)=1$ for all $z$.

Condition (iii) imposes that all defendants who would be released by the most lenient judge under the status quo would be released with probability 1 under the counterfactual. This is trivially satisfied under a universal release program ($D(z,1)=1$), but is somewhat weaker. For example, consider a program that releases all defendants except those predicted to be high-risk. Then (iii) will be satisfied if the most lenient judge is not currently releasing any such high-risk defendants under the status quo. Our next result shows that monotonicity also has no identifying if we replace condition (ii) defined above with condition (iii) in the statement of (ref).

prop[No identifying power of monotonicity with strong encouragement] The result in (ref) holds if condition (ii) is replaced with condition (iii) in the definitions of $\mathcal{P}^*_{\alpha}$ and $\mathcal{P}^*_{\alpha, Mon}$.
remark[Interaction of IV monotonicity with other constraints] (ref) provide conditions under which IV monotonicity alone does not help tighten the identified set for the average counterfactual outcome. As shown by machado_instrumental_2019, if we impose additional restrictions ---such as assume that all units share the same sign of the treatment effect $Y(1)-Y(0)$---then adding monotonicity can help to further tighten the identified set.
remark[Weaker monotonicity conditions] fll23 propose a weakening of IV monotonicity called average monotonicity. (ref) provide conditions under which imposing IV monotonicity does not help tighten the identified set. It follows immediately that under the same conditions, imposing the weaker notion of average monotonicity also does not help tighten the identified set.
remark[Relationship to literature] (ref) can be viewed as an extension of some important existing results in the literature which show that IV monotonicity has no identifying power for the average treatment effect parameter or the marginal distributions of potential outcomes BaPe97,kitagawa21,bai_identifying_2024, RiRo14. An important difference vis-\`a-vis our result is that we focus on more general policy counterfactuals.
remark[Can monotonicity help?] In (ref), we give an example of a policy and a data distribution such that IV monotonicity sharpens the identified set for $\theta$. This point relates to an observation in kamat_identifying_2019, showing that the identified set for the joint distribution of $(Y(1), Y(0))$ shrinks under IV monotonicity even though the identified sets for the marginals of $Y(1)$ and $Y(0)$ do not. For policies that do not satisfy conditions (ii) or (iii) above, IV monotonicity can potentially help by restricting the possible couplings of the marginal distributions of $Y(0)$ and $Y(1)$, even though it has no identifying power for the marginal distributions.

What assumptions help tighten the identified set?

The previous section outlined sufficient conditions under which IV monotonicity alone does not help to tighten the identified set. We now explore other assumptions that can potentially help to tighten it. We begin by showing that policy invariance can be helpful in some settings where IV monotonicity is not. Policy invariance may be too strong an assumption in practice, however, and we therefore introduce relaxations of policy invariance that may be more plausible, but nevertheless tighten the identified set. We then briefly discuss other economically motivated restrictions that may be reasonable and further help tighten the identified set.

Policy invariance and its relaxations

We begin by providing an intuitive example where the conditions of (ref) are satisfied, so that IV monotonicity alone does not help tighten the identified set, but policy invariance is helpful.

\paragraph{Example: quota policy.} Suppose that $Y(0)=0$ (e.g., you cannot commit a crime while in jail). Consider a quota policy that requires all judges to increase their release rate to match that of the most lenient judge under the status quo, so that $E[D(z,1)] = E[D(z_{\max},0)]$ for all $z$, where again $z_{\max}$ denotes the identity of the most lenient judge. It seems reasonable to impose policy monotonicity in this setting ($D(z,1) \geq D(z,0)$), which states that defendants who would be released without the quota would also be released with the quota.

Strengthening policy monotonicity by imposing policy invariance leads to point identification. Intuitively, under policy invariance, all judges have the same ranking of defendants under both the status quo and counterfactual, but they disagree on the cutoff for when a defendant should be released. The quota policy forces them to use the same cutoff as the most lenient judge, so that the counterfactual outcomes for all judges will match that of the most lenient judge under the status quo: $E[Y(D(Z,1))] = E[Y(D(z_{\max},0))] = E[Y \mid Z =z_{\max}]$.

By contrast, (ref) implies that imposing IV monotonicity in addition to policy monotonicity does not help tighten the identified set. In fact, the identified set may be trivial in the sense that it implies trivial bounds for the outcomes of policy compliers, $E[Y(1)\mid D(Z,1)>D(Z,0)]$. Consider, for example, a setting where there are two judges, who respectively release 10% and 20% of defendants. Under IV monotonicity alone, the first judge could choose to match the quota by marginally releasing only status quo never-takers---those not released by either judge under the status quo. The data does not restrict $Y(1)$ for these individuals, so we only have trivial bounds $[0,1]$ on their treated outcomes, implying trivial bounds for the policy complier treatment effect. Correspondingly, by (ref), the identified set for $\theta - E[Y]$ under IV monotonicity is given by multiplying the unit interval by the mass of policy compliers. Thus, policy invariance---which restricts agreement under judges under both the counterfactual and the status quo---has strong identifying power, leading to point identification. By contrast, IV monotonicity---which only restricts agreement under the status quo---does not help tighten the trivial bounds on the identified set.

Of course, policy invariance may be implausibly strong in many applied settings. Indeed, since policy invariance is stronger than IV monotonicity, if researchers doubt the validity of IV monotonicity (or it is rejected by the data) then they must necessarily doubt the validity of policy invariance as well. Nevertheless, the fact that there are relevant settings where policy invariance can help tighten the identified set, but IV monotonicity cannot, suggests that considering relaxations of policy invariance may be a more natural starting point than considering relaxations of IV monotonicity.

\paragraph{Relaxing policy invariance with disagreement bounds.} To that end, we now introduce a relaxation of policy invariance that may be more plausible but nevertheless help to tighten the identified set. At a high level, policy invariance requires that judges perfectly agree on the ranking of defendants; we consider a relaxation that bounds the extent to which they can disagree. More concretely, consider two distinct judges $z$ and $z'$, and two policies $a, a'\in \{0,1\}$. Without loss of generality, suppose that $E[D(z, a)] \geq E[D(z', a')]$ so that judge $z$ releases more people under policy $a$ than $z'$ does under $a'$. Under policy invariance, judge $z$ under policy $a$ must release all defendants released by $z'$ under policy $a'$, so that $P^*(D(z, a)=1 \mid D(z', a')=1) = 1$. Such perfect agreement is likely to be too strong in many settings. Nevertheless, it also may be unreasonable to expect that the two judges perfectly disagree, so that judge $z$ under policy $a$ releases none of the defendants released by $z'$ under $a'$. A natural middle-ground is to impose

equation[equation omitted — 105 chars of source]

so that judge $z$ under policy $a$ disagrees with no more than $\delta_{z, z', a, a'}$ fraction of defendants released by judge $z'$ under policy $a'$. This nests policy invariance as the special case with $\delta_{z, z', a, a'} = 0$, but allows for non-trivial disagreement for $\delta_{z, z', a, a'} \in (0,1)$.\footnote{IcTa00 assume that $D(z,0) = D(z',1)$ (a.s.) for a known pair (or set of pairs) $z,z'$. This can be formalized by imposing (ref) with $\delta_{z,z',0,1} =0$ and $\delta_{z',z,1,0} = 0$.}

\paragraph{Calculating the identified set under disagreement bounds.} The disagreement bound in (ref) imposes restrictions on the joint distribution of $(D(z, a), D(z', a'))$. At first glance, such restrictions appear to be difficult to incorporate into the linear programming approach for calculating the identified set described in (ref), which only optimizes over the marginals $\{\pi_{z}(\cdot)\}$, and does not enumerate the joint distribution of $(D(z, a), D(z', a'))$. It turns out, however, that given a set of marginals $\{\pi_{z}(\cdot)\}$, there is a simple formula for the minimal value of $P^{*}(D(z, a)=0 \mid D(z', a')=1)$ consistent with a joint distribution of primitives that matches these marginals and satisfies policy monotonicity. This allows us to still tractably compute the identified set under policy monotonicity by optimizing only over the marginals $\{\pi_{z}(\cdot)\}$ even while imposing disagreement bounds such as (ref). This is formalized in the following results, which give analogs to (ref) and (ref) that allow for imposing disagreement bounds.

propSuppose that $\mathcal{Y}$ is finite. Fix disagreement bounds $\delta_{z, z', a, a'} \in [0,1]$ for all $z, z', a, a'$. Let $\mathcal{P}^*_{\textnormal{DB}}$ be the subset of distributions in $\mathcal{P}^{*}_{\textnormal{valid}}$ satisfying (ref) and the disagreement bounds in (ref) for all $z, z', a, a'$. Consider a collection $\{\pi_{z}(\cdot)\}_{z\in\mathcal{Z}}$ of marginals. There exists a joint distribution $P^* \in \mathcal{P}^*_{I}(P;\mathcal{P}^{*}_{\textnormal{DB}})$ with marginals $\{\pi_{z}(\cdot) \}_{z \in \mathcal{Z}}$ if and only if the marginals $\{\pi_{z}(\cdot) \}_{z \in \mathcal{Z}}$ are consistent with (ref) and satisfy conditions (ref)--(ref) in (ref) as well as \begin{itemize} • Disagreement bound: For all $z, z', a, a'$, \begin{multline} \sum_{y_1, y_0} \min\{\pi(y_0, y_1, D(z, a) = 1), \pi(y_0, y_1, D(z', a') = 1) \} \\ \geq (1-\delta_{z, z', a, a'}) \cdot \pi(D(z', a')=1), \end{multline} where for any $(y_0,y_1) \in \mathcal{Y}^2$ and $(a, z)\in \{0,1\} \times \mathcal{Z}$, \begin{align*} \pi(y_0, y_1, D(z, a)=1) &:= \sum_{d_a=1, d_{1-a} \in \{0,1\}} \pi_{z}(y_0,y_1, d_0, d_1), \\ \pi(D(z, a)=1) &:= \sum_{d_{a}=1, d_{1-a}\in \{0,1\}} \sum_{(y_0,y_1) \in \mathcal{Y}^2} \pi_{z}(y_0, y_1, d_0, d_1). \end{align*} \end{itemize}
corollarySuppose that $\mathcal{Y}$ is finite and let $\mathcal{P}^* = \mathcal{P}_{DB}^* \cap \{P^{*}\colon \{\pi_{z}(\cdot)\}_{z \in \mathcal{Z}}\in\mathcal{R}^{*}\}$ for some convex $\mathcal{R}^*$, where $\mathcal{P}_{DB}^*$ as defined in (ref). Then ${\Theta}_{I}(P;\mathcal{P}^{*})$ is given by an interval, with the upper endpoint given by the optimization \begin{equation} \sup_{\{\pi_{z} \}_{z \in \mathcal{Z}} \in \mathcal{R}^*} \sum_{z \in \mathcal{Z}} P(Z=z) \sum_{y_0,y_{1}, d_0, d_1 \in \mathcal{Y}^2 \times \{0,1\}^2} \left(d_1 y_1 + (1-d_1)y_0 \right) \cdot \pi_{z}(y_0, y_1, d_0, d_1) \end{equation} subject to the constraints (ref)--(ref) given in (ref) and subject to policy monotonicity ($\pi_z(y_0, y_1, 1, 0) = 0$ for all $z,y_0,y_1$); the lower endpoint is given by an analogous minimization.

The intuition for (ref) is as follows. Note that (ref) can equivalently be written as a lower bound on the joint probability $P^*(D(z, a) = 1, D(z', a') = 1)$,

equation[equation omitted — 134 chars of source]

By the law of total probability, the right-hand side equals

multline*[multline* omitted — 198 chars of source]

where the inequality uses the fact that $P(A \cap B) \leq \min\{ P(A), P(B) \}$. (ref) follows from replacing $P^*(D(z, a) = 1, D(z', a') = 1)$ in (ref) with the upper bound given in the previous display. This turns out not to come at the cost of sharpness, however, because given a set of marginals for $(Y(1), Y(0), D(z, \cdot))$, there always exists a coupling such that the inequality in the previous display holds with equality. In particular, the proof of (ref) shows that the upper bound is achieved under a latent threshold crossing model wherein judges agree on the rankings of defendants conditional on their potential outcomes, so that $D(z, a) = \1{\alpha_{z, y_0, y_1} + \beta_{z, y_0, y_1} a \geq V_{y_0, y_{1}}}$ for a latent index $V_{y_0, y_1}$ that depends only on the potential outcomes but not on $z$.\footnote{If one does not impose policy monotonicity, then (ref) is still implied by (ref), and so one can obtain an outer set for $\Theta_I$ by running the analog to the optimization in (ref) without the policy monotonicity constraint. Without policy monotonicity, however, the interval obtained may no longer be sharp.}

We note that it is straightforward to impose (ref) in a linear program by introducing auxiliary parameters corresponding to the $\min\{\cdot, \cdot\}$ terms. Specifically, (ref) holds if and only if there exist constants $\eta^{y_0,y_1}_{z, z'} \geq 0$ that are less than the arguments of the $\min\{\cdot, \cdot\}$,

align[align omitted — 206 chars of source]

such that

equation[equation omitted — 131 chars of source]

The above formulation is linear in both the marginals $\{\pi_{z}(\cdot)\}$ and the auxiliary parameter $\eta$. Therefore, we can implement the disagreement bound in (ref) by adding the constraints in (ref) to the linear programming formulation in (ref), and augmenting the parameter vector $\pi$ by $\eta$.\footnote{We note that if one bounds disagreement rates among all pairs $z,z'$, this introduces $O(K^2)$ auxilliary parameters, so the dimension of the LP scales quadratically in $K$ rather than linearly.}

\paragraph{Bounds on average disagreements.} The results above can easily be adapted if we wish to weaken (ref) by only bounding the average pairwise disagreement between groups of judges. For concreteness, consider the quota policy from our empirical applications in (ref), which asks the bottom 90% of judges to match the release rate of the most lenient decile. Let $\mathcal{Q}$ denote the set of judges in the bottom nine deciles of leniency, who are subject to the quota, and let $\mathcal{Q}^c$ denote the top decile of judges, who are not subject to the quota. To bound the average disagreement between the two groups of judges under the counterfactual, we can impose the average disagreement probability bound

equation[equation omitted — 174 chars of source]

where $Z_{\mathcal{Q}}$ is a randomly picked judge from the group $\mathcal{Q}$, and ${Z}_{\mathcal{Q}^{c}}$ is defined similarly. In other words, (ref) places an upper-bound $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ on the average probability that a judge in the quota group disagrees with the release decision of a judge in the top decile. Similarly to the pairwise disagreement bounds, this average disagreement probability bound can be implemented as a linear program without losing sharpness of the resulting identified set. In particular, we can introduce auxiliary parameters $\eta_{z, z'}^{y_{1}, y_{0}}$, for each $z\in\mathcal{Q}$ and $z'\in\mathcal{Q}^{c}$, and impose the linear constraints (ref), as well as the constraint $\sum_{y_{1}, y_{0}}\frac{1}{\abs{\mathcal{Q}}\abs{\mathcal{Q}^{c}}}\sum_{z\in\mathcal{Q}, z'\in\mathcal{Q}^{c}} \eta_{z, z'}^{y_{0}, y_{1}}\geq (1-\delta)P^*(D(Z_{\mathcal{Q}^{c}}, 1)=1)$. The attractive feature of the average disagreement probability bound is that it only requires calibration of a single parameter, $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$. In contrast, imposing pairwise disagreement bounds requires calibration of $\delta_{z, z'}$ for each pair of judges, which may be more difficult in practice.\footnote{Our average disagreement bound focuses on the case where we only bound disagreement under the counterfactual. This is sufficient for the quota counterfactual, since we assume judges in the top decile do not change their behavior under the counterfactual. For other policies, it may be informative to impose bounds on $P^*(D(Z_{\mathcal{Q}}, a)=0\mid D({Z}_{\mathcal{Q}^{c}},a')=1)$ for pairs of $(a, a')$ other than $(1,1)$.}

\paragraph{Calibrating disagreement bounds.} In certain special cases, we can observe the decisions of multiple judges (or other decision-makers) on the same cases, which allows us to directly measure how frequently they disagree. Such special settings can be used to calibrate reasonable bounds on disagreement rates. For example, sigstad_monotonicity_2023 studies settings where panels of judges rule on the same cases. Likewise, agarwal_combining_2023 ask multiple radiologists to provide diagnoses on the same x-rays, and thus observe the decisions of multiple doctors on the same cases.\footnote{In a similar vein, ady23 observe the recommendation of both a pretrial services officer and a judge. One could use this to calibrate disagreement bounds among pairs of judges if one were willing to assume that judges and pretrial services officers disagree at similar rates as pairs of judges.} In our application to bail judges in (ref) below, we calibrate the disagreement bound $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ using a Gaussian model parameterized to match the data in sigstad_monotonicity_2023.\footnote{In principle, one could use the sigstad_monotonicity_2023 data to directly calculate the sample analog of the disagreement probability $P^*(D(Z_{\mathcal{Q}},0)=0\mid D({Z}_{\mathcal{Q}^{c}},0)=1)$. However, the disagreement probability is not invariant to the marginal release rates (if two judges have release rates equal to 99%, they cannot disagree on more than 2% of cases; by contrast, two judges with release rates of 50% could disagree 100% of the time). We therefore use a Gaussian signal model to compute the implied disagreement probability when the marginal release rates are matched to those in our application (see (ref)). The Gaussian signal model implicitly assumes that judges serving on panels rule the same way as when making decisions on their own. To weaken this assumption, one could estimate a richer model that allows for spillovers IaSh12.}

\paragraph{Limitations of policy invariance.} The quota example given above showed that policy invariance can help tighten the identified set (and even restore point identification) in a setting where IV monotonicity alone is not informative. It turns out, however, that there are some situations in which neither IV monotonicity nor policy invariance helps to tighten the identified set. To gain intuition, consider our earlier example of a universal release policy such that $D(Z,1)=1$ with probability 1. In that case, by assumption, all judges perfectly agree on their decisions under the counterfactual policy. It follows that imposing policy invariance---which imposes IV monotonicity and restricts agreement under the counterfactual---is equivalent to imposing IV monotonicity in this setting, which we know is not helpful by (ref). In settings like this, researchers will therefore have to turn to alternative restrictions to try to tighten the identified set.

Outcome restrictions

We now outline several other restrictions that researchers may consider imposing to help tighten the identified set.

\paragraph{Policy complier (PC) bounds.} In the pretrial release setting, bail judges are legally instructed to release defendants based on pretrial misconduct potential if released ($Y(1)$). Consider a policy that encourages judges to release more defendants (e.g., a quota). Unless the judges' estimates of pretrial misconduct risk are terribly miscalibrated, policy compliers---individuals only released under the quota---should be on average higher risk than those currently released, but lower risk than those who are never released:

equation[equation omitted — 113 chars of source]

When the marginal release rates under the counterfactual ($E[D(z,1)]$) are point-identified, (ref) can be written as a linear restriction on the marginals $\pi_{z}(\cdot)$, and is thus is straightforward to incorporate in the linear program for the identified set bounds.

\paragraph{Treatment effect bounds.} It may sometimes be reasonable to impose bounds on the sign or magnitude of the treatment effects. For example, in some settings, it may be natural to impose the monotone treatment response assumption that $Y(1) \geq Y(0)$, as in manski97 and machado_instrumental_2019. This is straightforward to impose by setting $\pi_{z}(y_0, y_1, d_0, d_1) =0$ whenever $y_1 < y_0$. Similarly, in some settings it may be reasonable to impose that the treatment effect for the policy compliers is not too large, e.g. $E[Y(1)-Y(0) \mid D(z,0)<D(z,1)] \leq c$, which can be implemented as a linear constraint when the share of policy compliers is identified, similar to (ref).\footnote{fll23 argue, for example, that it may be reasonable to restrict the magnitude of the treatment effect of incarceration among compliers between different sets of judges; here we impose an analogous constraint on the compliers with respect to a policy of interest.}

\paragraph{Outcome disparity bounds.} Some counterfactual policies directly encourage a subset $\mathcal{Q}$ of decision-makers to act more like a benchmark group $\mathcal{Q}^{c}$. For example, in implementing a quota policy directive that asks the set $\mathcal{Q}$ to match the treatment rate of the group $\mathcal{Q}^{c}$, a policy-maker may encourage the decision-makers to also emulate their outcomes. We consider such a setting in our empirical application in (ref) below. In such cases, an alternative to bounding the disagreements between their treatment decisions is to bound the disparity in the outcomes between the two sets of decision-makers by imposing

equation[equation omitted — 172 chars of source]

where, again, $Z_{\mathcal{Q}}$ is a randomly picked judge from the group $\mathcal{Q}$, and ${Z}_{\mathcal{Q}^{c}}$ is defined similarly.

To interpret this bound, it can be helpful to relate it to the disagreement probability bound in (ref) above. To this end, let $\mathcal{C}_{z, z'}=\1{D(z, 1)=0, D(z', 1)=1}$ denote the event that $z'$ would release an individual but not $z$ under the counterfactual. Assume that the treatment rates of both groups are equal, $P^*(D(Z_{\mathcal{Q}^c},1)=1)=P^{*}(D(Z_{\mathcal{Q}}, 1)=1)=q$ for some $q$. Then $E[\mathcal{C}_{{Z}_{Q}, {Z}_{\mathcal{Q}^{c}}}]=E[\mathcal{C}_{{Z}_{Q^{c}}, {Z}_{\mathcal{Q}}}]=P^*(D(Z_{\mathcal{Q}}, 1)=0, D(Z_{\mathcal{Q}^c},1)=1)$. Using this, we may write the left-hand side of (ref) as

multline*[multline* omitted — 304 chars of source]

This is the product of the aggregate disagreement probability, $P^*(D(Z_{\mathcal{Q}}, 1)=0, D(Z_{\mathcal{Q}^c},1)=1)=P^*(D(Z_{\mathcal{Q}}, 1)=0\mid D(Z_{\mathcal{Q}^c},1)=1)q$ times the difference between treatment effects for individuals about whom there is disagreement. Thus, if we bound the treatment effect difference by $\Delta_{TE}$, imposing (ref) implies that (ref) holds with

equation[equation omitted — 157 chars of source]

This relationship can be used to gauge the strength of the outcome disparity restriction, as we illustrate in (ref) below.

Empirical applications

NYC bail judges

\paragraph{Background and motivation.} We study pretrial release in New York City (NYC) using data from adh22. A judge ($Z$) decides whether to release a defendant or hold them in pretrial detention. The defendant is categorized as released ($D=1$) if they are released without preconditions or after paying money bail; otherwise they are categorized as detained ($D=0$). We then see whether a defendant commits pretrial misconduct ($Y \in \{0,1\}$) by either failing to appear for a court hearing or by being arrested for a new crime. By construction, a defendant cannot commit pretrial misconduct if they are detained, so $Y(0) = 0$. We also observe the defendant's race, denoted by $R \in \{b, w\}$, where $b$ and $w$ denote black and white.

We consider two counterfactual policy exercises in this context. First, we consider a policy that releases all defendants. \Textcite{albright22} studies a universal release program for defendants (meeting certain conditions) in Kentucky, so this policy counterfactual directly addresses what would happen if such a program were implemented for ADH22's sample of defendants in NYC\@. Moreover, as we describe below, identification of the parameter of interest in ADH22, the disparate impact parameter, is equivalent to identification of race-specific average outcomes under this counterfactual. As a result, our analysis also delivers bounds on the disparate impact parameter without imposing the parametric restrictions that allowed ADH22 to point-identify this parameter by extrapolating from the observed variation in the data using identification-at-infinity type arguments.

Second, we consider a quota policy that asks the judges in the bottom 90% of leniency to match the release rate of the top 10%. The original ADH22 analysis considered counterfactuals similar to this using a hierarchical MTE model, whereby a subset of the judges were subject to release quotas. Furthermore, several policies discussed in the literature encourage stricter judges to release more defendants while still giving the judges discretion. For example, beginning in 2016, New York City began encouraging judges to release some defendants through a supervised release program rather than requiring them to post bail skemer_pursuing_2020. Likewise, Kentucky passed a law in 2011 that, among other things, changed the presumptive default from monetary bail to non-monetary release, yet provided judges with the discretion to override the default stevenson_assessing_2018. Both policies were found to increase the fraction of defendants who were released, but did not increase this rate to one. Our quota exercise can be thought of as a simple approximation to reforms like these if one assumes that their main practical impact is to induce the bottom 90% of judges to increase their release rate to match the most lenient decile. Of course, if one had domain-specific knowledge about the likely impact of such reforms on release rates, the assumptions about their “first-stage” impact could be modified accordingly to make our quota counterfactual resemble them more closely.

\paragraph{Disparate impact in ADH22.} ADH22's primary goal is to estimate the “disparate impact” of pretrial release decisions by race. Their disparate impact parameter can be written as a function of observable probabilities and $E[Y(1) \mid R]$. Hence, learning about the disparate impact parameter is isomorphic to learning about the mean outcome (for each race) under the counterfactual in which everyone is released. To make this precise, let $FPR_r := P(D=1 \mid Y(1) =1, R=r)$ be the false positive rate for race $r$, i.e., the fraction of defendants who are released despite having $Y(1)=1$. Observe that we can write $FPR_r = \frac{P(D=1, Y=1 \mid R=r)}{P^*(Y(1) =1 \mid R=r)}$. The numerator is an observable probability, and thus the only challenge in learning about $FPR_{r}$ is learning about the misconduct rate under the counterfactual in which everyone is released, $P^*(Y(1) =1 \mid R=r)$. ADH22 similarly define $TPR_r := P(D=1 \mid Y(1) =0, R=r)$ to be the true positive rate, which we can write as $TPR_r = \frac{P(D=1, Y=0 \mid R=r)}{1-P^*(Y(1) =1 \mid R=r)}$. Again, the numerator is an observable probability. ADH22 define the disparate impact parameter as a weighted average of the differences in $FPR_r$ and $TPR_r$ across races,

equation*[equation* omitted — 99 chars of source]

where $E[Y(1)]=\sum_{r\in\{b, w\}}P(R=r)P^{*}(Y(1)=1\mid R=r)$. This relationship allows us to infer bounds on $\Delta$ from bounds on the race-specific counterfactual outcomes if everyone is released, $E[Y(1) \mid R=r]$.

\paragraph{Identification in ADH22.} ADH22 consider an identification-at-infinity type argument for $E[Y(1) \mid R]$. Let $P_{z, r} = E[D(z,0) \mid R=r]$ be the judge-specific release rate for race $r$. ADH22 consider various functional forms for $E[Y(1) \mid D=1,P_{z, r}=p, R=r]$ and then extrapolate to $p=1$. Specifically, they report linear, local linear, and quadratic extrapolations. For comparison with our nonparametric approach, we replicate the point estimates and confidence intervals based on these three parametric extrapolation approaches.

\paragraph{Data.} We analyze aggregated statistics from the main estimation sample in ADH22, which consists of the universe of arraignments made in NYC between 2008--2013 involving white or black defendants charged with a felony or misdemeanor, where the defendant is not already serving jail time for an unrelated charge. The sample comprises 284,598\ cases involving white defendants and 310,588\ cases involving black defendants. We do not have access to the individual microdata, and only observe estimates of the judge-specific release rates ($E[D(z,0)\mid R=r]$) and misconduct rates $E[Y(1)\mid D(z,0)=1,R=r]$ that are obtained from linear regressions that adjust for court-by-time fixed effects, along with the number of observations for each judge and race (these estimates are plotted in Figure 2 in ADH22). The covariate adjustment accounts for the conditional random assignment of the judges; see (ref) for details. Because we do not have access to the microdata, our inference ignores the statistical uncertainty stemming from the covariate adjustment: we treat the covariate-adjusted rates as if they were sample means. We show below that our replication of the results in ADH22 yields very similar standard errors as in the original, suggesting that the impact of the covariate adjustment on inference is minimal. To avoid small-sample issues stemming from observing only a few cases per judge, we pool all judges with fewer than 300 race-specific cases into a single judge, which leaves us with $K=167$ and $K=160$ distinct values for the instrument for the samples of black and white defendants, respectively.\footnote{Since the quota policy we consider below applies a quota to judges below the 90th percentile of leniency, we separately pool judges with fewer than 300 cases who lie below the 90th percentile of leniency, and those lying above it. (ref) show that our inference is virtually unchanged when we do not pool.}

\paragraph{Summary statistics.}

(ref) plots the judge-specific misconduct rates against their release rates. We see that the misconduct rates increase sharply with the release rate, which is intuitive since by construction $Y=0$ for non-released defendants. Correspondingly, the \ac{TSLS} estimate of the effect of release on misconduct, which is equivalent to fitting a weighted least squares line through this scatterplot, is quite high, and equals 49.7% for blacks and 48.1% for whites (with tight standard errors: 1.2\ and 2.0). The average release rate equals 69.7% for blacks and 76.5% for whites. Unconditionally, the average misconduct rate equals 22.9% for blacks and 20.6% for whites; conditional on release, these rates equal 32.9% and 26.9%.

figure[figure omitted — 51,866 chars of source]

\paragraph{Results for universal release policy.} We compute plug-in estimates of the identified set for $E[Y(1) \mid R]$, the race-specific average outcomes under a universal release program, along with 95% confidence intervals computed by projection (see (ref) for inference details). Because $Y(0) = 0$, the sharp bounds on $E[Y(1) \mid R=r]$ imposing only IV validity (i.e., randomization and exclusion) are given by the intersection of judge-specific intervals given in (ref) above. For illustration, (ref) shows the form of the bounds (along with simultaneous CIs) for black defendants: we compute sample analogs of the judge-specific bounds in (ref), and intersect them to estimate the identified set. While under monotonicity all the information is obtained from the most lenient judge, we see that in practice the estimate of the identified set exploits information from other judges and thus is tighter than the interval obtained by looking at the most lenient judge alone.

figure[figure omitted — 14,402 chars of source]
table[table omitted — 2,207 chars of source]

Columns (1) and (3) of (ref) give the estimates and confidence intervals for the race-specific average misconduct rates computed using our linear program as well as the parametric extrapolation methods in ADH22\@. For direct comparability to the confidence intervals using our approach, we compute standard errors without accounting for covariate adjustment.\footnote{ADH22 also weight their extrapolations by the inverse of the estimated sampling variance for the judge-level covariate-adjusted outcome means. Since we do not have these weights, we instead weight by the number of defendants the judge released.} (ref) in the appendix shows that the results are very similar to those reported in ADH22, suggesting that accounting for covariate adjustment doesn't substantively affect inference.

To help interpret the estimates, we also report estimates of the treatment effects for policy compliers, $E[Y(1)-Y(0)\mid D(Z,1)>D(Z,0)]=E[Y(1)-Y(0)\mid D(Z,0)=0]$ in columns (2) and (4). The lower bound for policy compliers is above the misconduct rate for white defendants released under the status (31.9% vs.\ 26.9%) and very similar to the status quo rate for black defendants (30.8% vs.\ 32.9%), suggesting that the marginally-released defendants will have similar or higher crime rates to those released under the status quo. This is intuitive, since we expect judges to be releasing defendants who they deem to be the lowest risk. The confidence intervals are also reasonably tight, allowing us to conclude that a universal release program will lead to a substantial increase in the misconduct rates relative to the status quo: for blacks, they imply that the misconduct rate will increase by at least $6.3$ percentage points, which is a substantial increase relative to the status quo rate, $22.9\%$.

We also compute an estimate of the identified set and confidence intervals for the disparate impact parameter $\Delta$. Comparing the resulting estimates to those from the parametric extrapolation methods, we see in (ref) that our estimate of the identified set, $4.0$--$8.2$%, is somewhat wider than the range of estimates we obtain based on the parametric extrapolations, 4.3--5.4%. Our 95% CIs are also informative: they equal $(1.0, 9.6)$%, so we can reject the null hypothesis of no disparate impact. Thus, the conclusion in ADH22 that $\Delta > 0$ is robust to dropping their functional form assumptions.\footnote{In a robustness check, ADH22 construct non-sharp bounds on $E[Y(1) \mid R]$ and $\Delta$ using the aggregate mean outcome among released defendants, rather than constructing the sharp bounds by intersecting the judge-specific intervals. This yields wider bounds than what we obtain. For example, the plug-in bounds on $\Delta$ from this approach are $-1$% to 10%.}

\paragraph{Results for the quota policy.}

table[table omitted — 2,685 chars of source]

To benchmark the effect of the quota policy, we first compute the policy effect of a simpler policy that achieves the same release rates: reallocating all cases to the most lenient decile. Here the policy effect is point identified under random assignment of the instrument, and it is given by the difference between the current average misconduct rate versus the average misconduct rate for individuals assigned to judges in the most lenient decile.\footnote{We note that for analyzing the impact of such a reallocation policy, we do not actually require IV exclusion.} We also consider a simple parametric benchmark that assumes homogeneous treatment effects, which also leads to point identification: the policy effect is identified by the TSLS estimate, multiplied by the change in the release rate relative to the status quo. As shown in (ref), both benchmarks imply a similar policy effect, implying a misconduct rate for policy compliers equal to about 50% for both whites and blacks.

We then estimate the policy effect of the quota under a conservative specification that assumes only \ac{IV} validity and policy monotonicity. Unlike for the universal release policy, the bounds for this policy are trivial in the sense that they imply the trivial bounds $[0, 100]$ for the policy complier treatment effect. The reason for this is that we cannot rule out that there is a sufficient mass of instrument never-takers under the status quo (those with $D(z,0)=0$ for all $z$) for whom we only have trivial bounds on $Y(1)$, and without further restrictions, policy compliers may be drawn from this pool in an adversarial way.

However, we can tighten these bounds by imposing two additional sets of restrictions that exploit the institutional details. First, we can impose the aggregate disagreement probability bound in (ref). As described in detail in (ref), we set the upper bound on disagreement, $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$, by calibrating a Gaussian signal model similar to that considered in cgy22 to sigstad_monotonicity_2023's data on panels of judges ruling on criminal cases in the São Paulo Appeal Court. We assume that the judges observe correlated Gaussian signals $U_{z}$, releasing the individual if the signal is high enough. To calibrate the correlation matrix, we focus on the first three judges to rule in each case in sigstad_marginal_2024's sample.\footnote{At least three judges rule on each case. Additional judges may provide an opinion in some cases, but only if the initial three judges disagree. To estimate disagreement rates, we therefore focus on the first three judges to avoid selection bias related to the initial three judges' decisions.} The implied signal correlation between a randomly selected judge in the bottom 90% and one in the most lenient decile is 0.989.\footnote{The raw disagreement rates are as follows: for a randomly-selected pair of judges $A$ and $B$ with judge $A$ in the top decile of leniency and judge $B$ in the bottom nine deciles in Sigstad's data, we have $P(D_A=1, D_B=0) = 3.8\%$ and $P(D_A=1, D_B=0) = 0.6\%$, where $D_A, D_B$ are the decisions of judges $A, B$, respectively.} To calibrate $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ in the NYC data, we set the correlation between any pair of signals $U_{z}$ and $U_{z'}$, where $z$ is a judge in the bottom 90% and $z'$ is in the top decile, to this value in our Gaussian signal model, and set the judges' cutoffs for release to match those under our quota counterfactual. We calculate the value of $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ as the average disagreement probability in this model, which yields $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu=0.025$ blacks and 0.021\ for whites.

Second, the sole legal objective of bail judges is to allow most defendants to be released before trial while minimizing the risk of pretrial misconduct. In line with this objective, we can impose the policy complier bounds given by (ref) for each judge $z$ in the bottom 90%.

(ref) gives the results for the quota policy when we impose either the calibrated disagreement bound or the policy complier bounds (or both). Confidence intervals are computed by projection, as detailed in (ref).

With the $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ bound imposed, the identified set estimate is fairly tight, implying that the treatment effects for policy compliers are not too different from the TSLS estimates. The estimates of the bounds exceed the status quo misconduct rates among those currently released, $E[Y(1)\mid D(Z,1)=1]$: these rates equal $32.9$% for blacks and $26.9$% for whites, while the lower bounds on the average value of $Y(1)$ for policy compliers, reported in column (2), equal $44.9$% and $35.7$%, respectively. While the confidence intervals are wider, the lower endpoints are close to the status quo misconduct rates, particularly for blacks ($26.3$% vs.\ $32.9$%), implying the policy compliers have similar or higher misconduct rates than the inframarginal individuals (those currently released). Additionally, imposing policy complier bounds helps tighten the point estimates, but not the confidence intervals once the $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{DP}\mkern-1.5mu} \mkern 1.5mu$ bound is imposed.

Suffolk County prosecutors

\paragraph{Background.} We reanalyze the data from adh23, who are interested in the effect of non-prosecution of nonviolent misdemeanor offenses on the subsequent criminal justice involvement of the defendants. Here the decision-makers ($Z$) are assistant district attorneys (ADAs), who decide whether to prosecute a case after an initial arraignment ($D=1$) or drop it ($D=0$). We then see whether a criminal complaint against the defendant has been filed within two years of arraignment ($Y\in\{0,1\}$). ADH23 use the jackknife IV estimator to estimate the treatment effect of non-prosecution, and find robustly negative IV estimates; they also estimate an MTE curve that is downward-sloping and negative for nearly all of its support. Based on this analysis, they conclude (ADH23, p. 1455):

footnotesize\begin{quote} The results of our analysis imply that if all arraigning ADAs acted more like the most lenient ADAs in our sample when deciding which cases to prosecute, Suffolk County would likely see a reduction in criminal justice involvement for these nonviolent misdemeanor defendants. \end{quote}

If we assume homogeneous treatment effects, then the negative IV estimates do indeed imply that increasing non-prosecution would lower criminal complaints. Likewise, if we impose policy invariance, then the negative MTE curve implies that increasing the non-prosecution rate would lower criminal complaints. We are interested in evaluating the robustness of this conclusion to dropping the assumptions of homogeneous treatment effects and policy invariance. Importantly, when ADH23 test IV monotonicity ((ref)) for each of the 9 courts in their dataset using the test developed by fll23, the test rejects IV monotonicity (and hence also policy invariance) in three of the courts. Correspondingly, we conduct the analysis without imposing (ref) or (ref).

To apply our policy evaluation approach, we need to make precise the way in which one would encourage stricter ADAs to act “more like the most lenient ADAs”. As a benchmark, we start by analyzing a policy that reallocates cases to the more lenient ADAs. We next consider a scenario in which, rather than literally re-allocating cases, the policy encourages the stricter ADAs to increase their release rates to match that of the more lenient ones. For comparability with our previous application, we consider a quota on the bottom 90% of ADAs to match the non-prosecution rate of the most lenient decile. As with the previous application, we can view this as an approximation to other encouragement policies that have the effect of increasing the non-prosecution rates of the bottom 90% of ADAs to match that of the top 10%. For example, ADH23 discuss how in 2019, a new district attorney increased non-prosecution rates by issuing a memo that instituted a presumption of non-prosecution for certain misdemeanor cases. Our results approximate the impact of such a policy to the extent that its impact on release rates was to make the bottom 90% of ADAs match the top 10%.

\paragraph{Data.} The data comes from the Suffolk County District Attorney’s Office in Massachusetts. We focus on the main analysis main sample in ADH23, which restricts to cases between 2004 and 2008, drops felony charges, violent crimes, and prosecutors with fewer than 30 cases. This yields a sample with $67,060$ cases. ADH23 argue that conditional on court-by-time fixed effects, the ADAs are as-good-as-randomly assigned to cases, and they control for these fixed effects linearly in all their specifications. In their main specification, they also consider an extended set of controls that include case and defendant characteristics. Correspondingly, to estimate ADA-specific probabilities $P^*(Y(d)=y, D(z,0)=d)$ we also linearly adjust for the court-by-time fixed effects and case and defendant characteristics, as detailed in (ref). Since we have access to the microdata, our inference accounts for this covariate adjustment. As in the NYC bail judge application, we pool ADAs with fewer than 300 cases to mitigate finite-sample issues (we do this separately by whether leniency is above or below the 90th percentile). (ref) shows that our inference results remain the same when we do not pool. This leaves us with $K=69$ distinct values for the instrument.

\paragraph{Summary statistics.}

(ref) plots the ADA-specific non-prosecution rates against the criminal complaint rates. Relative to (ref), there is much more variation in the outcomes for ADAs with similar non-prosecution rates. An implication of IV monotonicity is that ADAs with the same prosecution rate must have the same outcomes, since they must be prosecuting the same set of individuals; if the prosecution rates differ by $x$, the outcomes cannot differ by more than $x$ times the range of the support for the outcome. The IV monotonicity test of fll23 tests precisely this implication. Thus, the large outcome variability in (ref) can be seen as giving visual evidence against IV monotonicity, in line with the rejection of the fll23 test in several subsets of the data.

figure[figure omitted — 14,406 chars of source]

(ref) also shows that the relationship between non-prosecution and criminal complaints is negative: the IV estimate correspondingly equals $-28.8$%, albeit the estimate is an order of magnitude noisier than that in ADH22, with standard error $10.0$ (this replicates Table III, col. (4) in ADH23). The overall non-prosecution rate in the data is $20.4$%, and the overall criminal complaint rate is 34.1%.

\paragraph{Results for the quota policy.}

To benchmark our estimates, we again start out by computing the policy effect of a reallocation policy that assigns all cases to the most lenient ADAs\@, and a parametric benchmark that assumes homogeneous treatment effects, which allows for extrapolation of the IV estimates. As shown in (ref), these imply that the quota policy would lead to a reduction in the probability of criminal complaints by about 1--3 percentage points. Notably, however, under the reallocation policy the confidence interval only marginally excludes zero.

We then estimate the policy effect under a specification that only assumes \ac{IV} validity and policy monotonicity. Like the NYC bail judge data, the data is consistent with there being many instrument never-takers. The policy only moves treatment by $9.0$ percentage points, so without further assumptions, we cannot rule out that the individuals marginally released---the policy compliers---are picked from the set of instrument never-takers in an adversarial way. As a result, the estimate of the identified set (row 3 of (ref)) implies trivial bounds for the policy complier treatment effect.

table[table omitted — 1,351 chars of source]

We next consider an additional assumption that can help us tighten these bounds. In contrast to our previous application, imposing bounds on the outcomes for policy compliers as in (ref) is hard to motivate in this setting, since prosecution decisions are based in large part on the strength of evidence against the defendant, which is not directly related to their recidivism risk. Likewise, we do not have external data from panels of prosecutors to help calibrate disagreement probabilities across ADAs.

We therefore pursue a simple approach that directly restricts the counterfactual outcome disparity between ADAs in the bottom 90% and those in the top decile as in (ref) for some bound $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu$. Imposing $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu=0$ amounts to assuming that the policy effect matches that under the reallocation policy. As a simple benchmark, we calibrate $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu$ to the status quo outcome disparity between the two groups of ADAs. Since the criminal complaint rate for those in the top decile is $1.9$ percentage points lower than for those in the bottom 90%, we set $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu=0.019$. This is motivated by the informal discussion in ADH23 who consider a counterfactual where the stricter ADAs are asked to act “more like the most lenient ADAs.” If the policy directive asks them to act this way, one can reasonably expect the outcome disparity to be lower under the counterfactual, so that this value of $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu$ can be interpreted as a conservative bound. As (ref) shows, imposing this bound substantially tightens the estimate of the identified set. As column (2) indicates, the implied treatment effect for policy compliers lies between $0$ and about 130% of the IV estimate.

There is an alternative view of this bound that is useful. In particular, setting $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu$ to the status quo disparity is the maximal value of $\mkern 1.5mu\overline{\mkern-1.5mu\mathsf{OD}\mkern-1.5mu} \mkern 1.5mu$ that one can allow for while still guaranteeing identification of the sign of the policy effect.\footnote{Since ADAs above the quota are assumed not to change their behavior, the policy effect is proportional to the difference between the counterfactual and status quo outcome disparities, $\theta-E[Y]=n_{\mathcal{Q}}/n\cdot (E[Y(D(Z_{\mathcal{Q}}, 1))- Y(D(Z_{\mathcal{Q}^c}, 1))]-(E[Y\mid Z\in\mathcal{Q}]-E[Y\mid Z\in\mathcal{Q}^c]))$, with $n_{\mathcal{Q}}$ denoting the number of cases assigned to the ADAs who are subject to the quota. Since the status quo outcome disparity is negative, assuming that the counterfactual outcome disparity is no larger in magnitude than the status quo disparity implies that the policy effect is weakly negative.} Using (ref), we can thus ask if the implied disagreement probability is reasonable: if so, this suggests that identification of the sign of the treatment effect is possible under reasonable restrictions on the disagreement probability. If we make no restrictions on treatment effect heterogeneity, the implied value of $\mkern 1.5mu\overline{\mkern-1.5muDP\mkern-1.5mu} \mkern 1.5mu$ is 0.032, which corresponds to a signal correlation equal to $\rho=0.998$. This is substantially higher than the corresponding value calibrated from sigstad_monotonicity_2023 ($0.989$), implying that prosecutors in Massachusetts would have to be much more strongly in agreement than the judges in Sigstad's data. While it is possible that agreement rates differ across contexts, this strikes us as a stringent requirement, especially given the rejection of IV monotonicity, which implies disagreement under the status quo. We can, however, learn about the sign of the policy impact under weaker restrictions on disagreement if we also impose some restrictions on treatment effect heterogeneity. If we impose that $\Delta_{TE}$ is at most twice the IV estimate, then we only require $\mkern 1.5mu\overline{\mkern-1.5muDP\mkern-1.5mu} \mkern 1.5mu = 0.112$, which corresponds to $\rho = 0.972$ in the Gaussian signal model. This is lower than the value in Sigstad, and thus seems more reasonable.

\printbibliography