Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
105,317 characters · 17 sections · 79 citation commands
A Sharp Test for the Judge Leniency Design
We propose a novel sharp test to assess the validity of the judge leniency design, which has emerged as a prominent instrumental variable (IV) approach in recent years, particularly in empirical research exploring causal effects within the criminal justice system. This design has proven beneficial in investigating the impacts of various interactions with the legal system, such as pretrial detentions and incarcerations, on subsequent outcomes, including recidivism rates, conviction probabilities, and employment prospects. What sets the judge leniency design apart is its distinctive feature of randomly assigning judges to different cases, with each judge handling a significant number of cases while having discretion over the final decision. The random assignment of judges enhances the credibility of this IV approach and has led to its increasing popularity among researchers kling2006incarceration,di2013criminal,aizer2015juvenile,mueller2015criminal.\footnote{kling2006incarceration exploits randomized judge assignment along with judge propensities to instrument for incarceration length, aiming to investigate the causal impact of incarceration on labor market outcomes.} Importantly, the judge leniency design's random assignment feature extends beyond the context of criminal justice, making it a valuable methodology in diverse research contexts, including medicine, patents and startups, bankruptcy protection, evictions, and access to foster care doyle2015measuring,farre2020patent,dobbie2017consumer,gross2022temporary.\footnote{For example, doyle2015measuring employs the judge leniency design in the medical context to examine the impact of ambulance companies on patients in emergencies, relying on the pseudo-random assignment of ambulance companies to patients. Similarly, dobbie2017consumer uses the leniency of randomly assigned bankruptcy judges as an instrument to study the implications of Chapter 13 bankruptcy protection on future financial events.}
However, in addition to the random assignment, an instrumental variable must adhere to two additional crucial conditions: (i) an exclusion restriction, which means that judges' actions should only influence the treatment and should not have any direct influence on the defendant's future outcomes; and (ii) a monotonicity restriction, which means that judges should consistently exhibit more or less leniency. This means that if a defendant was treated (detained) by one judge, she would always be treated (detained) by a less lenient judge. Trial decisions (treatment) are often multidimensional, including incarceration, fines, community service, sentence length, and others johnson2014judges. These decisions impact future outcomes. Because different judges may have varying attitudes on these decisions, the exclusion restriction can be violated if some of the decisions are unobserved or uncontrolled. Furthermore, abrams2012judges and stevenson2018distortion argue there is considerable heterogeneity in how judges rank defendants when considering various types of offenses. If this heterogeneity is not observed, then it is possible that judges exhibit varying levels of leniency under different circumstances, and the monotonicity assumption would be violated. These observations align with mogstad2019identification, who demonstrate that, in general, monotonicity effectively requires homogeneous choice behavior for economic agents when there are multiple instruments. Therefore, offering a statistical test to evaluate the validity of the judge leniency design becomes a highly relevant empirical question.
In this paper, we characterize the sharp testable implications of the judge leniency design as a set of inequality restrictions on the distribution of the observed data. Our result is novel and contributes to the testable implications derived in the seminal work of heckman2005structural in two ways. First, our implications belong to a tractable subset of the constraints of heckman2005structural and are easier to implement in practice. Second, we establish the sharpness of our testable implication, that is, they possess the unique quality of exploiting all available information within the data distribution that is useful to refute the validity of the judge leniency design.
Numerous efforts have been made to test the judge leniency design in the existing literature. A common approach involves providing separate evidence for the validity of the individual assumptions made in the judge leniency design. For instance, to assess the random assignment of judges, dobbie2018intergenerational examines whether a measure of judge stringency (the instrumental variable) correlates with baseline cases and family characteristics of criminal defendants. Regarding the monotonicity assumption, they test an implication that requires the first-stage estimates to be non-negative for all subsamples. bhuller2018incarceration and norris2021effects employ similar individual testing approaches. Assessing the assumptions individually is effective in empirical scenarios where researchers know which assumption to test and have prior knowledge that other assumptions hold. Our approach adds to the existing body of knowledge by introducing a test that does not depend on prior information. In fact, the three key assumptions may collectively impose certain constraints on the observable data-generating process (DGP), which could not be detected by examining only the testable implications of each assumption in isolation.
Unlike individually testing each assumption, frandsen2023judging proposes a joint test for all assumptions underlying the judge leniency design. Their test leverages the property that, in the judge leniency design, the average outcome at the judge level should exhibit a smooth relationship with the propensity score (or the judge-level treatment probability). It ought to have a bounded slope, where the bounds depend on the limits of the outcome variable’s support. Although frandsen2023judging's testable implication has the desirable property that it assesses all the assumptions simultaneously, we show there is still significant relevant information in the data distribution essential for evaluating the judge leniency design's validity, but not used in frandsen2023judging's testable implication. This difference is also demonstrated by numerical examples and empirical studies reported in (ref).
To the best of our knowledge, our test is the only sharp test available for assessing the validity of the judge leniency design. In other words, our testable implications exhaust all the information in the observed data distribution. As seen in previous methods, non-sharp tests have practical virtue when there is no easily tractable characterization of the sharp testable implications of a model's assumptions. If a non-sharp rejects, it conveys an informative result that the assumptions should be rejected. However, there are also important trade-offs to consider. First, a non-sharp test can have no power against certain violations since it does not consider all possible constraints on the data distribution. Second, different non-sharp tests can lead to discordant empirical results and potentially misleading interpretations of the estimand of interest li2024discordant. For instance, two different non-sharp tests may produce conflicting results because they consider different aspects of the observed data distribution. Our sharp test addresses both issues as it is a consistent test built upon sharp testable implications and, therefore, a useful complement to the existing literature.
We construct valid and consistent semi-nonparametric and semiparametric tests based on these tractable testable implications. Our asymptotic tests support a diverse range of data structures. For example, we can apply our tests to the empirical context in which each judge handles a large number of defendants, and the number of judges can be either large or small (as in our empirical application). As we will further elaborate in (ref), our asymptotic tests are also applicable when the number of defendants for each judge is small, as long as the data regime permits a root-n estimation of the propensity score.
We also provide an easy-to-compute finite sample test for cases involving a small number of judges and a small number of defendants per judge. To the best of our knowledge, ours and the finite sample test of frandsen2023judging are the only finite sample specification tests in the judge leniency design literature. Like frandsen2023judging, our finite sample test also focuses on binary outcomes. Unlike frandsen2023judging's test, which ensures the finite sample validity by computing a “least favorable p-value" via a high-dimensional nonlinear optimization routine, we use Bonferroni correction. The computation for our test is very light and requires little more than simulating Bernoulli random variables. Therefore, it serves as a useful addition to the existing finite sample tests.
As a potential alternative to the existing non-sharp tests, one may consider testing the validity of the judge leniency design employing some of the existing sharp tests developed for the Local Average Treatment Effect (LATE) framework, i.e., kitagawa2015, huber2015testing, and MW2017. However, it is worth noting that these tests may over-reject since they are based on a priori direction in the monotonicity assumption and are not directly applicable in the context of judge leniency design. For instance, in the judge leniency design, the number of judges can be quite large, and in some cases, it might even be infinite, especially when judges' types are continuous. In such scenarios, the number of potential directions to consider becomes large, possibly infinite. Imposing a specific ex-ante direction in the judge leniency design is therefore overly restrictive, and considering all possible directions might be impractical or impossible. Furthermore, imposing an incorrect a priori direction bears an additional risk of model misspecification. These issues highlight the need for a more flexible testing approach, like the one proposed in this paper, which is free from making overly restrictive assumptions on the direction of monotonicity.
While our test is primarily motivated by testing judge leniency designs, it can also be applied to assess the identifying assumptions in a general Marginal Treatment Effect framework with continuous or discrete instrument variables, which has been applied to various empirical settings. See carneiro2011estimating,kowalski2016doing, Brinch2017, among many others. In the context of judge leniency designs, this also means that our test does not require observing a judge's identity and accommodates continuous judge types. Finally, motivated by mogstad2019identification, we propose to relax monotonicity and exclusion assumptions to partial monotonicity and partial exclusion, respectively, when our test rejects the null hypothesis.
We organize the rest of the paper as follows. (ref) presents the analytical framework and the sharp testable implications of the judge leniency design. (ref) presents the testing procedures. In (ref), we show the results of the simulations and discuss our empirical illustration. In (ref), we explore approaches to salvage the judge leniency design when its sharp testable implications are violated. The last section concludes the paper, and the proofs are collected in the online supplementary materials.
We adopt the potential outcomes framework. Let the observed treatment indicator be $D\in \{0,1\}$. For example, in the judge leniency design, the unit of observation is defendants. Hence, $D=1$ indicates that a defendant is incarcerated. Let $Z\in \mathcal Z \subseteq \mathbb R^{d_z}$ be the type of the judge assigned to the defendant. $Y_{d}(z)\in\mathcal Y \subseteq \mathbb R$ denotes the potential outcome of interest (e.g., recidivism) when the treatment and the judge's type are externally set to $D=d$, and $Z=z$, respectively. Similarly, $D_z$ denotes the potential treatment when the judge's type is externally set to $Z=z$. Let $Y = Y_{1}(Z)D + Y_{0}(Z)(1-D)$ be the observed outcome. For the moment, we omit observed defendant and case covariates $X$ (such as time and courtroom of the trial) for ease of notation. The identification analysis in this section can be extended by conditioning on $X$. We will also discuss the implementation of our test in the presence of $X$ in (ref).
In our setting, $Z$ can be multidimensional, continuous, discrete, or a combination of both. For example, if there is a group of judges $\mathcal J$, and if their identities are observed, then $Z \in \mathcal J$ can be chosen as the identity of the judge assigned to the defendant. This is the instrumental variable that frandsen2023judging consider. On the other hand, we allow scenarios in which the judge's identity is unobserved but with observed characteristics. In this case, $Z$ may contain a set of continuous or discrete variables, such as the judge's experience, gender, and race.
The literature mainly relies on the following assumptions to evaluate the causal effects of treatment $D$ on outcome $Y$.
A particular feature of the judge leniency design is that judges are usually randomly assigned to different cases, making the random assignment assumption likely to hold in practice. However, (ref) are usually less credible. (ref) means the effect of judges on the potential outcomes must necessarily transit through their effect on treatment assignment. (ref) requires that any defendants treated (incarcerated) by a more lenient judge be also treated if assigned to a less lenient one. heckman2005structural refers to the monotonicity assumption as a uniformity condition since it restricts that the treatment on all the defendants must vary in a uniform direction when externally assigned to another judge. vytlacil2002independence provides an equivalent characterization of the monotonicity assumption, which can be stated as follows:
Under (ref), we can rewrite the threshold crossing model without loss of generality as follows: \[ D = 1\left\{F_{U}(\nu(Z))\geq F_{U}(U)\right\} \equiv 1\left\{P(Z)\geq V\right\}, \] where $F_{U}(\cdot)$ is the distribution function of $U$, $P(\cdot)\equiv F_{U}(\nu(\cdot))$ is identified from the observed variables $(D,Z)$ by $P(z) = \mathbb P(D=1|Z=z)$, and $V\equiv F_{U}(U)\sim Uniform[0,1]$. Hereafter, we will write $P(Z)$ as $P$ when it causes no confusion. Let $\mathcal P \subseteq [0,1]$ denote the support of $P(Z)$. It is worth noting that the STC does not impose a priori direction in $z$ in the monotonicity condition since (ref) is equivalent to (ref) vytlacil2002independence. Under (ref), the judge leniency design model can be equivalently written as:
(ref) (equivalently (ref)) impose some restrictions on the joint distribution of the observed variables $(Y,D,P(Z))$, which we will characterize in (ref). But before stating the theorem, we will discuss the intuition of the testable implications. Let $g: \mathcal Y\rightarrow \mathbb R^{+}$ be a nonnegative real integrable function such that $\mathbb E \vert g(Y_d)\vert < \infty$. Taking $d=0$ as an illustration. For any pair $(p,p') \in \mathcal P \times \mathcal P $ such that $p \leq p'$, we have:
The first and fourth equalities hold by (ref) (STC) and (ref) (exclusion); the second and third equalities hold because of (ref) (random assignment), and the inequality holds because $p\leq p'$. Intuitively, under the assumptions of the judge leniency design, if a defendant is released by judge $p'$, then he/she would necessarily be released by judge $p$ since judge $p$ is more lenient than judge $p'$. On the other hand, there can exist a set of defendants who were released by a type $p$ judge, but not by a type $p'$ judge: a group of “compliers". Because $g(Y_0)$ is nonnegative, the average $g(Y_0)$ for this group of compliers is also nonnegative, delivering the inequality we see from the displayed equation above. The discussion is formalized in the following theorem.
The proof of (ref) is collected in (ref). The testable implications in (ref)(i) are a subset of the implications previously derived in heckman2005structural, who show for any non-negative integrable function, i.e. $g(\cdot): \mathcal Y \rightarrow \mathbb R^{+}$, $\mathbb E [g(Y)D \vert P=p]$ and $-\mathbb E [g(Y)(1-D) \vert P=p]$ are non-decreasing in $p$ under (ref). The contribution of (ref)-(i) is that it shows we do not need to visit every single non-negative measurable function. It is sufficient to restrict our attention to a tractable subclass of these functions to screen all possible observable violations. This tractable characterization provides a basis for constructing a formal statistical test to verify the validity of the assumptions.\footnote{ We note that use the the half-interval class $g(Y)=1\{Y\leq y\}, y\in\mathcal Y$ will result in loss of power. To see this, suppose the support is finite, that is, $\mathcal Y=\{y_1,y_2,\cdots, y_K\}$, then it is without loss of information to consider the class of singletons: $g(Y)=1\{Y=y_k\}$, $k=1,2,\cdots, K$. However, if one considers $g(Y)=1\{Y\leq y_k\}$, then it is possible that both $\mathbb P(Y\leq y_1, D=1 \vert P=p)$ and $\mathbb P(Y\leq y_2, D=1 \vert P=p)$ are non-decreasing function, but $\mathbb P(Y=y_2, D=1 \vert P=p)$ is not.}
The second part of (ref) is new, and it shows that the testable implications in (ref)(i) are the most informative way to detect all observable violations of the random assignment, the exclusion restriction, and the monotonicity assumption (without an ex-ante imposed direction). These testable implications cannot be strengthened without making additional assumptions. Various tests or testable implications are used in the literature to screen violations of the judge leniency design assumptions; for instance, dobbie2018intergenerational,bhuller2018incarceration,norris2021effects,frandsen2023judging. However, to the best of our knowledge, only (ref) provides sharp testable implications without imposing an a priori direction in the monotonicity assumption.
Tests based on sharp testable implications have empirical virtue. In practice, one may use tests developed from non-sharp testable implications for the sake of traceability. However, as recently discussed in li2024discordant, non-sharp tests can lead to discordant empirical results and misleading interpretations of the estimand of interest. It is possible that for the same data, two different non-sharp tests may generate contradictory results, as they may use different sets of information from the same observed DGP to screen violations of the model assumptions. Thus, the conclusion may largely depend on which test the empirical researcher implements.
Moreover, after implementing a specification test and obtaining a non-rejection result, one often proceeds and provides a causal interpretation of the estimand. For example, in judge leniency designs, the 2SLS or Local IV (LIV) estimand is interpreted as the LATE or MTE, respectively. However, since a non-sharp test only uses part of the observable information in the data and fails to reject the model when it is misspecified, we must be cautious about interpreting the 2SLS or the LIV estimand as identifying the LATE/MTE solely based on the result of a non-sharp test. Therefore, using a sharp test must be viewed not only as a theoretical exercise, but also as having an important empirical relevance. A sharp test provides the most informative way to detect all observable violations of a given model's assumptions and is more robust to possible misleading interpretations and discordant results.
Inspired by heckman2005structural, kitagawa2015 and MW2017 derive a set of sharp testable implications assuming an a priori direction in the monotonicity assumption. When judges' types are binary, i.e. $ Z\in \mathcal \{0,1\}$, there are only two potential directions, so it is not restrictive to assume the direction of the monotonicity. However, when the cardinality of the judges' types is large (or even infinite when the judges' types are continuous), imposing a specific ex-ante direction is extremely restrictive because the number of possible directions to consider can be rather large (or even infinite). One could implement their test by visiting all the possible directions, but this can be cumbersome or even computationally impossible if $Z$ takes many values.
One significant difference between the testable implication of kitagawa2015 and MW2017 and ours is we do not assume a prior direction. To illustrate this point, suppose $\mathcal Z=\{z_1,...,z_K\}$ and suppose we assume one of the $K!$ potential directions as: $$D_{z_{K}}\geq D_{z_{K-1}} \geq...\geq D_{z_1}$$ meaning that type $z_{K}$ judge is less lenient than type $z_{K-1}$ judge, which, in turn, is less lenient than $z_{K-2}$, $z_{K-3}$, $\cdots, z_{1}$ judge. Given this imposed ordering, (ref) imply the following testable implications studied in sun2023instrument:
A key point to note is that the above implications restrict $F_{Y,D|Z}(y,d|z)$ while the testable implications in (ref)(i) instead restrict $F_{Y,D|P}(y,d|p)$. In the first case, the induced direction of inequalities is with respect to the observed judge type $Z$, while in our case, the inequalities are with respect to the propensity score $P$, which is obtained without imposing a prior direction. Also, noteworthy is that if one takes $y=-\infty$ and $y'=\infty$, the testable implications in (ref)(i) no longer have any empirical content. But, the testable implications with an ex-ante monotonicity direction still restrict the propensity scores and the judges' types, i.e., $P$, and $Z$, such that
Therefore, implementing the testing approaches of kitagawa2015 and MW2017 may reject the judge leniency design assumptions even if (ref) hold, but just the ex-ante imposed direction of monotonicity is wrong.
FLL proposes a set of testable implications for (ref). Their testable implication has sound features of not relying on the ex-ante specified direction of monotonicity and assessing all the assumptions jointly. Their testable implication, however, is not sharp and can fail to screen some {\color{blue} non-negligible} observable violations of the judge leniency design. To see this, consider any integrable function $g(\cdot): \mathcal Y \rightarrow \mathbb R$, and let $p\neq p' \in \mathcal P$. Under (ref), we can derive the following equality:
If we denote by $L_g$ and $U_g$ the known lower bound and upper bound of the support of $g(Y)$, the latter equality implies:
where the inequality in (ref) is the main testable implication used by FLL (see Theorem 1 and Equation (2) therein) to implement their test. However, under (ref), we should also have:
where those two latter equalities lead to the following observable restrictions:
One can easily observe that the testable restrictions in (ref) and (ref) could be violated, whereas the restriction used by FLL, i.e. inequality (ref) still holds. Hence, implementing FLL's statistical testing procedure based on inequalities (ref) or (ref) could provide a different result compared to their test based on inequality (ref) alone. These discordant implications confirm the concern about developing a statistical test based on non-sharp restrictions. (ref) provides a concrete numerical example.
The intuition behind (ref) is not pathological and is reflected in the derivation in (ref). Because $\mathbb E[Y|P=p]=\mathbb E[Y_1D|P=p]+\mathbb E[Y_0(1-D)|P=p]$, it is possible the violations on the $\mathbb E[Y_1D|P=p]$ and $\mathbb E[Y_0(1-D)|P=p]$ “cancel” out. As a consequence, the quantity $\mathbb E[Y|P=p]$ provides no power to detect violations in these cases.
Another evident reason why FLL's implications cannot exhaust all violations of the judge leniency design is that they only focus on $g(Y)=Y$, whereas the inequality in (ref) should hold for any integrable function $g$ and for any pair $p\neq p' \in \mathcal P$. $g(Y)=Y$ is not a sufficient class of functions to screen all violations of the model.
Finally, we note our testable implications in (ref) do not rely on the known support of $g(Y)$, whereas to test inequality ((ref)), one needs to know the bounds of the support $(U_g,L_g)$. If the support of $g(Y)$ is unbounded, i.e., $U_g=+\infty$ and $L_g=-\infty$, then the testable implication in ((ref)) holds trivially and FLL's test does not have any power in detecting violations to the identification assumptions.
In the next section, we propose a testing procedure based on the sharp testable implications of (ref). We will show that in large samples, our test is consistent against all the violations of our testable implication and is, therefore, more powerful asymptotically than the existing ones.
In this section, we construct tests based on (ref). For the defendant $i\in\{1,2,\cdots, n$\}, researchers observed a vector $(Y_i,D_i,Z_i,X_i)$, where $Y_i$, $D_i$, $Z_i$, and $X_i$ represent his/her observed outcome, observed treatment status, the vector of characteristics of the judge that $i$ was assigned to, and the vector of additional control variables, respectively.
We first present our baseline semi-nonparametric test in (ref) without the presence of control variables $X_i$. For this test, we make no functional form or distributional assumptions about potential outcomes. We do need to estimate the propensity score $P(z)\equiv \mathbb P(D_i=1|Z_i=z)$ first, for which our procedure can accommodate different data scenarios. If $Z_i$ contains continuous variables, we follow the common practice in the literature to employ a parametric model so that $P(z)=P(z,\theta_0)$ for all $z\in\mathcal{Z}$ and for a finite-dimensional parameter vector $\theta_0\in\Theta$. Popular choices include the Probit or Logit model with a linear index $z'\theta_0$ carneiro2011estimating,kowalski2016doing.\footnote{When $Z$ is continuous, the rejection result of our semi-nonparametric test can be interpreted as rejecting the joint assumption of the judge leniency design and the parametric form imposed on the propensity score. In our simulation studies, we always keep the propensity score correctly specified. In these studies, therefore, the rejection shows the power of our test to reject false judge leniency assumptions.} When $Z_i$ only contains discrete variables, such as judge's gender, we can estimate $\mathbb P(D_i=1|Z_i=z)$ by the sample averages of $D$ conditioning on each possible value of $z\in\mathcal Z$. In this case, our test is indeed nonparametric.
We should emphasize that, in both cases above, we do not require any knowledge of the identity of judges, nor do we need the number of defendants handled by each judge to diverge to infinity. For example, when $Z$ is gender, we only need the number of defendants for each judge's gender to go to infinity. This can happen when the number of judges is large, but each judge handles a finite number (or even one) of defendants. Therefore, our test can also be applied to other empirical contexts than the judge leniency design. There is another scenario in which $Z_i$ is the judge's identity. Suppose the number of defendants handled by each judge is large, as in our empirical application. In this case, we can also consistently estimate judge $j$'s propensity score $\mathbb P (D=1|Z_i=j)$ by the sample frequency estimator for each judge $j$, regardless of whether the number of judges is small or large.
In practice, the number of defendants that a judge handles can be small. In this case, one can not estimate $\mathbb P (D=1|Z_i=j)$ consistently without additional assumptions and conventional inference methods can be invalid. This phenomenon has received attention from the literature, see discussions in jochmans2023many, ren2024extrapolating, sithole2024locally, and yap2024inference. These papers, however, focus on inference on the parameters instead of testing model specification. To account for data scenarios with small numbers of defendants per judge, we design a test that does not require a consistent estimator for the propensity score and controls the size at any sample sizes for the case of a binary outcome. To the best of our knowledge, this test and FLL's finite sample test are the only ones for testing judge leniency design specification with finite samples, and both focus on binary outcome variables. Our test uses upper bounds of the null distribution to calculate the critical values, and hence is very easy to implement. It only requires simulating Bernoulli random variables, and no nonlinear optimization is involved. For the purpose of exposition, we collect the finite sample test in (ref) and focus on the cases in which the propensity score can be consistently estimated in this section.
In practice, researchers may observe a set of defendant and case covariates $X$ and assume the randomization and monotonicity hold conditioning on $X$ (see (ref) below). In the presence of covariates, researchers can use the semi-nonparametric test introduced in (ref) when the dimension of covariates is small or the number of support points in $\mathcal X$ is not large; please see (ref) below. In other cases, the semi-nonparametric test may encounter challenges associated with the curse of dimensionality. To address this concern, we introduce an alternative semiparametric test designed to accommodate situations with a large (but fixed) number of covariates in (ref).
For the convenience of the exposition, we restate the testable implications as the null hypothesis $H_0$. That is, for all $p_1\geq p_2 \text{ with } p_1,p_2\in \mathcal P \text{ and all } y,y' \in \mathcal Y$,
The alternative hypothesis $H_1$ is then inequality ((ref)) or ((ref)) fails to hold for some $(p_1,p_2)$ and $(y,y')$. Without loss of generality, we assume the support of $Y$ is $[0,1]$.\footnote{We can always apply a transformation to ensure the support of $Y$ is $[0,1]$. If $Y$ has a finite support $[a,b]$, we can apply an affine transformation $\tilde Y=(Y-a)/(b-a)$. If $Y$'s support is the whole real line, we can apply standard normal CDF after rescaling and re-centering: $\tilde Y= \Phi\left(\frac{Y-\bar Y}{\hat{std}(Y)}\right)$, where $\bar Y$ is the sample average and $\hat{std}(Y)$ is the sample standard deviation.} Testing inequalities ((ref)) and ((ref)) involves two features; first, it is a set of inequality restrictions defined on conditional moments where the conditioning variable is possibly continuous. We deal with the first difficulty by employing the method of hsu2019testing to transform them into an equivalent set of restrictions on unconditional moments. The second feature is that the conditioning variable $P$ is not directly observed from the data. We derive the new influence functions and show that the first-stage estimation error is properly accounted for.
To be more specific, we define a collection of functions $\{\nu_d(\ell):\ell\in \mathcal L, d=0,1\}$ as follows:
and
where the index $\ell \in \mathcal L$ is defined as
Then, following the same calculation as in hsu2019testing, we can formulate the null hypothesis in inequalities ((ref)) and ((ref)) as the following:
against the alternative hypothesis $H_1$ that inequality ((ref)) fails to hold for some $\ell\in \mathcal L$ and for $d=0$ or $d=1$. Consequently, testing the original sharp implication in (ref) is equivalent to testing the set of inequalities indexed by $\ell\in\mathcal L$, a class of cubes. There is no loss of information for such transformation AndrewsShi2013. Under $H_0$, we expect to see $T \equiv \sum_{d=0,1} \sum_{\ell\in\mathcal L} \max\{\nu_d(\ell),0\}^2 \Omega(\ell) = 0$, where $\Omega(\cdot)$ is a positive weighting function. On the other hand, $T>0$ under $H_1$. Our test statistics are based on the appropriately rescaled and standardized sample analog of $T$.
In the expression of $\nu_d(\ell)$, the propensity score $P(Z_i)$ is unknown, but can be replaced by its root-n consistent estimate $\hat P_i$. When we estimate the propensity score by a parametric model, we denote it as $\hat P_i \equiv P(Z_i,\hat\theta)$, where $\hat\theta$ is the MLE. When $Z_i$ is the judge's identity and the number of defendants for each judge is large, we simply use the frequency estimator $\hat P_i=\frac{\sum_{j=1}^n D_j1\{Z_j=Z_i\}}{\sum_{j=1}^n 1\{Z_j=Z_i\}}$. (ref) below summarizes the semi-nonparametric test's implementation procedure. Please see (ref) for detailed equations and expressions.
(ref) shows that the test $\phi_n$ has its size controlled asymptotically and is consistent. The proof for (ref) is collected in (ref) of the online supplementary material. We also list all the technical assumptions, such as conditions that ensure the first-stage estimator converges at a sufficiently fast rate, in that section for the sake of exposition.
In this section, we introduce a semiparametric test in the presence of covariates $X$. We begin by introducing the following assumptions.
When (ref) hold, the testable implications can be written as follows. For all $x\in\mathcal X$, $p_1,p_2\in\mathcal P$ and $p_1\geq p_2$, and all $y,y'\in\mathcal Y$
When the dimension of $X$ is high, an alternative approach is to include the covariates parametrically, as in carr2021testing, which we state below:
carr2021testing show if (ref) is strengthened to (ref), then the testable implications in ((ref)) and ((ref)) can be characterized as
for $y,y'\in \mathcal Y$, and \[ \tilde Y= D(U_1+\alpha_1) + (1-D)(U_0+\alpha_0) = D(Y_1 - X'\beta_1)+(1-D)(Y_0-X'\beta_0) = Y - X'(D\beta_1 +(1-D)\beta_0). \] The advantage of using ((ref)) and ((ref)) is that both inequalities are only conditional on the scalar-valued propensity score. The effect of covariates has been filtered out by constructing a new outcome variable $\tilde Y$. (ref) is a common assumption made in the literature for estimating the MTE, see for instance CarneiroLee2009,carneiro2010evaluating,kowalski2016doing. Nevertheless, we do acknowledge it is subject to the potential risk of model mis-specification. Under the null hypothesis of the model being correctly specified, parameters $\beta_0$ and $\beta_1$ can be estimated by partial linear regression of $Y$ on $X$ and propensity score $P$ separately for the sample of $D=1$ and $D=0$: \[ \mathbb E[Y|X=x,P=p,D=d]=x'\beta_d + K_d(p),\quad d\in\{0,1\}, \] where $K_d(p) = \mathbb E[\alpha_d+U_d|X=x,D=d,P=p]$ only depends on $p$ under (ref)-(ii). The following algorithm summarizes the steps for implementation.
In this subsection, we provide two sets of simulations to assess the size and power properties of our sharp test under various DGPs in finite samples. Throughout this section, we ran $1000$ replications for each simulation design, and the bootstrap sample size is chosen to be $B=800$. We set $a_n=0.15\ln n$ and $B_n={0.85\ln n}/{\ln\ln n}$, as in hsu2019testing. We choose $Q_P=5$ and $Q_Y=5$ (for continuous $Y$) or $Q_Y=2$ (for binary $Y$). We set the infinitesimal constant $\eta=10^{-6}$ and the constant $\epsilon=10^{-6}$ (see the definition of $\hat\sigma^2_{d,\epsilon}(\ell)$ in (ref)-4).
The first set of simulations is based on a DGP introduced in FLL (online appendix, page 22). In this set of simulations, we mimic the random assignment of $n$ defendants to a pool of $J$ judges, ensuring an equitable distribution of $\frac{n}{J}$ defendants to each judge. As in FLL, the severity probability of each judge $j$ is set as follows: $$ p_{j} = p_{a} + \frac{ j-1}{J-1}(1-p_{a} - p_{n}) $$ Here, $p_{a}$ and $p_{n}$ stand for the fraction of always and never treated defendants, respectively. FLL consider a binary outcome model where the outcome $Y \in \{0,1\}$ satisfies the following condition: $$ \mathbb{E} \left[ Y \mid p_{j} \right] = \frac{1 - (1-\lambda)(p_n + p_a)}{1 - (p_n + p_a)} p_j - \frac{\lambda}{1-(p_n + p_a)} p_a. $$ The parameter $\lambda$ dictates the extent of deviation from the exclusion restriction assumption. When $\lambda = 0$, there is no violation of the judge leniency design assumptions. Consequently, for $\lambda = 0$, the simulations aim at assessing the size property of the two different tests. On the other hand, $\lambda > 0$ signifies a departure from the judge leniency design assumption, with higher (absolute) values indicating a more pronounced deviation. Like in the original paper, we adopt the parametrization for the fraction of always and never treated $p_n = p_a = 0.2$. Meanwhile, we vary the value of $\lambda$ within the range of 0 to 1. Note that the parameter $\lambda$ directly governs the shape of the function $\mathbb E[Y|P=p]$. The nonzero value of $\lambda$ can potentially be generated by violations of one of the three assumptions (or their combinations).
(ref) visually illustrates our testable implications of the judge leniency design for the specific function $g(Y) = 1\{0<Y\leq 1\} = Y$ (because $Y$ is binary). The left and right panels of the figure, respectively, depict $ \mathbb E [-Y(1-D) \vert P=p]$ and $ \mathbb E [YD \vert P=p]$. These population quantities are approximated by a large number of defendants (1 million) for each judge. Intuitively, it is expected that $ \mathbb E [YD \vert P=p]$ and $\mathbb E [-Y(1-D) \vert P=p]$ should be non-decreasing when the judge leniency design holds. When the exclusion restriction holds, as shown in both figures with $\lambda=0$, $ \mathbb E [YD \vert P=p]$ and $ \mathbb E [-Y(1-D) \vert P=p]$ behave as expected. However, for a violation of the exclusion restriction ($\lambda = 0.4$ or $\lambda=0.8$), despite that $\mathbb E[YD|P=p]$ remains to be increasing, the other function $\mathbb E [-Y(1-D) \vert P=p]$ decreases for higher values of the propensity score. This discrepancy starkly contrasts with the implications of the judge leniency design assumptions.
In (ref)-(a), we report the size property for our sharp test and FFL's test at 5% significance level (when $\lambda = 0$). The simulation designs involve twenty judges and varying sample sizes, ranging from 500 defendants (equivalent to 50 defendants per judge) to 5500 defendants (equivalent to 550 defendants per judge). The plot reveals that both tests control size well in the aforementioned DGP. Specifically, it is evident from the graph that the rejection rate of our sharp test is controlled by and close to the nominal level of 5%. Conversely, the nonparametric test proposed in FLL consistently yields rejection rates close to zero when setting the tuning parameter $K=1$.\footnote{Recall the outcome variable is binary; hence, the largest possible absolute value for the treatment effect is $1$. These results correspond to Figures 9 and 10 in the online appendix of FLL, where the rejection probabilities are nearly zero for various sample sizes when $\lambda=0$.}
FLL discuss how one can improve the power of their testing methodology by considering more stringent upper bounds on the largest possible treatment effects (i.e., using a smaller value of $K$). For instance, in their empirical application of a binary outcome model--where the maximum treatment effect is set at 1--they advocate exploring smaller permissible maximum treatment effect values. However, if $K$ is set to be too small, then FLL's test can have server size distortion. Indeed, (ref)-(b) graphically represents this situation by plotting the rejection rate associated with FLL's nonparametric test under two additional cases: when the maximum allowable treatment effect $K$ is set at 0.8 and 0.4, respectively. The striking observation is that the conclusions drawn from these scenarios can be misleading, as they suggest an excessive over-rejection of the assumptions even when those assumptions are indeed satisfied. For example, if one sets $K=0.4$, then the rejection rate is always $100\%$ whenever the sample size is greater or equal to $1000$. As a matter of fact, the rejection we observe from (ref)-(b) reflects that the ad-hoc imposed magnitude of the treatment effect is not correct, but the underlying exclusion restriction holds. Our test is immune to this problem since it does not require pre-specifying the magnitude order of the unknown treatment effect.
To assess and compare the power property of the two nonparametric tests, (ref)-(c) plots the rejection rate as a function of $\lambda $ for 10 judges and 1000 defendants (100 defendants per judge). The solid line is the rejection rate of the FLL test, which is nearly the same as what is plotted in FLL (Appendix, Figure 10). The rejection rate achieved by our sharp test consistently surpasses that of the FLL test across the entire spectrum of exclusion restriction violations, as indicated by varying degrees of $\lambda$. As shown, the power improvement can be substantial.
The second set of simulations aims to show the performance of our test in detecting violations of the judge leniency design when the outcome is continuous and unbounded. Let $(U_0,U_1,U,Z^*)\sim N(\pmb{\mu},\Sigma)$, where $\pmb{\mu} =(\mu_0 ,\mu_1 ,\mu_U ,\mu_Z)'$ is a vector of means, and $\Sigma $ is a covariance matrix. For generic random variables $A$ and $B$, let $\sigma_A^2$ be the variance of $A$ and $\rho_{A,B}$ be the correlation coefficient between $A$ and $B$. In this design, we set $\sigma_A=1$ for all $A\in\{U_1,U_0, U, Z^*\}$. We let $\rho_{U_0,U}=-0.5$, $\rho_{U_1,U}=0.5$, $\rho_{U,Z}=0$, $\rho_{U_1,U_0}=0$, $\rho_{U_1,Z}=\delta_1$, and $\rho_{U_0,Z}=\delta_1$. To create discrete judges or IV, we set \[ Z=F_{Z^*}^{-1}\left(\frac{\ell(Z^*)}{L}\right), \quad \ell(Z^*) = \mathop{\mathrm{argmin}}_{\ell=1,2,\cdots,L-1} \left|F_{Z^*}(Z^*) - \frac{\ell}{L}\right|. \] That is, we divide the support of $Z^*$ by $L$ equal-probability intervals and concentrate the mass over each interval to its nearest cutoff points. Let the potential outcomes and treatment assignment be
and \[ Y_{d}(z) = \alpha_d+X\beta_d + \delta_3 z + U_d,\quad Y_d =\sum_{z\in \mathcal Z} Y_{d}(z) 1\{Z=z\}. \] where $X\sim N(0,1)$ is independent of all the other random variables. We let $\nu(x,z)=z$ and set $\alpha_0=0$, and $\alpha_1=1$.The $\delta$ parameters, however, are set to be different values to capture different violations of the judge leniency design. More specifically,
(ref) plots $\mathbb E[g(Y)D|P=p]$ as a function of $p$ when $g(Y)=1\{Y\geq 0.5\}$ and 20 judges for a simple illustration. The graphs were simulated with a large sample size (over three million) and approximated the population quantity. The function is non-decreasing when all assumptions are met, as shown in the upper-left panel. In contrast, $\mathbb E [g(Y)D \vert P=p]$ deviates from the expected pattern when the judge leniency design assumptions are violated in different ways.
(ref), on the other hand, plots the testable implication used in FLL. The left side panels plot $\mathbb E[Y|P=p]$ for each of the $p\in\{p_1,p_2,\cdots, p_{20}\}$ (sorted in increasing order) for each of the four designs. The right panels plot the “numerical derivative” of the form $\frac{\mathbb E[Y|P=p_j]-\mathbb E[Y|P=p_{j-1}]}{p_{j}-p_{j-1}}$ against $\{p_2,\cdots, p_{20}\}$. The FLL testable implications require that the curves in the right-hand side panels be bounded between $[-K, K]$, where $K$ again is the difference between the upper and lower bounds of the support. Note that in this example, the outcomes have unbounded support and, therefore, $K=+\infty$. If we choose $K$ as a large number, then it is apparent that all four designs satisfy FLL's testable implication. Hence, we expect no rejection for designs 2-4, albeit they violate the identifying assumptions unless $K$ is set to be relatively small.
We proceed by implementing our sharp test and FLL's nonparametric test. This comparison is conducted across various parameter values and sample sizes. Specifically, we consider a size design (Size $\delta_1=\delta_2=\delta_3=0$), violation of independence (Power1 $\delta_1=-0.5,\delta_2=\delta_3=0$), violation of monotonicity (Power2 $\delta_2\neq 0, \delta_1=\delta_3=0$), and violation of exclusion (Power3 $\delta_3=-0.5, \delta_1=\delta_2=0$). For each violation, we consider situations with covariates ($\beta_1=\beta_0=1$) or without covariates ($\beta_1=\beta_0=0$ ). When there are covariates, we use carr2021testing's method to control for covariates, as discussed in the previous section. To implement FLL's test, we set $K$ to be the difference between sample maximum ($y_{max}$) and minimum $(y_{min})$: $\Delta_y\equiv y_{max} - y_{min}$. We also consider $K=\frac{\Delta_y}{8}$ and $K=\frac{\Delta_y}{16}$. The results are summarized in (ref).
Regarding the size property, all tests control the size except the FLL test when $K$ is set to be very small. Our test and the FLL test with $K=\Delta_y$ and $K=\frac{\Delta_y}{8}$ are conservative. When one sets $K=\frac{\Delta_y}{16}$, the rejection probability of FLL's test increases quickly even when all the assumptions are satisfied (the first design). This is unsurprising because a very small $K$ essentially introduced another severe misspecification to the model. However, when examining the power property of the three tests, we see clearly that our test outperforms FLL's tests by a large margin. The proposed sharp test has enough power to detect the violation of any of the three assumptions (independence, exclusion, and monotonicity). In particular, the rejection rates for our sharp test quickly increase with sample size, surpassing 90% for all cases when the sample size reaches $2000$ (or $100$ cases per judge). Note that in this simulation, the parametric form of the propensity score is correctly specified (except for Power2 when monotonicity is violated); hence, the high power of our test is not because of misspecification of $P(z,\theta_0)$. In contrast, FLL's test has low power performance unless we set $K$ as a small value, which, on the other hand, induces size distortion.
(ref) further examines how the rejection frequency varies as the “magnitude of violation varies" for independence and exclusion. For this exercise, we focus on sample size $n=1000$ (50 cases per judge). Not surprisingly, when the magnitude of the violation is small, all tests have low power. However, as the degree of violation increases, the power of our sharp test rises quickly, even quicker than the FLL's nonparametric test with $K=\frac{\Delta_y}{16}$. On the other hand, when $K=\Delta_y$, FLL's nonparametric test does not reject even if the degree of violation is substantial. Again, this table demonstrates that sharp testable implications are desirable in practice.
In this subsection, we employ our test to assess the validity of the judge leniency designs using data from stevenson2018distortion; see also cunningham2021causal, who studies the impact of pretrial detention on conviction. Using Philadelphia court records and leveraging the varying leniency of bail magistrates as an instrumental variable, the author discovers that pretrial detention leads to a 13% increase in the likelihood of conviction.
In the Philadelphia court system, following an arrest, individuals are taken to one of seven city police stations for a video conference interview by Pretrial Services, which assesses risk factors and financial details for public defense eligibility. Utilizing this information, Pretrial Services assigns arrestees to a bail recommendation grid. Bail hearings, conducted by magistrates every four hours via video conference, involve a brief process where charges are explained, next court appearances are specified, eligibility for a court-appointed defense attorney is determined, and bail amounts are set based on arrest details, interviews, criminal history, guidelines, and input from representatives. Magistrates hold broad authority to assign bail, which can fall into categories such as release without payment, cash bail with a 10% deposit, or no bail at all.
stevenson2018distortion's research design leverages the varying magistrate tendencies to assign affordable bail as an instrument to study detention's impact on case outcomes. To answer the research questions, the author utilizes data from the court records of the Pennsylvania Unified Judicial System, obtained through web scraping of public records in PDF format, which are then transformed for statistical analysis. The dataset encompasses arrests in Philadelphia, where charges were filed between September 13, 2006, and February 18, 2013. The final dataset includes 331,971 cases and eight randomly assigned judges, with each observation pertaining to a specific criminal case. As noted in stevenson2018distortion, the shift-rotation system at the Philadelphia court forms the basis for such randomness.
In what follows, we focus on the aggregate dataset (all criminal cases together) and four primary categories of criminal cases in the data: aggressive assault, robbery, drug sale, and drug possession. These four criminal cases we consider in isolation constitute 43% of the total cases. In (ref), we present two scatter plots for each crime category: $\{(p_j, \mathbb E[YD|P=p_j])\}_{j=1}^8$ and $\{(p_j, \mathbb -E[Y(1-D)|P=p_j])\}_{j=1}^8$, along with a fitted polynomial to illustrate whether the anticipated implications of the judge leniency design framework are satisfied for the considered categories of criminal cases. The graphs indicate $\mathbb E [YD \vert P=p]$ and $\mathbb E [-Y(1-D) \vert P=p]$ are most likely to be non-decreasing for the aggressive assault case.\footnote{Note all the outcome variables are binary. Therefore, the close interval we use for the (ref) is $1\{0<Y\leq 1\}$, which equals to $Y$.} The non-decreasing shape of the functions is unclear for the other types of criminal categories. Although this graphical representation does not constitute a formal test, it offers an intuitive insight. Specifically, it suggests that the assumptions are the least likely to be violated in the aggressive assault case, while the drug possession case shows the highest likelihood of violating the assumptions of the judge leniency design.
We observe a relatively large set of covariates, including fixed effects for year, month, and day of the week. We, consequently, implement the semi-parametric version of our test. For comparison, we also implement FLL's nonparametric and semi-parametric tests. The results of the three tests are presented in (ref) for both the aggregate dataset and separately for each of the four crime categories aforementioned. The nonparametric test introduced by FLL indicates the validity of judge leniency design cannot be rejected either conditioning on each crime category or the aggregate data set at 10% level, despite that the shape of $\mathbb E [YD \vert P=p]$ and $\mathbb E [-Y(1-D) \vert P=p]$ for the drug possession type suggests the opposite. In contrast, our novel test yields results that align with expectations. For instance, the sharp test does not indicate a rejection of the validity of the judge leniency design assumptions for the aggressive assault. However, for all three other types of offenses, our test rejects the validity of the judge leniency design. Meanwhile, FLL's semi-parametric test rejects the category of aggregate assault.\footnote{For FLL's semi-parametric test, we fit the regression function $\mathbb E[Y|P=p]$ by B-spline with three knots. The results for other numbers of knots are reported in the appendix. The reported p-value is the “combined p-value” of the fit component and slope component of the test, and we can see from (ref) in the online appendix that the rejection is mostly generated by the fit component.} These results suggest that using the Wald estimand or the MTE approach for those cases will lead to inconsistent estimates of the causal effects of interest.
Finally, we see no evidence to refute the assumptions underpinning the judge leniency design when applying our sharp test to the aggregate dataset. This outcome may be influenced by the notably high proportion of aggressive assault cases within the dataset compared to other categories. Our result also ascertains that the exclusion restriction or monotonicity can hold for some crime categories but not others, suggesting that controlling the crime type is important in practice.
The rejection of the sharp test means that the judge's leniency design assumptions are too stringent for the data. In this case, relaxing some of these assumptions is required to salvage the model. There are different ways to relax a model's assumptions. One way is to maintain the same estimand used in the stringent model and ask under what conditions this estimand can still be interpreted causally. The model relaxation recently entertained by FLL falls into this second approach, providing alternative conditions under which the 2SLS could still have a causal interpretation when (ref) are too stringent for the data. In (ref), we revisit the average exclusion assumption proposed by FLL and show it is a special case of a zero-covariance condition: a restriction that may not always be justifiable in all empirical settings.
There is another approach that focuses on a well-defined policy-relevant parameter and examines how this parameter could be point-identified or set-identified using weaker and more credible assumptions. In such a case, the parameter of interest remains the same, but the (set) estimands may vary depending on the credible assumptions one would be willing to maintain. We will discuss this approach in (ref).\footnote{In this section, we mainly focus on the case in which the exclusion or monotonicity assumption is violated. When the random assignment assumption is violated, one can consider a partial identification, see mourifie2025layered.}
We first revisit the average exclusion and monotonicity conditions. For simplicity, suppose $Z$ has finite support as in FLL such that $\mathcal Z=\{1,2,\cdots, J\}$. The general form of the potential outcome model is, \[ Y=\tilde{Y}_{1} D + \tilde{Y}_{0} (1-D) , \quad \tilde{Y}_{d}= \sum_{z \in \mathcal Z}Y_{d}(z) 1\{Z=z\} ,\quad D=\sum_{z \in \mathcal Z}D_{z}1\{Z=z\}. \]
FLL proposes to relax (ref) and (ref) with the average exclusion restriction and the average monotonicity assumption, respectively:
Under (ref), FLL's Theorem 3 shows that the 2SLS estimand (of using $P(Z)$ as IV) has a causal interpretation since it can be written as a weighted average of the following treatment effect $\delta = \bar{Y}_{1} - \bar{Y}_{0}$ i.e.,
Note that the $\delta$ in (ref) is a deterministic function of the collection of potential outcomes $\{Y_{d}(z)\}_{d=0,1,z\in\mathcal Z}$. The 2SLS estimand is causal because it is a weighted average of $\delta$ and the weight $\omega$ is positive by the average monotonicity ((ref)-(b)).
(ref) below provides a generalization and more transparent discussion of the FLL's Theorem 3. First, we demonstrate that the average exclusion assumption is essentially equivalent to a zero covariance condition. Second, we show that (ref) indeed holds for any deterministic function of $\{Y_{d}(z)\}_{d=0,1,z\in\mathcal Z}$, not just for $\delta$.
To clarify these points, let us define $\alpha_{z} \equiv Y_{1}(z) - Y_{0}(z) $, and $ \tilde{\alpha} = \sum_{z \in \mathcal Z} \alpha_{z} 1\{Z=z\} =Y_1(Z)-Y_0(Z)$. Let $\alpha \equiv h(Y_{1}(1),...,Y_{1}(J),Y_{0}(1),...,Y_{0}(J))$ be an arbitrary measurable deterministic function of the collection of potential outcomes. $\delta$ defined in (ref) is a special case when we pick $h(Y_{1}(1),...,Y_{1}(J),Y_{0}(1),...,Y_{0}(J)) = \bar{Y}_{1} - \bar{Y}_{0} $. One could instead be interested in different treatment effects specific to each judge: $\alpha = \alpha_z$, $z=1,2,\cdots, J$. $\alpha$ can also be a quantity without clear economic interpretation such as $\alpha=\sum_{z \in \mathcal Z}zY_{1}(z)$.
The proof for the proposition is collected in (ref). Under (ref) (independence), (ref)(b) shows that the average exclusion restriction of FLL is, indeed, a special case of the zero-covariance assumption when $\alpha=\delta$. (ref)(a) further shows that if one targets an arbitrary quantity $\alpha \equiv h(Y_{1}(1),...,Y_{1}(J),Y_{0}(1),...,Y_{0}(J))$, and if one is willing to impose the same zero-covariance assumption on $\alpha$:
then one can always interpret 2SLS estimand as the weighted average of $\alpha$ with positive weights under the average monotonicity (ref)(b). This happens because the above zero-covariance condition in (ref) is a reduced-form condition, which assumes that the correlation between a reduced-form error (involving the parameter of interest) and the propensity score, i.e. $\text{Cov}\left( Y - \alpha D,P(Z)\right)=0$.
How does one assess the plausibility of the average exclusion condition? FLL provides a heuristic argument.\footnote{frandsen2023judging: “Average exclusion can be probed by examining the correlation between judge-level treatment propensity and judge-level averages of alternative channels through which judges may affect outcomes if such channels are observed. Average exclusion may be more plausible if these correlations are near zero."} However, this argument could also be invoked by anyone who wants to impose that (ref) holds for other $\alpha\neq \delta$. Also, it is difficult to justify why $\text{Cov}\left( (\tilde{\delta} - \delta) D + \tilde{Y}_{0},P(Z)\right)=0$ but $\text{Cov}\left( (\tilde{\alpha} - \alpha) D + \tilde{Y}_{0},P(Z)\right) \neq 0$ for other $\alpha\neq \delta$.
Furthermore, it is worth noting that the average exclusion restriction is not invariant to a relabelling of the treatment. In other terms, this assumption may hold if the researcher defines the treatment as $D$ equals $1$ if incarceration and $0$ if not, while it may not hold if the researcher recodes the treatment as $D$ equals $1$ if no incarceration and $0$ if incarceration. Indeed, after a relabelling, the zero-covariance in (ref) becomes: $\text{Cov}\left( (\tilde{\alpha} - \alpha) D + \tilde{Y}_{1},P(Z)\right)=0$. It follows that the average exclusion assumption is invariant to a relabelling if and only if $\text{Cov}\left(\tilde{Y}_{1},P(Z)\right)=\text{Cov}\left(\tilde{Y}_{0},P(Z)\right)$.
Despite all those discussed above, we do observe a direct way to assess the validity of (ref). In fact, under these assumptions, there is:
where $U$ and $L$ are, respectively, the upper and lower bounds of $\mathcal Y$. Therefore, if the support of the outcome is bounded from both above and below, then the absolute value of the 2SLS estimand must also be bounded.
In practice, it is not uncommon for researchers to have good reason to believe the assumptions hold after controlling for the judge's specific characteristics. We explore this idea in this section and demonstrate that it is closely related to the partial exclusion assumption (defined below) and the partial monotonicity assumption made in mogstad2019identification. Specifically, we decompose $Z$ into two components: $Z_I$ and $Z_c$, and we assume the monotonicity and exclusion restriction hold conditionally on $Z_c$. Here, $Z_c$ can be a judge's race or political party, and $Z_I$ is a vector of the remaining characteristics.
The partial exclusion assumption relaxes (ref) and allows the potential outcomes to depend on the subvector $Z_c$. For instance, when the treatment of interest is incarceration, judges could assign and differ in other punishments, such as probation, fines, or sentence length. These other punishments could directly affect potential outcomes, making (ref) unlikely. Minority judges may be less lenient in their sentence length than their majority counterparts johnson2014judges. Beyond the decision to incarcerate, different sentence lengths may have divergent effects on future labor market outcomes. If the sentence length is not observed or controlled, we would expect the potential outcome to depend on whether a judge is a minority judge through this channel. The partial exclusion assumption states that whether and how the judge assigns other types of punishment depends only on a subset of the judge's observable characteristics ($Z_c$), but not on others ($Z_I$). In other words, a defendant will end up with the same pair of potential outcomes $(Y_1(z_c), Y_0(z_c))$ as long as he or she is assigned to judges with the same observed characteristics $Z_c=z_c$. Finally, when the only instrument variable we observe in the data is the identity of the judge $Z_I$, then the partial exclusion assumption is equivalent to the original exclusion (ref).
The partial monotonicity (ref) was initially introduced in mogstad2019identification. It significantly weakens (ref) since it does not require comparing the level of leniency across judges with different observable characteristics. For instance, let $Z_c=(Z^R_c, Z^P_c)$ be composed of the following binary variables: $Z^R_c$ equal to $1$ if the judge is black or Hispanic and 0 if not, while $Z^P_c$ is $1$ if the judge is from the Republican party and $0$ if from the Democratic party. Imposing (ref) means it is not possible to have a black democrat judge be more lenient than a white republican judge for some defendants while being less lenient for other defendants, i.e., these two judges may have different cut-off points, but they rank all the defendants in the same order. Mathematically, we can not have both $\mathbb P(D(z_I,1,0)=1,D(z_I',0,1)=0)>0$ and $\mathbb P(D(z_I',0,1)=1, D(z_I,1,0)=0)>0$. However, there is a large body of empirical evidence of heterogeneity in the ranking of judges' leniency across different types of offense or defendants abrams2012judges,stevenson2018distortion. This is, however, compatible with the partial monotonicity. Its main advantage is that it no longer requires a uniform ranking of defendants across different judges. Judges' rankings are allowed to vary with their characteristics $Z_c$. Applying the result of vytlacil2002independence, the partial monotonicity condition can be characterized as a partial single threshold-crossing restriction under the independence assumption (ref), which we restated below.
Under (ref), we can apply the standard normalization, \[ D(z_I,z_c) = 1\left\{F_{U_{z_c}|Z_c}(\nu(z_I,z_c)|z_c)\geq F_{U_{z_c}|Z_c}(U_{z_c}|z_c)\right\} \equiv 1\left\{P(z_I,z_c)\geq V_{z_c}\right\}, \] where $F_{U_{z_c}}(\cdot)$ is the distribution function of $U_{z_c}$, $P(z_I,z_c)$ is identified from the observed $(D,Z)$ by $P(z_I,z_c) \equiv \mathbb P(D=1|Z_I=z_I, Z_c=z_c)$. Note by construction, $V_{z_c}$ follows $Uniform[0,1]$ distribution because the distribution of $U_{z_c}$ is absolute continuous; also, $V_{z_c}$ is independent with $(Z_I,Z_c)$.
The key difference between the STC and the Partial STC is even though $V_{z_c}$ follows $Uniform[0,1]$ distribution, each defendant does not face a single $V$. Instead, he or she faces a collection of $\{V_{z_c},z_c\in\mathcal Z_c\}$. This unobserved latent variable is now different for judges with distinct observable characteristics. The partial STC has a natural interpretation as an extension of the Roy model canay2024use. We can interpret $P(z_I,z_c) $ as the perceived gain of incarcerating a defendant by a type $z=(z_I,z_c)$ judge, and $V_{z_c}$ as the expected cost (but unobserved to the econometrician) of incarcerating a defendant. The particularity of the partial STC is that the expected cost can vary across judges with distinct observable characteristics $z_c$, but is fixed within judges with the same $z_c$. In the standard monotonicity assumption, the cost $V$ would be the same regardless of the characteristics $(z_I,z_c)$. For the same reason, the partial STC is also more reasonable in settings where decision-makers (judges) differ in their preferences and skills chan2022selection.
Here, we provide an example of eight judges deciding whether to incarcerate a given defendant to elucidate further the richer heterogeneity enabled by the partial monotonicity (or, equivalently, the partial STC) assumption. We consider the two observable characteristics of the judges introduced earlier, $Z_c \equiv (Z_c^{R}, Z_c^{P}) \in \{0,1\} \times \{0,1\}$. These two binary observable characteristics result in four types of judges. The eight judges are evenly allocated across these four types.
The left rectangle of (ref) shows the benefit and the expected cost of incarcerating the defendant in a separate unit segment for each judge. For example, $p_{11}$ and $p_{11}'$ are the benefits of the two black democratic judges with type $Z_c=(1,1)$ to incarcerate the defendant. The right rectangle of (ref) plots the benefit numbers of all eight judges on the same unit segment. Similarly, $U_{11}$ represents the expected cost of incarcerating the defendant by a black democratic judge: they share the same expected cost or skills. A judge incarcerates the defendant when the corresponding benefit is higher than the expected cost of incarceration. In (ref), the judges who incarcerate the defendant are blue-colored, while those who release the defendant are red-colored.
The behavior of the eight judges does not violate (ref) or (ref). However, the standard monotonicity (ref) is clearly violated (right rectangle of (ref)). Indeed, the judge with propensity score $p_{11}^{\prime}$ incarcerates the defendant (blue-colored), whereas judges with higher propensity scores $p_{01}$, $p_{10}$, or $p_{00}$ do not incarcerate the defendant (red-colored). Note that (ref) would not be violated for this group of judges only under one of these two conditions: (i) all four $V_{z_c}$ are greater than the maximum of the eight propensities or smaller than the minimum of all eight properties. In other words, when all judges make the same decision regarding this defendant, or (ii) judges who do not incarcerate the defendant must have lower benefit scores than judges who incarcerate the defendant. Moreover, one of these two conditions must hold for all defendants when we impose (ref) (or (ref)).
However, (ref) (or (ref)) does not require such a binding restriction. In particular, under the partial monotonicity assumption, defendants are allowed to be defiers across judges with distinct observed characteristics $Z_C$. For instance, in (ref) and using propensity scores to identify judges, the defendant is a $p_{11}^{\prime}-p_{01}$ defier, a $p_{11}^{\prime}-p_{10}$ defier, a $p_{11}^{\prime}-p_{00}$ defier, and a $p_{01}^{\prime}-p_{00}$ defier.
(ref) are weaker than (ref). We show that under these weaker conditions, it is still possible to identify meaningful treatment effect parameters.
The proof of (ref) is similar to (ref) after conditioning on $Z_c=z_c$ and therefore omitted. The identification results stated in (ref) (i)-(ii) demonstrate that whenever there are two judges with distinct $Z_I$ but share the same observed characteristics $Z_c=z_c$, the conditional Wald estimand identifies the LATE provided the propensity scores for these two judges are different. Moreover, when the distribution of $Z_I|Z_c=z_c$ allows one to take the derivative of $\mathbb E [g(Y)\vert P=\cdot, Z_c=z_c]$, the conditional LIV estimand identifies the MTE. This identification result is a local version of the standard LATE and MTE identification.
(ref) (iii) presents the testable implications of the weaker monotonicity and exclusion assumptions. The testable implications in (ref) (iii) are weaker than those in (ref) (i). To see this, let us consider the same example of the eight judges discussed above, where the outcome of interest is recidivism ($Y \in \{0,1\}$). We consider the same two observable characteristics of the judges, $Z_c \equiv (Z_c^{R}, Z_c^{P}) \in \{0,1\} \times \{0,1\}$. Let $ \theta^{d}(p) = \mathbb P(Y=0 , D=d \vert P=p) $ for $d \in \{0,1\}$. In this simple case, the sharp testable implications under the standard judge leniency design, i.e. (ref) are:
which is a total of fourteen inequalities. However, when invoking our weaker set of assumptions, we have only eight inequalities that characterize the sharp testable implications:
The comparison of the testable implications in (ref) confirms that the judge leniency design is more stringent than the conditional judge leniency design. Hence, whenever the standard judge leniency design is rejected, the researcher may rely on its relaxed versions as long as the testable implications derived in (ref) are satisfied.
In this paper, we derive the sharp testable implications for identifying assumptions for the judge's leniency design in a general framework where the instruments can be either discrete or continuous and propose a consistent test for the implications. Our simulation study and empirical results highlight the importance of considering sharp implications for a better use of information in the data. While we focus on the primary application of testing the validity of judge leniency design, our method can be readily applied to a broad range of other applications.