Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,322 characters · 23 sections · 8 citation commands
Detecting Grouped Local Average Treatment Effects and Selecting True Instruments With an Application to Estimating the Effect of Prison on Recidivism
\onehalfspacing \pagenumbering{gobble}
\pagenumbering{arabic}
Instrumental variables (IV) analysis is a well-established method to estimate the effect of a treatment on an outcome in presence of unobserved confounding. In a model with a single endogenous treatment and multiple IVs, different instruments can identify heterogeneous local average treatment effects (LATE) given that they satisfy the following LATE assumptions of \citet*{Imbens+94} and \citet*{Angrist+96}:
Recent advances in parametric statistics focused on separating valid instruments from invalid ones that violate (b) and (c), but did not consider the case of heterogeneous treatment effects. In this paper, we propose a method for detecting IVs that satisfy the LATE assumptions (b), (c), and (d) in heterogeneous treatment effect models among a set of candidate instruments that may also contain invalid IVs not satisfying the LATE assumptions. Our approach relies on the following two conditions. First, there must exist clusters of IVs which generate identical first stages, i.e., conditional treatment probabilities given the instrument. Second, the LATE assumptions must hold for a plurality (i.e., relative majority) of IVs within such clusters.
Our methodological contribution to the literature on valid IV selection consists of allowing for the presence of heterogeneous first stage effects (of the instruments on the treatment) and of heterogeneous treatment effects on the outcome and proceeds in two steps. In the first step, we make use of Agglomerative Hierarchical Clustering \citep*[AHC,][]{Ward1963Hierarchical} to find clusters --which we call clubs-- of IVs with identical conditional treatment probabilities, henceforth referred to as propensity scores, in order to generate homogeneous complier groups when pairing these clubs. The underlying intuition is that if two distinct IV pairs entail the same higher and lower propensity scores, then the complier groups in either pair must be identical in terms of unobserved characteristics if the LATE assumptions hold. Therefore, our approach permits clustering IVs into club-pairs with homogeneous LATEs if the IV exclusion restriction, exogeneity, and montonicity hold. We stress that the key assumption underlying this step is that propensity scores defining the first-stage effect of the instrument on the treatment are indeed clustered in the population. In the second step, within a specific club, we distinguish valid and invalid IVs, which violate one or multiple LATE assumptions, based on the plurality assumption. The latter states that the largest group of IVs within a club satisfies the LATE assumptions and thus, entails the same LATE, where a group is defined as a set of IVs with identical reduced form estimands, i.e., the expected outcomes conditional on the IVs. We determine the set of valid IVs by hierarchically clustering reduced form estimates within clubs (defined on first stages) and selecting the cluster containing the relative majority of IVs.
Our two-step procedure broadly fits into the fast evolving literature on the integration of machine learning methods in causal inference, by applying unsupervised machine learning for clustering compliers based on the first stage effects of IVs and selecting IVs that satisfy the identifying assumptions, respectively. We show that under certain regularity conditions, the first-step selection procedure consistently classifies propensity scores into clubs (if clubs exist), while the second-step selection procedure consistently detects those IVs which satisfy the LATE assumptions (if plurality holds). Furthermore, we investigate the finite sample behavior of our method in a simulation study and find it to perform well even under a relatively moderate number of observations per instrument in terms of first- and second-step clustering, when a subset of instruments violates the exclusion restriction.
Our study contributes to a growing literature on detecting and selecting valid and invalid IVs \citep*{Kang2016Instrumental, Windmeijer2019Use, Guo2018Confidence, Windmeijer2021Confidence, Apfel2021Agglomerative} under the assumption that at least a subset of instruments is valid, following the seminal work of Andrews1999Consistent'\ (Andrews1999Consistent). A drawback of previously available IV selection approaches is that they impose homogeneous treatment effects and thus, LATEs and can for this reason not distinguish violations of an identifying assumption from effect heterogeneity. Furthermore, the methods do generally not allow for simultaneous violations of the exclusion restriction and IV exogeneity \citep*{Apfel2022Falsification}. Finally, they do not consider violations of treatment monotonicity in the instrument as discussed in \citet*{Imbens+94}. We are the first to consider violations of both the IV assumptions and the homogeneity of treatment effects. More specifically, our method aims to simultaneously detect groups with distinct LATEs and instruments subject to violations of any of the LATE assumptions (while other contributions focus on a subset of IV assumptions like the exclusion restriction). \citet*{Sun2022Pairwise} propose a technique exploiting testable conditions of the validity of the instruments to remove invalid variation in the instruments. However, their conditions are necessary, but not sufficient and therefore, their procedure might select a different set of IVs than the true one. \citet*{Masten2021Salvaging} propose the falsification adaptive set that reflects model uncertainty when the baseline model with all instruments valid is falsified. Our method can also be regarded as an approach to salvage a falsified IV model, by detecting a valid subset of IVs.
As an empirical contribution, we apply our method in the well-studied context of whether incarceration actually decreases offenders' probability to reoffend. This question is policy-relevant, given that incarceration rates in the US reach 1 percent of the population in a given year. Starting with the seminal paper by \citet*{Kling2006}, a common estimation strategy in the literature on the effects of incarceration is to use the assignment of judges to cases as IVs which are associated with incarceration but not directly with the outcome. Many studies find zero or crime-inducing effects. However, the effects might vary with geography and the institutional context (e.g.\ conditions under imprisonment), as, for instance \citet*{Bhuller2020Incarceration} find crime-reducing effects of imprisonment in Norway. Moreover, the effect might also vary inside each country, even by court. Each IV, defined by a pair of judges who differ in terms of stringency, permits estimating a possibly distinct LATE. The LATE corresponds to the average effect of incarceration on recidivism among those offenders who comply with the IV in the sense that they are incarcerated under the more stringent judge, but not under the less stringent judge when considering a specific pair of judges as IV. As these LATEs may vary because they refer to different complier groups who are in general distinct in terms of unobservables, a conventional 2SLS based on using all IVs simultaneously might hide interesting heterogeneity in the effect of incarceration.
An additional concern with judge IVs is that they might not satisfy the LATE assumptions, for example by directly affecting the outcome, thus violating the so-called exclusion restriction, or through an association with unobserved characteristics affecting the outcome, thus violating IV exogeneity. For example, judges could violate the exclusion restriction by varying in terms of the use of alternative measures to incarceration, such as electronic monitoring \citep*{Loeffler2021Impact}. Furthermore, the monotonicity assumption may be violated, which imposes that more stringent judges are stricter across all types of offenses and offenders than less stringent judges such that any offender incarcerated under a less stringent judge would also be incarcerated under a more stringent judge. \citet*{sigstad2023monotonicity} provides empirical evidence that a substantial share of judge IVs may violate monotonicity, which for instance occurs if a judge who is rather stringent on average (i.e., has a high treatment propensity score) is more lenient towards specific types of offenses. Our method permits investigating effect heterogeneity depending on whether offenders are imprisoned under more or less stringent judges and thus, likely differ in terms of the probability of having committed relatively less or more severe offenses. At the same time, we allow for violations of the LATE assumptions by multiple judges.
In the empirical application, we consider a newly collected data set on minor crimes in Minnesota for the years 2009 to 2017 to estimate the effect of incarceration on recidivism in the US. To benchmark our results against conventional approaches, we first consider ordinary least squares (OLS) regression, which suggests crime-inducing effects, as well as linear IV regression based on two stage least squares (2SLS) when simultaneously relying on all judge IVs, which also yields crime-inducing effects. However, when using our first-step method for clustering judge-based IVs, we find that LATEs estimated by club-pairs are sometimes considerably lower, sometimes higher and statistically significant. This finding points to substantial effect heterogeneity of the effects of imprisonment on recidivism across stringency levels of judges, which is masked by 2SLS. Within a certain club-pair (which consists of several judges), constructing LATE estimators based on multiple judge-specific IVs, taking one club as basis, permits testing the overidentifying restriction that all LATE estimates are equal. Running joint tests on an overidentifying set of IVs within a club-pair mostly rejects the overall validity of IVs. For this reason, we also apply our second-step procedure for selecting those instruments that satisfy the LATE assumption under the plurality rule. After applying the second-step selection, the overidentification tests no longer reject the LATE assumptions for the remaining IVs.
The remainder of this paper is organized as follows. In Section (ref), we introduce the treatment effects model, the LATE, and the identifying assumptions. Furthermore, we demonstrate that two pairs of judges with identical propensity scores have the same LATE under these assumptions. Section (ref) introduces our two-step procedure of (i) detecting club-pairs of IVs (defined on judges) with equal lower and higher propensity scores and (ii) selecting those IVs not violating the LATE assumptions under the condition that the largest group of IVs satisfies the assumptions (plurality). In Section (ref), we provide a simulation study which points to a strong finite sample performance of our method with regard to detecting clubs with equal propensity scores and appropriately selecting club-pair-specific instruments that satisfy the LATE assumptions. Section (ref) presents our empirical application considering the effect of incarceration on recidivism in the US. Section (ref) concludes.
For modelling the causal effects of interest and discussing the assumptions required for their identification, we make use of the potential outcome framework, see for instance \citet*{Neyman23} and \citet*{Rubin74}. To this end, we denote by $D$ a binary treatment, e.g.\ incarceration, which can take values $d \in \{0,1\}$. In general, we will refer to random variables by capital letters and specific values thereof by lower case letters. We are interested in the causal effect of the treatment on a discrete or continuous outcome $Y$, in our empirical problem on a binary indicator on recidivism. Furthermore, we denote by $Z$ an instrumental variable (IV), which needs to satisfy certain conditions outlined below, which is assumably multivalued and discrete. That is, $Z$ may take values $z \in \{1, ..., J\}$, with $J$ corresponding to the total number of values of the instrument. Let $\mathcal{J} = \{1, ..., J\}$ be the full set of possible IV values. $N_z$ denotes the number of observations with a specific value $z$ and $\sum_{z} N_z = n$ is the overall number of observations in the sample. Our empirical application, the instrument values indicate the assignment to alternative judges, which differ in terms of incarceration rates, i.e.\ the likelihood to issue prison sentences. Furthermore, $D(z)$ corresponds to the potential treatment state when setting the IV to a hypothetical value $Z=z$ in the support of $Z$, while $Y(d,z)$ is the potential outcome under treatment value $d$ (either 0 or 1) and IV value $z$. Finally, let $\mathbbm{1}(\cdot)$ denote the indicator function, which is equal to one if its argument is satisfied and zero otherwise. We denote matrices in upper case and bold, $\mathbf{X}$, and vectors in lower case and bold $\mathbf{x}$.
Following \citet*{Angrist+96}, individuals can be classified into compliance types as function of how someone's treatment state depends on or reacts to switching the value of the instrument. Let us to this end consider two distinct IV values $z$ and $z'$, which can assumably be ordered in a sensible way, e.g.\ two judges with a higher and lower incarceration rate, respectively. Individuals satisfying $(D(z)=1,D(z')=0)$ are compliers in the sense that they are treated, i.e.\ incarcerated under the higher value of the IV ($Z=z$), i.e.\ when being assigned to a more stringent judge, while remaining non-treated under the lower value ($Z=z'$) when being assigned to a more lenient judge. Never takers neither receive the treatment under a higher, nor under a lower IV value, thus satisfying $(D(z)=D(z')=0)$, while always takers are treated under either IV value, i.e.\ $(D(z)=D(z')=1)$. Finally, defiers counteract the instrument assignment by being treated under the lower IV value implying a more lenient judge, while remaining untreated under the higher IV value implying a stricter judge, such that $(D(z)=0,D(z')=1)$ holds. It is important to note that different choices of values $z$ and $z'$ generally entail different proportions of complier types, see e.g.\ the discussion in \citet*{Froe02a}. For instance, the share of compliers is likely larger when considering two judges that greatly differ in terms of their stringency rather than two judges among which one is only slightly more stringent than the other.
\citet*{Imbens+94} discuss assumptions under which an IV-based evaluation approach nonparametrically identifies a local average treatment effect (LATE) on the compliers under heterogeneous treatment effects. For two instrument values $z>z'$, this LATE is formally defined as
In terms of assumptions, the instrument must be as good as randomly assigned, which rules out statistical associations with background characteristics affecting the outcome, and must not have a direct effect on the outcome other than through the treatment. Furthermore, defiers must not exist, while compliers must exist in the population. We subsequently formally state these IV assumptions, which permit recovering the LATE for a specific pair of instrument values $z$ and $z'$.
Assumption (ref), consists of two conditions. The first one states that the IV values $z$ and $z'$ are as good as randomly assigned, i.e.\ independent of the potential treatments as well as the potential outcomes. This rules out confounders jointly affecting $Z$ on the one hand and $D$ and/or $Y$ on the other hand. This is often referred to as exogeneity. The second condition states that the instrument constructed from values $z$ and $z'$ does not affect the outcome conditional on the treatment, implying that $Z$ does not have a direct effect on $Y$ other than through $D$, which is known as the exclusion restriction. For this reason, we define the potential outcome as function of the treatment only (rather than also the instrument) as long as we assume no violation of the exclusion restriction, i.e.\ $Y(d)$. Assumption (ref) requires that the potential treatment state can never decrease when shifting the instrument from $z'$ to $z$ and for this reason rules out the existence of defiers. Assumption (ref) imposes a non-zero (first stage) effect of the instrument shift on the treatment. Together with Assumption (ref), this necessarily implies the existence of compliers.
Under Assumptions (ref) to (ref), the LATE on compliers who are responsive to IV shifts from $z'$ to $z$ corresponds to a so-called \citet*{Wald40}-estimand. The latter consists of the (reduced form) average effect of the instrument shift on the outcome scaled by the (first stage) effect of the instrument shift on the treatment:
For the case of a discretely distributed IV with mass points at $z$ and $z'$, (ref) is equivalent to the probability limit of a two stage least squares regression (TSLS) in which only observations with $Z \in \{z,z'\}$ are considered. It is also worth noting that $\Pr(D=1| Z=z)-\Pr(D=1| Z=z')$ identifies the complier share $\Pr(D(z)=1,D(z')=0)$.
As an important matter of fact for our method suggested below, \citet*{Vy02} proves that Assumption (ref) is equivalent to imposing a so-called threshold-crossing model on the treatment. In this model, the treatment is an additively separable function of the instrument and an unobserved term and takes the value one whenever the function on the instrument is larger than or equal to the unobserved term. Formally,
where $V$ is a scalar (index of) unobservable(s), and $\psi(Z)$ is a nonparametric function of $Z$. This allows us to characterize the compliers (and other compliance types) in terms of the distribution of $V$. Because $D(z)=1$ implies that $\psi(z)\geq V$ in (ref) and $D(z')=0$ implies that $\psi(z')< V$, it is easy to see that the distribution of $V$ among compliers satisfies $\psi(z)\geq V > \psi(z')$. For this reason, it follows that
for values $v=\psi(z)$ and $v'=\psi(z')$ for the unobservable $V$.
Let us now assume that there exists another pair of instrument values $z^*,z''$, in our case two further judges, which also satisfies Assumptions (ref) to (ref) and generates exactly the same compliance behaviour as the previous pair $z,z'$. Formally, $v=\psi(z^*)=\psi(z)$ and $v'=\psi(z')=\psi(z'')$. It follows that
As both pairs of instruments refer to the very same complier group in terms of the distribution of unobservables $V$, the LATE identified by these pairs is identical. This in turn implies that the Wald estimand $\frac{E(Y| Z=z^*)-E(Y| Z=z'')}{\Pr(D=1| Z=z^*)-\Pr(D=1| Z=z'')}$ is equivalent to that in (ref) and that both denominators $\Pr(D=1| Z=z)-\Pr(D=1| Z=z')$ and $\Pr(D=1| Z=z^*)-\Pr(D=1| Z=z'')$ identify the share of compliers satisfying $\Pr(v \geq V > v')$.
The identification of an identical complier effect under either pair of instrument values is driven by the fact that the satisfaction of $\psi(z)=\psi(z^*)$ for any pair $z,z^*$ implies that $\Pr(D=1| Z=z)=\Pr(D=1| Z=z^*)$. This is the case because for any value $z$, $\Pr(D=1| Z=z)$ is the probability of the event $(\psi(z)\geq V)$: $\Pr(D=1| Z=z)=\Pr(\psi(Z)\geq V|Z=z)=\Pr(\psi(z)\geq V)$, where the first equation follows from equation (ref) and the second from the independence of $D(z)$ and $Z$ implied by Assumption (ref). The converse holds as well, i.e.\ $\Pr(D=1| Z=z)=\Pr(D=1| Z=z^*)$ implies that $\psi(z)=\psi(z^*)$. To see this, let us normalize $V$ such that it only takes values between 0 and 1. That is, rather than considering the unobservable $V$, we take its cumulative distribution function (cdf), denoted by $F_V$ which is bounded between 0 and 1. Formally, $F_V\sim Unif[0,1]$. As noticed in the literature on marginal treatment effects, see e.g.\ \citet*{HeckVytlacil00} and \citet*{HeVy05}, this normalization is innocuous in the sense that it does not affect the results of our treatment model postulated in Equation (ref). The latter can without loss of generality be reparametrized in the following way when expressing the treatment as a function of $F_V$ rather than $V$:
where $F_V(\psi(Z))$ is the cdf of $V$ evaluated at the value $\psi(Z)$. By the definition of a cdf and treatment equation (ref), $F_V(\psi(Z))=\Pr(V\leq \psi(Z))=\Pr(D=1|Z)$, which is the treatment propensity score. For this reason, there exists a one-to-one correspondence between $\psi(z)=\psi(z^*)$ and $\Pr(D=1|Z=z)=\Pr(D=1|Z=z^*)$. This implies that distinct pairs of instruments entailing identical pairs of upper and lower propensity scores necessarily identify the LATE for the very same complier group (i.e.\ with the same distribution of unobserved characteristics), if Assumptions (ref) to (ref) hold.
To ease notation, we subsequently denote the treatment propensity score as a function of the instrument by $p_z=\Pr(D=1 | Z=z) = E(D | Z=z)$, where the second equality holds because $D$ is binary. In a next step, we assume that propensity scores follow a specific cluster structure, such that the propensity scores within a cluster are homogeneous, even across different instrument values (e.g.\ judges). To this end, we introduce some further notation. Let $C_k$ denote some cluster $k$ of instruments and $p_k^0$ the constant propensity score of all instruments in this cluster. $K^0$ is the total number of clusters of IVs with the same propensity score and without the zero superscript $K$ is a number of clusters, but not necessarily the correct one. Formally, we make the following cluster assumption, in analogy to \citet*{Su2016Identifying} (who look at panel data models more generally):
Assumption (ref) implies that there exist sets of $z$ with the very same propensity score, which we call clubs, in line with the growth literature where this type of method is usually applied, where $k$ defines the club identity. This assumption seems restrictive at first, but as we will see soon, it is not particularly controversial as there is a way to verify it in the data. Specifically, one might have a case where each judge belongs to its own club, i.e. we have singleton clubs. The prevalence of several clubs permits forming pairs of clubs with distinct propensity scores $p_k^0$ and $p_{k'}^0$ in order to construct IVs with a non-zero first stage for LATE evaluation. In fact, IV methods relying on any judge in club $k$ and any judge in a different club $k'$ identify the very same LATE, if assumptions (ref) to (ref) are satisfied, which follows from ((ref)). The following theorem states the identification result for the LATE in a setting with clubs, adding assumption (ref).
In a next step, we refine our approach such that it allows for the possibility that some instruments violate any of assumptions (ref) and (ref) (henceforth “the LATE assumptions”), based on the previous insight that all IVs in the same club-pair which satisfy the LATE assumptions yield the same LATE. To this end, we exploit the fact that for any pairs of IVs with the same pairs of propensity scores, the denominator in (ref) is equal. This in turn implies that by Theorem (ref), the reduced form effect in the numerator in (ref) is equal, too, if the instruments satisfy the LATE assumptions and thus, have the same Wald estimand. More formally, for any pairs of instruments $z, z'$ and $z^*, z''$ satisfying the LATE assumptions and $p_z=p_{z^*}$ and $p_{z'}=p_{z''}$ such that $p_z-p_{z'}=p_{z^*}-p_{z''}=c$, it holds that
Hence using any of the judge-specific means of the outcome $E(Y | Z=z) := r_z$ within the same pair of clubs must yield the same reduced form effect. We maintain relevance, given that existence of the first stage ($c\neq0$) can be tested.
In many empirical applications, however, the assumption that all IVs satisfy the LATE assumptions might be challenged. Let us for instance assume that a pair of judges violates the exclusion restriction postulated in Assumption (ref). In this case, one judge directly affects recidivism relative to another one, for instance by making use of distinct alternative measures to imprisonment, such as electronic monitoring. This generally implies for a pair of judges $z, z'$ that the average direct effect of the instrument conditional on the compliance type $(D(z)=d,D(z')=d')$, denoted by $\gamma_{zz'}^{dd'}$, is non-zero:
In this case, the Wald estimand defined in (ref) no longer identifies the LATE but is biased, as for instance discussed in \citet*{Huber2014Sensitivity}, who demonstrates that under a violation of the exclusion restriction, the LATE corresponds to
Therefore,
where the expression to the right of the LATE $\Delta_{z,z'}$ characterizes the bias of the Wald estimand. The occurrence of such biases is not constrained to cases with violations of the exclusion restriction, but they generally arise under violations of any of the LATE assumptions. For instance, the bias occurring under a violation of the monotonicity assumption has been characterized in \citet*{Angrist+96}.
Next, we define the validity set, $\mathcal{V}$, which is defined as follows:
In words, this corresponds to the set of all instrument values from different clubs, which produce valid instruments if paired together. This appears similar to the set in \citet*{Sun2022Pairwise}. However, the difference is that \citet*{Sun2022Pairwise} define the validity pair set, i.e.\ the set of pairs which fulfil the LATE assumptions, while our set is defined in terms of single IV values, in our case judges. In our example, the validity set is the subset of judges that when paired together give judge-pair or club-pair LATEs without using invalid judges. These could be those judges that influence recidivism through other dimensions of their decision than imprisonment and judges where monotonicity is violated because a usually severe judge does not incarcerate individuals who would have been incarcerated under a more lenient judge.
To find single judges satisfying the IV assumptions, we consider expected recidivism by judge, denoted by $r_z=E(Y|Z=z)$. Next, we define groups:
A group $\mathcal{R}_g^k$ therefore corresponds to a set of IV values $z$ with the same mean outcome $g$ within a club $k$ as defined by Assumption (ref). In our case, these are judges with the same incarceration and recidivism rates. Let us denote by $\mathcal{R} = \{\mathcal{R}_g^1,..., \mathcal{R}_g^K\}$ the union of groups of reduced form averages across all clubs and note that $\mathcal{R}_g^k \subseteq C_k$. Next, we define the reduced form-based group with the highest number of IV values within a propensity score-based club as $\mathcal{R}_{max}^k = \mathcal{R}_{g}^k$, s.t. $|\mathcal{R}_{g}^k| > \underset{g' \neq g}{max}|\mathcal{R}_{g'}^k|$ and the union of such dominant groups across clubs by $\mathcal{R}_{max} = \{\mathcal{R}_{max}^1,..., \mathcal{R}_{max}^K\}$. $\mathcal{R}_{max}^k$ thus corresponds to the largest number of judges with the same recidivism rate within a club $k$. To identify LATEs based on such dominant sets of IVs, we assume that the latter correspond to the validity set provided in Definition (ref), which imposes an IV plurality assumption:
Plurality implies that the respective largest groups of judges with the same recidivism rates within any club of homogeneous incarceration rates satisfy the LATE assumptions and, thus, belong to the validity set. We point out that methodologically, our approach is somewhat different from existing studies aiming at selecting valid instruments, which pre-select relevant instruments and where the first-stage relationship affects group membership. In contrast, our approach relies on first forming clubs with a homogeneous first-stage association in order to detect violations by investigating the heterogeneity in reduced form parameters within clubs. Comparing between clubs, we automatically have a difference in propensity scores and hence relevance is given. Moreover, the first-stage does not affect group membership, which is solely defined in terms of the reduced form parameter, in contrast to studies invoking parametric assumptions when selecting valid instruments; such as \citet*{Kang2016Instrumental}, \citet*{Guo2018Confidence} and \citet*{Windmeijer2021Confidence}.
Before presenting the methodological details of our method, we briefly discuss the general idea of our approach. We recall that by Proposition (ref), pairs $z, z'$ and $z*, z''$ which satisfy the LATE assumptions and yield the same propensity scores $p_z=p_z^*$ and $p_z'=p_z^{''}$ identify an identical LATE (for the very same complier group in terms of unobservables). We write the Wald estimator, $\beta$, as
The key idea is to first detect clubs with equal treatment propensity scores $p$ and then isolate groups with equal reduced form outcome means $r$ in the data. We may then pair such groups to estimate LATEs. Because the quantities $p$ and $r$ in each group are homogeneous, then so are the LATEs of specific group-pairs. As discussed in more detail in the following sections, we suggest a two-step clustering procedure for detecting first stage-based clubs as well as instruments that satisfy the LATE assumptions. In the first step, we cluster propensity score estimates into clubs. In the second step, we use clustering within each club to determine the largest group of judge-based instruments with homogeneous estimates of the reduced form outcome. After this, we pair such largest groups with distinct propensity scores to obtain LATE estimates for specific club-pairs. This approach is illustrated in Figure (ref).
We use the Agglomerative Hierarchical Clustering (AHC) algorithm of \citet*{Ward1963Hierarchical} to perform both steps of our procedure. In order to further improve the clustering, we also combine AHC with the clustering algorithm in regression via data-driven segmentation \citep*[CARDS,][]{Ke2015Homogeneity}. We report the results of the CARDS approach in Appendix (ref). However, we cannot achieve any improvement using the CARDS approach compared to AHC. Furthermore, CARDS is not computationally efficient, while AHC is much faster and therefore, overall preferable. In a previous version of the paper, we also considered so-called Classifier-Lasso \citep*{Su2016Identifying} for the estimation of club identities. However, again we could not improve the results compared to AHC.
The first step for estimating LATEs consists of detecting clubs of treatment propensity scores based on Assumption (ref). We estimate club membership based on the Agglomerative Hierarchical Clustering (AHC) algorithm of \citet*{Ward1963Hierarchical}. This method was first applied to the selection of valid instruments by \citet*{Apfel2021Agglomerative} in the context of over-identified parametric models where the number of valid instruments exceeds the number of regressors and plurality is fulfilled. Here, we instead use it for clustering the treatment propensity scores as a function of the instruments. We denote by $\mathcal{C}$ a partition into clusters with $C_k \in \mathcal{C}$ and by $\hat{\mathcal{C}}$ the estimated partition of the data. The true partition is denoted by $\mathcal{C}^0$. In the following, we provide a description of the algorithm, which is closely related to that given in \citet*{Apfel2021Agglomerative}:
We note that distinct weighing schemes in the Euclidean distance correspond to distinct objective functions to be minimized. We follow the classical approach in \citet*{Ward1963Hierarchical} and define the sum of within-cluster variance as objective function. As discussed in Apfel2021Agglomerative, the results, however, do not rely on this specific choice of the dissimilarity metric. When running the algorithm, we obtain a path of $S = J-1$ steps, with clusters of size $|\hat{C}_k| \in \{1, ..., J\}$ at each step. Note that each number of clusters, $K$, entails a different estimated partition $\hat{\mathcal{C}}(K)$.
To apply Algorithm (ref) in practice, a central question is at which step one should stop the merging process of clusters. The penalty parameter to choose in this context is the number of clusters $K$. We propose the following procedure based on F-tests for equality of parameters to optimally select $K$. Starting with $K=1$, we test whether all of the propensity scores are equal. If the test rejects the null, we proceed to the step of the AHC with $K=2$ and test for equality of all propensity scores within each cluster (or club).
The asymptotic F-statistic for the test applied in Algorithm (ref) is defined as
where $\mathbf{R}(\mathcal{C})$ is a hypothesis matrix with dimension $(J-K) \times J$. For the IV of the first club to be tested for equality, the matrix entry is 1 and for the remaining IVs it is zero, except for one other IV in the club, whose entry is -1.\footnote{In this way, each matrix row has an entry of 1 and another entry of -1.} In this way, we test the equality of all coefficients. $\mathbf{0}_J$ is a $(J \times 1)$ vector of zeros. Under the $H0$, the test statistic follows a $\chi^2$-distribution with $J-K$ degrees of freedom.
After identifying clubs of IVs with comparable propensity scores, clubs with distinct propensity score values may be paired to obtain the first stage effect of the instrument on the treatment as the difference across club-specific propensity scores. In this context, one club of judge IVs in the pair may be considered as reference category and the other as focal category. Estimating the LATE within a club-pair can be based on the subsample which exclusively contains observations (court rulings) coming from either the reference judges or the focal judges in that pair. In this case, the club-pair-specific instrument can be defined as a dummy variable, which takes the value one whenever a defendant is assigned to one of the judges from the focal category and zero otherwise. This approach yields just-identified IV models for each club-pair (due to using subsamples with reference and focal judges), which we may consistently estimate by 2SLS under the satisfaction of the LATE assumptions.
One may also create multiple instruments within a club-pair based on individual judges in that pair and run a 2SLS regression using these instruments. More concisely, we can for instance consider all judges in the club with the lower propensity score as reference category and generate separate instrumental dummy variables for each judge in the focal category. Asymptotically, all pair-specific IV dummies must yield identical LATEs if the LATE assumptions as well as Assumption (ref) on first-stage clusters hold, as they refer to the same complier population in terms of unobservables. Finding heterogeneous LATEs within club-pairs therefore points to a violation of the identifying assumptions, which motivates the procedure outlined in the next section. In contrast, LATEs that are based on distinct club-pairs refer to distinct complier populations and can for this reason differ under our heterogeneous treatment effect model.
We then evaluate the information criterion $IC(\lambda)=\ln \left(\sigma_{N}^2(\lambda)\right)+ 0.5 \hat{K}(\lambda) (N)^{-1/2}$ at a fine grid of values $\boldsymbol{\lambda} = (\lambda_1, \lambda_2)$ and choose the combination of tuning parameters with the lowest value of $IC$. As an alternative procedure, we evaluate the F-test as in ((ref)) and choose the model that passes the test at a pre-specified significance level, with the minimal number of clubs.
where $\boldsymbol{\lambda}_{pass} = \{\boldsymbol{\lambda}: p_F(\boldsymbol{\lambda}) > \alpha_F \}$.
Theorem 6 in \citet*{Ke2015Homogeneity} and Theorem 2 in \citet*{Wang2018Homogeneity} state that if the preliminary segmentation is such that the classification algorithm wrongly assigns members to adjacent clubs, so that misclassification is not severe in this sense, and regularity conditions hold, the clustering algorithm finds the correct clubs with high probability. We expect that finite sample performance will improve through combined use of AHC and CARDS. We call this combined algorithm HiCARDS. \fi
After clustering on the first-stage propensity score, any remaining heterogeneity in the LATE must come from the violation of one or several LATE assumptions outlined in Section (ref). Put differently, if all judge-specific IVs constructed within a club-pair satisfy the LATE assumptions, then IV estimates based on the various judge IVs should yield comparable LATEs and tests of overidentifying restrictions in IV models should not reject. For instance, one may apply a Hansen-Sargan-type test to multiple IVs created from two paired clubs of homogeneous propensity scores. In classical parametric IV models which impose homogeneous treatment effects and do not rely on first-stage clustering for generating clubs, such tests may not only have power against violations of IV assumptions, but also against a violation of effect homogeneity. In our framework allowing for heterogenous treatment effects, however, the application of Hansen-Sargan-type tests within club-pairs tailors it to testing the LATE assumptions rather than functional form restrictions.
To allow for a violation of the LATE assumptions, we proceed to the second step of our procedure, a clustering approach for determining the validity set $\mathcal{V}$ as provided in Definition (ref). To this end, we apply Algorithms (ref) and (ref) another time, but now for detecting groups within a club that are homogeneous in terms of reduced form estimates of judge-specific mean outcomes $r_z$. This requires first estimating the reduced form parameters by a linear regression of the outcome $Y$ on IV dummies for all judges (but no constant) within a club with homogeneous propensity scores. Therefore, the instrument $Z$ satisfies $Z \subseteq C^k$ in club $k$. Then, the identical cluster procedure as previously used for detecting first-stage clubs is applied to the reduced form estimates of $r_z$, in order to classify them into homogeneous groups. Finally, we select the groups with the largest number of instruments within a club as the validity set, relying on the plurality assumption introduced earlier.
We note that our procedure is not the only feasible approach for selecting true instruments in heterogeneous treatment effect models under the plurality assumption. As an alternative to selecting instruments within first-stage clubs based on the reduced form mean, one can also estimate all possible judge-pair-specific LATEs within a club-pair (or union) and group them by our clustering procedure. More concisely, for two first-stage clubs $a$ and $b$ containing $J_a$ and $J_b$ judges, respectively, there are $J_a \cdot J_b$ possible judge pairs for computing the LATE. We may apply the AHC algorithm to cluster the LATE estimates which are closest to each other in terms of weighted Euclidean distances. After each clustering iteration, we can extract the estimates in the largest cluster of estimates and verify their homogeneity based on a Hansen-Sargan-type test. In simulations (not reported), we found this approach to yield very similar results as our suggested procedure of clustering the reduced form means. For simplicity, we focus on the reduced form-based approach. One advantage of the latter approach is that it permits testing individual judge IVs rather than pairs of judges. This implies that for two first-stage clubs $a$ and $b$, the clustering method is based on only $J_a+J_b$ first stage estimates, rather than $J_a \cdot J_b$ LATE estimates when pairing the judges.
Clustering first-stage clubs and reduced-form groups within clubs provides us with an estimate of clubs $\hat{\mathcal{C}}$ as well as the validity set $\hat{\mathcal{V}}$, which consists of the largest club-specific groups. Based on these estimates, we compute the LATE by the sample analog of equation (ref) or equivalently, by 2SLS, using observations in the validity set that come from two different clubs (and thus, differ in terms of propensity scores).
More formally, we define $Z_k$ to be an aggregated instrument indicating whether a judge IV $z$ in estimated club $k$ belongs to the largest group of reduced form mean outcomes:
where the subscript $max$ refers to the largest reduced-form group in first-stage club $k$. For two clubs $k$ and $k'$ with distinct propensity scores, the LATE estimator exclusively using observations from the two largest groups in two clubs (such that an observation $i$ satisfies $i: z \in \hat{\mathcal{R}}_{max}^{k} \cup \hat{\mathcal{R}}_{max}^{k'}$), denoted by $\hat{\beta}(\hat{\mathcal{R}}_{max}^k, \hat{\mathcal{R}}_{max}^{k'})$, then corresponds to
where the IV which sets observations to zero for members of one group and to one for members of the other group is
We term such an estimator a group-pair IV estimator (GPIV).
In the following, we prove that our procedure consistently classifies propensity scores and selects the validity set. Then, we go on to show that the IV-estimator using the selected groups is asymptotically normal. We first introduce estimands, which compare two clubs or groups
We have that $plim(\hat{\beta}({\hat{\mathcal{R}}_{max}^k,\hat{\mathcal{R}}_{max}^{k'}}))=\beta(\mathcal{C}, \mathcal{R}, k, k')$. The estimand for the GPIV estimator which uses the correct club identities and validity set is
We call the corresponding estimator the oracle GPIV estimator for $k, k'$. $K=K^0$ means that the number of clusters is equal to the true number of clubs. There could also be estimators with $K \neq K^0$. This GPIV estimator identifies the true underlying group-pair specific LATE, $\Delta(k,k') = E[Y(k')-Y(k)|v \geq V > v']$, i.e. $\hat{\beta}(\mathcal{C}_0,\mathcal{V}, k, k') \overset{p}{\rightarrow} \Delta(k,k')$ for $k, k' \in \{1, ..., K^0\}$. The following result will help establishing consistent classification.
This follows directly from Lemma (ref) in the Appendix. The idea is that we start with $K=J$ and then merge only clusters which contain propensity scores from the same club. Only when all $p_z$ from each club are in a cluster respectively, the algorithm starts to merge clusters with members of different clubs.
This theorem shows that asymptotically, algorithms (ref) and (ref) correctly select the club identities, i.e. the correct partition $\mathcal{C}^0$. The next corollary states that the analogous algorithms for the maximal group selection inside each cluster will select the validity set.
The proof of this theorem follows the ones of Theorem (ref) in this paper and of Theorem 1 in \citet*{Apfel2021Agglomerative} closely and is hence omitted. The main difference is that the correct selection of the validity set does not rely on correctly assigning each judge to its group. As long as the invalid judges are not in the largest group, the algorithms still correctly select the validity set, even for cases where $K \neq K^0$. By theorems (ref) and (ref) we now have that as $n \rightarrow \infty$, $\hat{\mathcal{C}} = \mathcal{C}^0$ and $\hat{\mathcal{R}}_{max} = \mathcal{V}$ with probability approaching one. The next theorem states that the GPIV is asymptotically normally distributed.
\fi
We investigate the finite sample properties of our clustering and testing procedure in a simulation study consisting of two settings, with and without invalid judge IVs, where invalidity refers to a violation of the IV exclusion restriction. We consider three different sample sizes to analyze consistency.
We analyse settings with unbalanced panels, with $J=10$ judges. To determine the number of cases per judge we draw from a uniform distribution $Unif(3, 5)$, multiply this value by 20, 60, or 100, respectively, and round the number to the next integer. We repeat this approach for all $J$ judges and construct dummy variables for each judge. Furthermore, we define the first-stage model for the treatment of an observation $i$ in the following way: $$D_i = \mathbbm{1}(\mathbf{z}_i\cdot \bm{\pi} > V_i)$$ where $\mathbf{Z}$ is the $(N_z\cdot J) \times J$ matrix of judge dummies and $\mathbf{z}_i$ is a row vector. $\bm{\pi} = (0.8, 0.5,0.3)'$ is the vector of first stage coefficients (or propensity scores) of the judge dummies. $V_i$ is the first-stage error and follows a uniform distribution, $V_i \sim Unif(0,1)$. In this setup we create three clubs with 4, 4, and 2 different judges per club, respectively.
To avoid the correlation between the steps of clustering and estimating the LATEs, we apply sample splitting. That is, we divide the cases per judge into a training and a test set. We use the training set to cluster the judges into clubs and apply the second step of the procedure to detect invalid judges. We use the groupings from this step to estimate the group-pair wise LATEs.
The outcome model is given by the nonlinear function $$Y_i = (D_i \cdot 0.5 + D_i \cdot U_i + \mathbf{Z} \gamma + U_i)^4.$$ $U_i$ is the error in the outcome equation and modelled as $U_i = 0.5\cdot V_i + W_i$, where $W_i = Unif(0,1)$ is a uniform random variable which is independent of $V_i$. As the first-stage error affects the error in the outcome equation, the treatment is endogenous, which motivates the application of our IV approach. $\bm{\gamma}$ is a vector of direct effects of the instruments on the outcome. Therefore, a nonzero entry $\bm{\gamma}$ implies a violation of the exclusion restriction for the related judge dummy and thus, of one of the LATE assumptions. We define two different settings regarding the invalidity of the judges, to be able to identify the accuracy of the LATEs in both stages, before and after the identification of the invalid IVs. In the first setting we set $\bm{\gamma} = (0\bm{i}_{10})$. This allows us to investigate the club assignment in detail. In the second setting, we set $\bm{\gamma} = (0.5, 0\bm{i}_{3}, 0\bm{i}_{2}, 0.4, 0.6, 0\bm{i}_{2})$, so that in club 1, three judges are valid and one is invalid and hence majority is still fulfilled, in club 2, two judges are invalid and two valid, and hence plurality is fulfilled, but majority is violated, and in club 3 all judges are valid. While we focus on violations of the exclusion restriction to investigate the finite sample performance of our method, we notice that one could also consider violations of other LATE assumptions, namely treatment monotonicity or random IV assignment.
We report the following statistics for the simulations. Under $\#clubs$, we report the mean number of clubs, under $\#corr$ we report the fraction of times the correct number of clubs has been selected. We report the estimated LATEs for each pair comparison, when the number of clubs selected is the true number of clubs ($1-2$, $1-3$, $2-3$). For comparison we also report the true (or oracle) LATE. This fully informed estimator can be computed directly from the data generating process of our simulation. We take the average of Hansen p-values for each repetition (over pair-comparisons) and then average these means over repetitions ($Hansen$ $p$). We also calculate the probability ($Power$) of detecting a non-zero effect by testing the beta coefficients against zero. Respectively we calculate the average coverage rate ($Cover$), by testing whether the estimated 95 percent confidence intervals contain the true LATE. We simulate the true LATE by calculating the mean of 5000 oracle estimates, estimated using data independent from the one used to illustrate the selection and estimation procedures.
We report the mean normalized mutual information (NMI), which is an indicator for the quality of the clustering, as used in Ke2015Homogeneity and Ana2003Robust. The NMI is defined for the comparison of a clustering $\mathcal{C}$ and the oracle clustering $\mathcal{C}_0$: $$\mathrm{NMI}(\mathcal{C}, \mathcal{C}_0)=\frac{I(\mathcal{C}; \mathcal{C}_0)}{[H(\mathcal{C})+H(\mathcal{C}_0)] / 2}$$ where $I(\mathcal{C} ; \mathcal{C}_0)=\sum_{k, j}\left(\left|C_k \cap C_{0j}\right| / \right) \log \left(\mid C_k \cap\right.\left.C_{0j}|/| C_k|| C_{0j} \mid\right)$ is the mutual information between the two clusterings and $H(\mathcal{C})=-\sum_k\left(\left|C_k\right| / \right) \log \left(\left|C_k\right| \right)$ is the entropy of $\mathcal{C}$. The NMI takes values between 0 and 1, with larger NMIs indicating more similar clusterings and a value of 1 meaning that they are equal.
In a separate table, we report the results for the second-step AHC: the average of the fraction of valid IVs detected ($ValDet$) and of invalid IVs detected ($InvDet$). Further we also report the percentage of repetitions in which all valid and invalid IVs were correctly identified ($CorVal$) and again $Hansen$ $p$, $Power$, $Cover$ as well as the estimated LATEs.
We first consider a setting without invalid IV dummies, implying that all entries of $\gamma$ in the outcome equation are equal to zero, and run 1000 Monte Carlo simulations. Table (ref) reports the mean OLS and 2SLS coefficient estimates and standard error, for the three sample size settings. In the smallest sample size setting (20) the OLS estimate is 18.87 with 1.02 as a standard error. The 2SLS coefficient estimate is slightly higher at 20.31 and 1.48 as standard error. We observe almost the same coefficient estimates for the two larger sample settings, with considerably lower standard errors for OLS and 2SLS. The F-statistic on the instruments, in all settings, indicate that judge dummies are strong instruments for incarceration.
Table (ref) shows the classification results after the first step of our method, for both settings, with and without invalidity. In the first step, we aim at detecting clubs of treatment propensity scores based on Assumption (ref). We estimate club membership based on our algorithms (ref) and (ref). For comparison, we also report the oracle LATE estimator. In Table (ref) we first have a look at the setting without any invalid IVs. Column 5 ($\#corr$) shows that using AHC we find the right number of clubs, respectively, in 33 percent, 95 percent, and 99 percent of the repetitions, based on the different sample sizes. Increasing the sample size greatly improves the clustering performance. The average of the Hansen p-value is considerably above any conventional significance level, indicating we cannot find evidence for invalid instruments. Therefore, the second classification step (based on the reduced form parameters) is not necessary. The CI coverage is close to its nominal level for all sample sizes.
In the setting with invalid IVs we observe a comparable performance, when it comes to the identification of the right number of clubs. However, the Hansen p-value is always below the 0.05 significance level, indicating a high rate of rejections of the Null. There are two possible reasons that lead to a rejection of the Hansen test: effect heterogeneity and invalidity. By classifying the judges into clubs with the same incarceration rate, we can rule out effect heterogeneity as potential cause of a rejection of the Hansen test. Therefore, if it still rejects after classifying judges into groups, the only reason left for rejecting is the presence of invalid judges.
To address invalidity, we apply the second step of our IV selection procedure. In Table (ref) we present the classification results after the second stage of our procedure, that is applied to identify and eliminate invalid judges from the estimation. We apply AHC in both steps of the procedure, using the classification results shown in Table (ref).
After applying the second step, Table (ref) shows that the average of the Hansen p-value is considerably above any conventional significance level for all methods, indicating we no longer find evidence of invalid instruments. The power is close to 1, depending on the sample sizes, while the Coverage is between 0.76 and 0.91 for AHC. Respectively 97 to 99 percent of the valid IVs have been correctly detected ($ValDet$). The fraction of correctly identified invalid IVs ($InvDet$) is slightly lower at respectively 69, 87, and 96 percent. In total, again respectively for the sample sizes, in 33, 68, and 90 percent of the repetitions AHC identified all valid and invalid IVs correctly ($CorVal$). An NMI very close to 1 also indicates that the correct grouping has been retrieved in many cases.
In Table (ref) we report the respective LATE estimates, averaged over all repetitions (in which the correct number of clubs was found). In the setting without invalidity, the estimated LATEs, using AHC, are extremely close to the oracle LATEs and the SE decrease with increasing sample size. In the setting with invalidity, we first evaluate the performance of LATE estimates after only the first (club assignment) step, without trying to detect invalid judges. As expected, given that invalid judges are still present in the estimation, mean coefficient estimates are far from the oracle estimates, irrespective of sample size. Applying the second, group selection step of the method, however, dramatically improves our estimation and now the estimates from our method and the oracle estimates are very close. With a small sample, when for each judge there are only 60 to 100 observations per judge, the variance can be very large, but the performance clearly improves in the medium sample setting, with 180 to 300 observations and it is best in the largest sample setting.
Overall, these simulations indicate that our method delivers reliable results, retrieving clubs and groups that fulfil the LATE assumptions and they indicate that the oracle properties shown earlier on in the paper indeed hold. Additionally, the methods can be expected to perform well already with samples of medium size. We expect that performance not only improves with the number of observations, but also with increasing value of violations, increasing degree of separation among propensity scores and decreasing error variance.
In recent years, several studies have addressed the question whether incarceration affects the likelihood of recidivism (future criminal behavior) or, in other words, whether sentencing offenders to prison affects the likelihood of them relapsing into criminal behavior. Incarceration could either decrease the likelihood of recidivism if prisons help rehabilitate offenders and reintegrate them into society. If, however, the time in prison integrates them into criminal networks or leads to a loss of human capital, incarceration might increase the likelihood of recidivism. In order to answer this question, authors typically aim at estimating the following empirical model:
where $Y_{jc}$ denotes the outcome, an indicator for recidivism, i.e.\ re-indictment or re-incarceration in a certain time window after conviction, $j$ is the judge-index and $c$ is the case-index. $D_{jc}$ is the treatment variable with the coefficient of interest $\beta$, indicating whether a judge has sentenced the convict to a term of imprisonment or has decided for a non-prison sentence, such as probation. $\mathbf{x}_{jc}$ are observed covariates potentially affecting the probability of recidivism, with the coefficient vector $\theta$. \citet*{Nagin2009Imprisonment} suggest to use prior record, offence type, age, race and sex as key control variables. $\varepsilon_{jc}$ reflects unobserved characteristics affecting the outcome.
Estimating equation (ref) by OLS is generally biased and inconsistent if the association between outcome and right-hand side variables is non-linear and/or unobserved confounders jointly affect $D$ and $Y$ even after controlling for $X$. For instance, the judge might have information (not available to the researcher) about the defendant's previous offenses or possible addictions, which may simultaneously influence the judge's sentencing and the likelihood of recidivism. Likewise, the sentence may be influenced by the defendant's behavior in court, which in turn may shed light on the defendant's potential to recidivate. In order to address this issue of unobserved confounders, several studies have leveraged the fact that in some judicial systems cases are randomly assigned to judges whose use of prison sentences varies systematically. They have considered dummies for being assigned to a particular judge or the incarceration rate by judge as instrumental variables. Following this literature, we also use a set of judge IV dummies.
Earlier evidence on the effect of imprisonment is mixed but studies that find crime-increasing and null effects are in the majority. \citet*{Loeffler2021Impact} provide a review of 13 published IV-based studies on the effect of incarceration on recidivism concluding that those based on data from U.S. courts mainly find either an insignificant or a significant recidivism-increasing effect. It appears that U.S. prisons are ineffective in terms of rehabilitation and resocialization, and thus do not serve to fight crime beyond general deterrence effects. In contrast, a study by \citet*{Bhuller2020Incarceration} finds that incarceration in Norwegian prisons has a significant recidivism-reducing effect and a positive impact on employment, pointing to considerable heterogeneity in the effect of incarceration across regions, possibly due to differences in prison infrastructure and support services, such as labor market training opportunities.\footnote{Many studies in this literature focus on US data. But there is also some research on data from other countries, such as Chile \citep*{Cortes2019juvenile}. Most studies examine the incarceration effect only among convicts, while some, such as \citet*{Cortes2019juvenile} and \citet*{Leslie2017unintended}, assess the impact of pre-trial detention among all individuals accused of a crime.}
Effect heterogeneity can also arise within the same institutional context, in that defendants sentenced to prison by different judges generally differ in their background characteristics (such as personality traits) and these in turn can influence the effect of incarceration. For this reason, the LATE, i.e.\ the effect of incarceration on recidivism among individuals who would be incarcerated under a stricter but not a more lenient judge might generally differ across distinct pairs of judges (or relatedly, distinct propensities of incarceration). Such effect heterogeneity would be masked when applying 2SLS simultaneously to all judge IVs, because the 2SLS estimator is a weighted combination of just-identified IV estimators, which in turn estimate LATEs. This is one motivation for using our approach.
Further, the instruments might fail to satisfy the identifying assumptions (ref) to (ref). For instance, judges could differ in terms of their use of alternative measures to imprisonment, such as electronic monitoring \citep*{Loeffler2021Impact}, which would violate the exclusion restriction postulated in Assumption (ref). Further, the IVs might violate Assumption (ref), requiring monotonicity of the treatment in the instrument, see for instance the discussion in \citet*{Frandsen2023Judging}. Depending on how different judges weigh the individual aspects of a case, a rather stringent judge with a relatively high rate of prison sentences could refrain from incarcerating a particular defendant, who would have been imprisoned by a judge with a lower incarceration rate. i.e. their (potential) sentences could contradict the judges' order of (average) severity. Such a situation entails the existence of defiers and thus a violation of weak monotonicity of the treatment in the instrument.
To nevertheless consistently estimate the LATEs of interest, we invoke the existence of clubs of judges with a similar propensity to incarcerate, as well as plurality, such that the largest group of judges constitutes IVs fulfilling the LATE assumptions. The cluster assumption appears realistic as long as judges can be plausibly categorized into a limited number of judge types, each with a distinct rate of prison sentences. As an example, \citet*{Green2010Using} exploit variation in judicial calendars, where it could be the case that judges assigned to the same calendar affect each other's decisions. Moreover, some judges have moved between courts and reappear in the data with different judge IDs. Judges might have been educated and practiced at the same institutions, establishing cultures of higher or lower use of prison sentences. Moreover, in the US, different severity of judges might reflect judges' allegiance to political parties.
The model developed in Section (ref) is applied to a data set of offenders in the U.S. state of Minnesota, which is composed of data from two primary sources: the first one is an extract from the Minnesota Judicial Branch case database containing information on the offender, including their full name, date of birth and place of residence, for all criminal cases from 2009 to 2020; the second source is a data set from the Minnesota Sentencing Guidelines Commission on all adult offender cases in Minnesota between 2001 and 2017, with information on the type and severity of the offense as well as the judge hearing the case. We link these two datasets using the case number as identifier in order to obtain a data set with years ranging from 2009 to 2017 on all criminal offense cases that includes information on offender, judge, offense and conviction.
For each offender in our data set, we identify all criminal cases in Minnesota in which they were involved between 2009 and 2017, using the offenders' full names and dates of birth as identifiers. Despite also having information on the offenders' place of residence, we do not include this information for identifying recidivism in order to account for changes of residence, which occur frequently, especially after returning from jail or prison. This way, we accept the low risk of incorrectly linking the cases of two individuals with the same name and date of birth, while reducing the risk of not detecting recidivism. In addition, we cannot detect recidivism if an offender committed a crime in another state, has moved to another state/country or has changed his or her name.
For estimating the effect of incarceration on offender recidivism, we consider recidivism within three years of sentencing as the outcome. Therefore, in order to observe the three-year post-conviction period of every offender, we must reduce our dataset to the years 2009 to 2014. According to the Minnesota Order for Assignment of Cases, all criminal cases in Minnesota are randomly assigned to a judge having jurisdiction in the county in which the crime is tried. After being assigned to a case, a judge in active service must preside over that case until its resolution, i.e., random judge assignment, as required for our IV approach, is guaranteed by Minnesota state case assignment rules. Senior judges, however, are permitted to opt out from hearing a case. We therefore remove all cases heard by a senior judge in order to ensure random judge assignment. Then, although the vast majority of judges in Minnesota remain in one and the same court throughout their tenure, there are some judges in the sample that have changed court during our observation period. These judges are assigned a different ID for each court such that the terms at different courts are treated as if belonging to different judges, in order to account for differences in county crime profiles, which in turn are reflected in the judges' sentencing practices.
To ensure that the vast majority of offenders have been able to re-offend in the data within the three years following sentencing, i.e., are not in prison for the entire three years during which we observe potential recidivism, we need to reduce the data set to minor crimes. An analysis of the recidivism effect based on the entire dataset would require strict control for crime and offender profiles in order to avoid bias in the estimated effect caused by the inclusion of observations for which recidivism is highly unlikely due to long-term incarceration. At the same time, including offenses that result in a long prison sentence would not improve the quality of the estimator for the recidivism effect. Recent studies on the effect of incarceration have reduced their datasets based on different rules: Loeffler2013does have concentrated on cases of the three lowest charge classes, Green2010Using on drug offenses. We follow Bhuller2020Incarceration by reducing the dataset based on the sentence lengths that an independent institution - in our case the Minnesota Sentencing Guidelines Commission - recommends for each observed offense given the type of crime, the severity of the offense and the offender's criminal history. We reduce our dataset to cases with presumptive sentences of up to three years (a detailed list of the included offenses can be found in Appendix (ref)). The resulting sample contains 48,849 cases involving 38,874 unique offenders. Only some 10 percent of the offenses in our sample did not result in incarceration.
Some 82 percent of the offenses in the final dataset have resulted in a sentence of up to one year in county jails, while in about 14 percent of the observed cases, offenders were sentenced to one to two years in state prison, meaning in 96 percent of the cases offenders had at least one year after their official release date to recividate. In only some 0.9 percent of the cases the offender was convicted to three or more years in state prison. Given that usually only two-thirds of a sentence are served in prison and the rest is on probation, the share of offenders having at least one year to recidivate is even higher than 96 percent.
Table (ref) in Appendix (ref) provides some descriptive statistics for our data, namely the mean of outcome and covariates in the total sample, as well as among those offenses resulting in a prison or jail sentence ($D = 1$) and those that are not sanctioned with a prison or jail sentence ($D = 0$). The descriptive statistics suggest that the distribution of prison/jail sentences differs not only in terms of crime type but also in terms of the offender's race and gender. The proportion of cases that resulted in a prison or jail sentence is lower among white or Hispanic offenders than among offenders of other races and higher among men than among women. The share of property crimes sanctioned with incarceration is smaller than that of crimes against persons, drug crimes, weapon offenses and sex offenses. The table also shows that the severity of crimes and the likelihood of incarceration are positively correlated, where the variable “Severity” is an indicator for the seriousness of a crime as defined by the Minnesota Sentencing Guidelines Commission, ranging from low (Severity = 1) to high (Severity = 11)\footnote{The Minnesota Sentencing Guidelines Commission defines a different severity grid for sex offenders. The severity levels of the sex offenders in our sample are translated into the standard severity level according to the sentence lengths the Minnesota Sentencing Guidelines Commission's suggests for each severity level.}.
We restrict the data to judges that heard at least 300 (results for 200 in (ref)) randomly assigned cases between 2009 and 2017 and, in addition, to cases tried in counties where no fewer than two of these judges are stationed in any given year. The resulting samples contain respectively 11,219 cases.
We start with our first-step AHC and choose the significance level for the F-test as 0.1/log(N). In table (ref), we show the results of the first stage clustering, for the sample with at least 300 cases per judge. We first partial out the controls from the outcome, treatment and the IVs. We then run the first-stage regression of the treatment (imprisonment) on all judge dummies. With that we get four clubs, of sizes 1, 3, 10, 11. We exclude the singleton club since the second stage selection would be pointless with a singleton club and we cannot be sure the judge is valid. The other propensity score means lie between 0.56 and 0.77.
In Table (ref) we run an OLS regression and a first baseline IV estimation, where we use judge dummies as instrumental variables. This IV approach is close to the approach by \citet*{Green2010Using} who use dummies for judicial calendars as instruments. We include race-, gender-, offense-type, year-dummies, severity, age, the squares of the latter two and race-gender interaction dummies. In practice, we partial these controls out of the outcome and the treatment variable. The OLS estimate is 0.06 and it is estimated with precision. The 2SLS estimate is 0.03, but it is insignificant, with an F-statistic at 76.50. The results of the baseline OLS and 2SLS estimations are in line with the IV literature on the effect of incarceration on recidivism \citep*[see e.g.][]{Loeffler2021Impact}.
The three further columns in Table (ref) show the estimates for the pairwise comparisons between the three remaining clubs. In Panel A we do not run an IV selection, but use all of the judges within the clubs to build an IV which compares two clubs. The coefficients range from -0.10 to 0.20. The number of judges involved here lies between 13 and 21. F-statistics are mostly high and range between 64 and 175. A red flag is that even though we found different estimates with relevant first stages, the tests of overidentifying restrictions still reject clearly, with p-values close to zero for two of the three comparisons.
If we look for the largest group of reduced form coefficients and reduce the number of IVs further via our second-step AHC, we slightly increase the first-stage F-statistics, which are between 68 and 176. For two of the three clubs the Hansen-Sargan tests now do not reject any longer at any conventional significance level. The estimates again range from -0.10 to 0.50. The estimate for the comparison between club 1 and club 2 is highly significant.
The first-stage clustering, as well as the validity set selection via clustering inside each club is presented in figure (ref). We can see that there are three clusters of propensity scores with the lowest and smallest one separated most clearly from the other two. Judge-specific propensity scores, ordered by mean propensity score are also shown in figure (ref) in the Appendix, along with their confidence interval. In the rightmost club, two judges are selected as invalid and their reduced form appear to be clear outliers as compared to the $r_z$ of the remaining judges.
The key takeaway of this application is that our method can discover different LATEs from a large set of IVs and can provide a list of LATEs that would otherwise be collapsed into a single 2SLS estimate. Using the second-step AHC we find subsets of IVs in the club-pairs that seem more likely to fulfil the exclusion restriction. Using the first-step AHC, we can find different clubs of propensity scores which yield very high first-stage F-statistics. Whether there is a monotonously increasing effect of the disaggregation of the 2SLS estimate into several LATEs on the first-stage F-statistic is unclear and is left for future research. Moreover, one might think about even more flexible ways to control for observables, such as through random forests or other machine learning methods.
In the Appendix, we have added two additional exercises to this application. First, we have repeated the analysis, but now looking at judges with at least 200 cases. When doing this, we find five clusters, of which one singleton cluster (Table (ref) and Figure (ref)). Invalid judges are now found in two of the clubs ((ref)). When inspecting pairwise LATEs, qualitative results are similar to before in that the positive 2SLS estimate seems to be driven by a few positive, large and significant LATEs. In section (ref), we tackle the problem of controlling for observed covariates in an alternative way, by considering a subset of the data that is similar in terms of observables.
\FloatBarrier
In this paper, we proposed a method based on hierarchical clustering for (i) estimating grouped LATEs under the condition that multiple instruments have the same firsts stage effect on the treatment and (ii) detecting instruments which satisfy the LATE assumptions, under the plurality condition that those assumptions hold for a relative majority of instruments with the same first stage. The method first groups instrument values based on their treatment propensity scores given the instrument and then groups the reduced form estimates (the mean outcomes given the instruments) within groups of propensity scores to determine the largest cluster of instruments as the one satisfying the LATE assumptions. This approach consistently selects clusters of identical first stages (and thus, complier groups) and of valid instruments to estimate the grouped LATEs under certain regularity conditions. Simulation results suggested a decent finite sample performance of our method even under a relatively moderate number of observations per instrument. Finally, we applied the procedure to assess heterogeneous effects of incarceration on recidivism when using randomly assigned judges as (binary) instruments, where it appears likely that the LATE assumptions fail for multiple judges. As a possible direction for future research, our method could also be extended to settings with non-binary instruments, like Mendelian randomization, where genetic markers are considered as potential instruments.