EconBase
← Back to paper

Multidimensional clustering in judge designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,424 characters · 13 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Multidimensional clustering in judge designs

abstractEstimates in judge designs run the risk of being biased due to the many judge identities that are implicitly or explicitly used as instrumental variables. The usual method to analyse judge designs, via a leave-out mean instrument, eliminates this many instrument bias only in case the data are clustered in at most one dimension. What is left out in the mean defines this clustering dimension. How most judge designs cluster their standard errors, however, implies that there are additional clustering dimensions, which makes that a many instrument bias remains. We propose two estimators that are many instrument bias free, also in multidimensional clustered judge designs. The first generalises the one dimensional cluster jackknife instrumental variable estimator, by removing from this estimator the additional bias terms due to the extra dependence in the data. The second models all but one clustering dimensions by fixed effects and we show how these numerous fixed effects can be removed without introducing extra bias. A Monte-Carlo experiment and the revisitation of two judge designs show the empirical relevance of properly accounting for multidimensional clustering in estimation. \\[1ex] Keywords: judge design, multidimensional clustering, fixed effect, instrumental variable, jackknife.\\ JEL codes: C21, C26, C36, K14.

Introduction

Due to the plausibility of their identification strategy, judge designs are popular to identify the effect of a judge's decision on an outcome variable. These studies use that cases are randomly assigned to judges and that the judges differ systematically in their decisions. Which judge presides can therefore serve as an instrumental variable (IV) for its endogenous verdict.

The number of judges and hence the number of IVs, is typically large and thus may cause a many instrument bias. The conventional method to analyse judge designs conceals that many IVs are used because, in a first step, it aggregates the judge identities into a single instrument. In general, however, combining the IVs does not remove the many instrument bias.

Most judge designs aggregate the IVs via a leave-out mean. This specific combination method overcomes the many instrument bias, but it does so only in case the data are clustered in at most one dimension. This dimension is implicitly defined by what is left out in the mean. The vast majority of judge designs, however, cluster their data in different dimensions when estimating the variance. The effect of these additional clustering dimensions on the many instrument bias is not well studied and is therefore the topic of this paper.

To show the empirical relevance of multidimensional clustering, we analyse 65 papers that use a judge design, listed in chyn2024examiner. (ref) gives the details of the overview to which we will return throughout the paper. For starters, we note that 58 papers use some form of leave-out, which therefore defines a clustering dimension. 51 out of these 58 cluster their standard errors in either the main specification or in the robustness checks on a different dimension, hence implying multidimensional clustering. Furthermore, eleven papers explicitly allow for multidimensional clustered data by using two-way clustered standard errors. Another seven papers use one way clustered standard errors, but are concerned with additional dependence, such that in robustness checks they cluster in other dimensions.

In this paper we study the effect of multidimensional clustering on the many instrument bias in judge designs. We start by analysing the effect of clustering in a single dimension. We show that clustering in the data exacerbates the many instrument bias in the conventional two stage least squares (2SLS) estimator compared to independent data. Moreover, clustering makes that the usual solution to the many instrument bias, the jackknife instrumental variable estimator angrist1999jackknife,blomquist1999small, no longer fully removes this bias. In case the data are clustered in a single dimension the cluster adaptation to the JIVE, the CJIVE ligtenberg2023inference,frandsen2023cluster, does remove the many instrument bias.

Furthermore, this CJIVE is equivalent to the usual approach to analyse judge designs which combines the judge dummies into a single instrument via a leave-out mean. This equivalence explains why there is no many instrument bias in judge designs with the data clustered in a single dimension. In addition, the equivalence allows us to focus on the model with an instrument for every judge.

The CJIVE only removes the bias due to clustering in a single dimension. Clustering in additional dimensions adds bias terms to the 2SLS estimator, which are not removed by the CJIVE. We remove these additional terms in a new estimator, the multidimensional (MD) CJIVE, which can be used for many instrument bias free estimation with multidimensional clustered data.

Judge designs often include control variables in their model. Removing these control variables, such that JIVE or one of its variants can be applied, can introduce extra clustered dependence in the data chao2023jackknife. In particular, in case the control variables take the form of a fixed effect (FE), the control variables are removed by subtracting group means from the observations. Since the group mean contains all observations from the group, this introduces additional clustering in the data. Our overview of 65 papers shows that FE are commonplace, since all papers use them in some form.

For the MD CJIVE this extra dependence means that additional terms need to be removed to eliminate the bias, but poses no further problems. Another way to account for the extra dependence caused by removing the FE is via the FE JIVE from chao2023jackknife. This estimator adapts the JIVE to correctly remove the FE in case the data are independent. The FE JIVE is therefore an alternative to the MD CJIVE in case all clustered dependence can be modelled by including FE. Since it is not always plausible that the clustering in every dimension can be captured by FE, we note that the FE CJIVE extension of the FE JIVE from ligtenberg2023inference allows for more general dependence in one dimension.

We briefly compare the MD CJIVE and the FE CJIVE. We find that the FE CJIVE may work well even if the clustering cannot be fully captured by FE. This happens when the elements of the instrument projection matrix which are used to weight the bias terms, do not covary systematically with the covariances of the errors. Furthermore, we show that in some models the more general MD CJIVE does not identify the parameter of interest, whereas the FE CJIVE does.

Our overview shows that the judge level is a particularly popular clustering dimension. Out of the 65 papers 30 cluster on the judge in the main specification. Another 8 papers do so in robustness checks. Clustering on this level, suggests that after accounting for all variables in the model, there remains dependence between the errors of the cases handled by the same judge. We show that this type of dependence is especially heavily weighted in the bias terms, but that in general neither the MD CJIVE nor the FE CJIVE can correct for clustering on the judge level.

We compare the different estimators in a Monte-Carlo simulation where we generate data with clustering in multiple dimensions. We find that 2SLS and JIVE are biased in this case. The bias of the CJIVE is close to zero in some simulation set ups, but larger in others. Furthermore, the FE (C)JIVE can be biased in case general clustering is modelled with FE. The estimates of the MD CJIVE perform especially well in case there is complex clustering on both dimensions.

Finally, we revisit two judge designs to study the effect of clustering on the many instrument bias in the coefficient estimates. The first is by di2013criminal who estimate the effect of electronic monitoring (EM) as an alternative to incarceration on recidivism. In this study there are possibly three more clustering levels beyond the individual level that is implied by the leave-out mean instrument. When we take these dimensions into account with the MD CJIVE, we find that the effect of EM on recidivism is slightly more negative than when we ignore the additional clustering. The FE CJIVE yields no interpretable results.

The second study we revisit is the one by agan2023misdemeanor. These authors estimate the effect of nonprosecution on recidivism. Due to FE there may be two additional clustering dimensions that are not taken into account in estimation. The MD CJIVE that removes from the estimator the terms associated to these dimensions, yields estimates that are less negative than when these dimensions are ignored. The FE CJIVE estimates are again less stable, but if we ignore the one outlier, we find that the estimates that correctly incorporate the control variables are more negative than those that naively remove the control variables.

\paragraph{Literature review} That many instruments bias conventional estimators has long been acknowledged. angrist1999jackknife for example used the theory from nagar1959bias to show how the bias of the 2SLS estimator increases in the number of instruments. angrist1999jackknife and blomquist1999small propose JIVEs that are robust against many instruments as they remove the terms in 2SLS that cause the bias. We take a similar approach for the MD CJIVE.

Also for judge designs a many instrument bias has been a concern from the very beginning. kling2006incarceration, who initiated the judge design literature, used JIVE on the separate judge identity dummies. In later studies it has been more common to use a leave-out mean instrument and 2SLS instead of JIVE, as also comes forth from our overview in (ref). Although with this estimation method it is not immediately clear that many instruments are used, the topic kept resurfacing in the judge design literature. See for example the notes in hull2017examiner and mikusheva2021inference.

The jackknife that is used in the JIVE is more broadly applicable to overcome many instrument problems in IV models. For instance, it has been applied to adapt LIML and Fuller type estimators in hausman2012instrumental and to derive identification robust tests in crudu2021inference, mikusheva2021inference and matsushita2020jackknife. In the latter setting the jackknife for independent data was generalised to allow for one-way clustered data by ligtenberg2023inference. Subsequently, frandsen2023cluster used this cluster jackknife in a cluster JIVE.

The cluster jackknife allows for data clustered in only a single dimension. Ever since cameron2011robust, however, two- or multi-way clustered standard errors are commonplace. In this paper we therefore generalise the cluster jackknife to a MD cluster jackknife. We furthermore note that the focus of the paper by cameron2011robust and the clustering literature in general, has mainly been on the effect of clustering on variance estimators and inference, whereas this paper studies the effect of clustering on bias in estimators for the coefficients.

A discussion on clustering specified to judge designs followed from the discussion on clustering in the treatment effect literature. There they showed that which clustering dimensions matter depends on which view one takes on the data generating process (DGP). In particular, they discern a sampling based view in which the data are drawn from an infinitely sized super population, and a design based view where a large part of the population is observed, but where the assignment of treatment differs over samples abadie2020sampling,abadie2023should. chyn2024examiner took the latter perspective for judge designs. They argued that the observations should be clustered, both for estimation and for inference, in how they are assigned to the judges. Here we abstract from the different views on the DGP and we take the clustering dimensions as given.

\paragraph{Outline} In (ref) we present the general judge design and the effects of clustering on 2SLS, JIVE and CJIVE. This section also shows the equivalence of CJIVE and 2SLS with a leave-out instrument. (ref) discusses multi-dimensional clustering and the MD CJIVE. The inclusion of FE and how the FE (C)JIVE can be an alternative to the MD CJIVE are covered in (ref). In the section that follows we compare the two. (ref) adds a note on clustering at the judge level. We assess the performance of the different estimators in the Monte-Carlo simulations in (ref) and the empirical applications in (ref). (ref) concludes.

\paragraph{Notation} Throughout we use the following notation. For any matrix $\bs A$, define the projection matrices $\bs P_{A}=\bs A(\bs A'\bs A)^{-1}\bs A'$ and $\bs M_{A}=\bs I-\bs P_{A}$. Let $\odot$ be the Hadamard product. Write $\bs D_{A}=\bs I\odot\bs A$ if $\bs A$ is a square matrix and $\bs D_{v}$ the diagonal matrix with the elements of $\bs v$ on its diagonal if $\bs v$ is a vector. Then define $\dot{\bs A}=\bs A-\bs D_{A}$ as the matrix with the diagonal elements set to zero. Finally, $\sum_{g\neq h}$ can be the double sum $\sum_{g=1}^G\sum_{h=1,h\neq g}^H$, whether this is the case and the values of $G$ and $H$ depend implicitly on the context. Other sums are defined similarly.

Model and clustering

In the general judge design, one is interested in the effect of a decision taken by an expert on an outcome variable. The expert and the decision can vary from judges needing to decide on conviction to patent examiners needing to decide on whether a patent is granted. See chyn2024examiner for an overview. For ease of exposition we stay close to the original setting of kling2006incarceration and call the expert a judge, what the judge needs to decide on a case and a positive decision a conviction.

The judge may be influenced by factors unobservable to the researcher that affect both his decision and the outcome variable. This makes the decision endogenous and consequently the effect of the decision cannot be estimated with ordinary least squares (OLS). Judge designs therefore take a different approach and leverage that judges differ systematically in their conviction rates. That is, some judges are stricter than others. Furthermore, judges are often randomly assigned to cases. These two points combined make that the identity of the judge is a valid and relevant IV for the decision on the outcome variable.

To be precise, we consider the following model to estimate the coefficient of interest $\beta$

equation[equation omitted — 142 chars of source]

where for case $i=1,\dots,n$, $y_i\in\mathbb{R}$ is the outcome variable, $X_i\in\mathbb{R}$ is a dummy variable indicating conviction, $\bs Z_i\in\mathbb{R}^k$ are mutually exclusive dummy variables indicating which of the $k$ judges handles the case and $\varepsilon_i\in\mathbb{R}$ and $\eta_i\in\mathbb{R}$ are the second and first stage errors. We stack the instruments per case in $\bs Z=(\bs Z_1,\dots,\bs Z_n)'$ and similarly for the other variables. For now we assume there are no control variables, but we come back to that in (ref).

Since each judge can handle only a limited amount of cases, we allow $k$ to grow with $n$ in the asymptotic approximations. We leave the dependence of $k$, $\bs\Pi$ and $\bs Z$ on $n$ implicit. We make the following assumptions on the judge identity dummies to ensure they are valid and relevant instruments.

assumption\begin{assenumerate*} • $\operatorname{E}(\varepsilon_i|\bs Z)=0$ and $\operatorname{E}(\eta_i|\bs Z)=0$ for all $i$. • $\bs\Pi\neq\bs 0$. \end{assenumerate*}

For now the data can be clustered on a single dimension, which we generalise to more dimensions later in the paper. That the data are clustered means that the data can be partitioned in groups or clusters, and that the observations within a cluster can have any dependence between them, whereas observations from different clusters are independent. We formalise the notion of clustering in the next assumption for which we use the notation from ligtenberg2023inference and that we repeat below.

The $n$ observations can be grouped in $G$ clusters with $n_g$, $g=1,\dots,G$, observations per cluster. Denote the indices of the observations in cluster $g$ by $[g]$ and for any $n\times m$ matrix $\bs A$ let $\bs A_{[g]}$ be the $n_g\times m$ matrix with only the rows indexed by $[g]$ selected. For later purposes, if $\bs A$ is $n\times n$ we also define $\bs A_{[g,h]}$ to be the $n_g\times n_h$ submatrix with the rows in $[g]$ and the columns in $[h]$ selected.

assumption\begin{assenumerate*} Conditional on $\bs Z$, $\{\bs\varepsilon_{[g]},\bs\eta_{[g]}\}_{g=1}^G$ is independent with mean zero. \end{assenumerate*}

In IV models, one usually estimates $\beta$ via 2SLS. When the number of instruments is large, 2SLS is inconsistent and there is a so-called many instrument bias. To see where this bias comes from and why clustered data can increase this bias relative to independent data, write the 2SLS estimator as $\hat{\beta}=(\bs X'\bs P_{Z}\bs X)^{-1}\bs X'\bs P_{Z}\bs y=\beta+(\bs X'\bs P_{Z}\bs X)^{-1}\bs X'\bs P_{Z}\bs\varepsilon$. Assume that a law of large numbers works on both the part in the inverse and the part after it. Then we can write the latter as

equation[equation omitted — 355 chars of source]

Since $\bs Z_i$ are judge dummies, we have an explicit expression for $\bs P_{Z}$. See (ref) for more details.

equation[equation omitted — 196 chars of source]

where $J(i)$ is the judge that handles case $i=1,\dots,n$ and $n_{j}$ is the number of cases that judge $j=1,\dots,k$ oversees. This shows that the elements of $\bs P_{Z}$ are well above zero, unless each judge handles an infinite amount of cases.

Therefore $\cref{eq:bias}$ becomes

equation[equation omitted — 498 chars of source]

where $\mathbbm{1}\{\cdot\}$ is the indicator function. The first term captures the usual many instrument bias that is also present in case the observations are independent. The second term is the additional bias to due to clustering in the data.

(ref) suggests that a solution to the many instrument bias is to set the elements in $\bs P_{Z}$ corresponding to the two terms to zero. The JIVE sets the elements corresponding to the first term to zero, which yields $\hat{\beta}_{J}=(\bs X'\dot{\bs P}_{Z}\bs X)^{-1}\bs X'\dot{\bs P}_{Z}\bs y$.

Clearly, in case the data are clustered JIVE does not remove all bias terms. This is solved by the CJIVE, which also sets the elements of $\bs P_Z$ corresponding to the second term in (ref) to zero. We write this as $\hat{\beta}_{CJ}=(\bs X'\ddot{\bs P}_{Z}\bs X)^{-1}\bs X'\ddot{\bs P}_{Z}\bs y$, where for any square matrix $\bs A$, $\ddot{\bs A}$ is the matrix with in the rows indexed by $[g]$ the elements in the columns indexed by $[g]$ set to zero, $g=1,\dots,G$. For instance, if the observations are sorted by cluster, the block diagonal elements of $\bs P_Z$ are set to zero.

In this section we have focused on the judge design that estimates $\beta$ with the judge identity dummies as instruments. An alternative, more common way of estimating $\beta$ is via 2SLS with a leave-out mean conviction rate as instrument. The leave-out mean conviction rate is a single instrument with as $i^{\text{th}}$ value the average conviction rate of judge $J(i)$, excluding any cases from the cluster that $i$ belongs to. For example, in judicial contexts it is common to exclude any case that involves the defendant of case $i$, as these cases are expected to be dependent. Therefore, the leave-out mean models the data as clustered on the defendant.

In (ref) we show that CJIVE and 2SLS with a leave-out conviction rate are equivalent up to weighting. This result was anticipated by french2014effect, aizer2015juvenile, dobbie2017consumer, doyle2017evaluating, arnold2018racial, dobbie2018effects, ribeiro2019pretrial, hyman2018can, baron2022there and chyn2024examiner, among others. This also explains why few authors worry about a many instrument bias in judge designs. Given this equivalence result, we focus also in what follows on the IV model with judge dummies as instruments.

Multidimensional clustering

Studies that use a leave-out mean conviction rate as instrument implicitly model the data as clustered on the dimension that is left out, often cases or individuals, and remove the many instrument bias in the estimators due to the dependence in this dimension. The overview of judge designs shows, however, that there are often more dimensions on which the data are dependent.

These judge designs adjust their standard errors to the additional dependence, but leave the point estimators unchanged. Similar to when clustering in a single dimension is ignored, this can lead to a many instrument bias. To formalise this, we introduce the following notation.

The data are clustered on $C$ dimensions with on dimension $c=1,\dots, C$, $G^{(c)}$ groups or clusters. On dimension $c$ there are $n_{g}^{(c)}$ observations in cluster $g=1,\dots,G^{(c)}$. The number of cases handled by judge $j=1,\dots,k$ that are in cluster $g$ on dimension $c$ is given by $n_{j,g}^{(c)}$. We denote the cluster that case $i=1,\dots,n$ is in on dimension $c$ as $G^{(c)}(i)$. Define $[g]^{(c)}$ as indices of the cases in cluster $g$ on dimension $c$. Take $[G(i)]=\bigcup_{c=1}^C[G^{(c)}(i)]^{(c)}$ as the indices of cases that share a cluster with case $i$ in some dimension.

Then multidimensional clustering is specified by the following assumption.

assumptionConditional on $\bs Z$ and for any $i=1,\dots, n$ and $j\notin [G(i)]$, $\{\varepsilon_{i},\eta_{i}\}$ is mean zero and independent from $\{\varepsilon_{j},\eta_{j}\}$.

For expositions purposes we will focus in the remainder of this section on the case in which there are $C=2$ clustering dimension. Our results extend straightforwardly to more dimensions.

We can then generalise the bias formula given in (ref) as

equation[equation omitted — 502 chars of source]

Compared to (ref) there is now a third term that captures the many instrument bias due to the dependence in the second clustering dimension. This additional term is not removed by a CJIVE that clusters on the first dimension.

We remove the additional bias term in a new estimator called the MD CJIVE. Similar to the CJIVE, the MD CJIVE uses a projection matrix where the elements corresponding to the bias terms are set to zero. To write this estimator succinctly, we define $\bs S^{(c)}$ to be an $n\times n$ selection matrix with in row $i$ zeroes everywhere except in the columns indexed by $[G^{(c)}(i)]^{(c)}$. Also, let $\bs S^{(1,2)}=\bs S^{(1)}\odot\bs S^{(2)}$ be the selection matrix with in row $i$ ones in the columns corresponding to $[G(i)]$. For any $n\times n$ matrix $\bs A$ we write $\bs B^{(c)}_{A}=\bs S^{(c)}\odot\bs A$ and $\bs B^{(1,2)}_{A}=\bs S^{(1,2)}\odot\bs A$. Then $\dddot{\bs P}_{Z}=\bs P_{Z}-\bs B^{(1)}_{P_Z}-\bs B^{(2)}_{P_Z}+\bs B^{(1,2)}_{P_Z}$ and the MD CJIVE $\hat{\beta}_{MDCJ}=(\bs X'\dddot{\bs P}_{Z}\bs X)^{-1}\bs X'\dddot{\bs P}_{Z}\bs y$. The terms that are subtracted remove the elements of $\bs P_{Z}$ causing a bias in both clustering dimensions. However, since the cases that are in the same cluster on both dimensions are subtracted twice, we need to add back these elements, which we do with the final term.

In (ref) we show consistency of the MD CJIVE under high-level assumptions similar to frandsen2023cluster. In (ref) we also present a variance estimator for $\hat{\beta}_{MDCJ}-\beta$, although we make no claims about asymptotic normality.

Fixed effects

So far we have considered a judge design model without exogenous control variables. In practice however, most studies include control variables in their model. The overview of judge designs showed that a particular popular type of control variable are FE, for example to control for time or location effects. A model with control variables is

equation[equation omitted — 191 chars of source]

where $\bs W\in\mathbb{R}^{n\times l}$ are the exogenous control variables and the other variables are as before. Similar to the number of instruments, the number of control variables, $l$, can increase with the sample size, but we again suppress dependence of $l$, $\bs W$, $\bs\Gamma_1$ and $\bs\Gamma_2$ on $n$.

Motivated by the Frisch-Waugh-Lowell theorem, the usual way to control for the exogenous variables $\bs W$, is to project them out. This is done by pre-multiplying both equations in (ref) by $\bs M_{W}$. Then conventional methods such as 2SLS or (C)JIVE, can estimate $\beta$ from this transformed model.

Projecting out a small number of exogenous control variables is in most cases an innocent transformation. When the control variables are numerous on the other hand, as is often the case with FE, projecting them out introduces extra dependence in the data. For FE this dependence takes the form of extra clustering dimensions, since the FE are removed by subtracting the group mean from every observation. If there are many FE, the group mean is taken over a small number of observations, making the contribution of each relatively large. Consequently, all observations within a group have nonnegligible dependence.

For the MD CJIVE the extra within group dependence induced by FE, is no obstacle to estimation without many instrument bias. It only adds an additional clustering dimension, with the clusters defined by the groups.

Ignoring the additional dependence in the transformed model on the other hand, can aggravate or introduce a many instrument bias, as chao2023jackknife show. They consider the JIVE for a IV model with many instruments, FE and independent observations. To see why projecting out the FE can lead to a many instrument bias, write JIVE for the transformed model as

equation[equation omitted — 180 chars of source]

For it to have no many instrument bias, we require $\operatorname{E}(\bs\eta'\bs M_{W}\dot{\bs P}_{\bs M_{W}Z}\bs M_{W}\bs\varepsilon|\bs Z)$ to be zero. Although $\dot{\bs P}_{\bs M_{W}Z}$ has zero diagonal, there is no guarantee that $\bs M_{W}\dot{\bs P}_{\bs M_{W}Z}\bs M_{W}$ has. Therefore the products of $\eta_i$ and $\varepsilon_i$ may have nonzero weight. Alternatively, $\{[\bs M_{W}\bs\eta]_{i},[\bs M_{W}\bs\varepsilon]_{i}\}_{i=1}^n$ is not a sequence of conditionally independent variables, which yields a many instrument bias as discussed in (ref).

chao2023jackknife provide an alternative JIVE, called the FE JIVE, that ensures that there is no many instrument bias, also after projecting out the control variables. They achieve this by solving for a diagonal matrix $\bs D_{\vartheta}$ which is such that $\bs P_{M_WZ}-\bs M_{W,Z}\bs D_{\vartheta}\bs M_{W,Z}$ has zero diagonal. The solution, $\bs D_{\vartheta}$ is given by the diagonal matrix with on its diagonal $\bs\vartheta=(\bs M_{W,Z}\odot\bs M_{W,Z})^{-1}\operatorname{diag}(\bs P_{M_WZ})$, where $\operatorname{diag}(\bs A)$ is a vector with the diagonal elements of $\bs A$. Using this adapted projection matrix instead of $\dot{\bs P}_{Z}$ in the FE JIVE yields $\hat{\beta}_{FEJ}=(\bs X'(\bs P_{M_WZ}-\bs M_{W,Z}\bs D_{\vartheta}\bs M_{W,Z})\bs X)^{-1}\bs X(\bs P_{M_WZ}-\bs M_{W,Z}\bs D_{\vartheta}\bs M_{W,Z})\bs y$. The bias term of the FE JIVE is given by $\operatorname{E}(\bs\eta'\bs M_{W}(\bs P_{M_WZ}-\bs M_{W,Z}\bs D_{\vartheta}\bs M_{W,Z})\bs M_{W}\bs\varepsilon|\bs Z)=\operatorname{E}(\bs\eta'(\bs P_{M_WZ}-\bs M_{W,Z}\bs D_{\vartheta}\bs M_{W,Z})\bs\varepsilon|\bs Z)=0$ by its zero diagonal, provided the observations are independent.

The FE JIVE is an alternative for the MD CJIVE to estimate $\beta$ many instrument bias free when there is clustering in multiple dimensions. Namely, if we assume that all clustering can be captured by cluster FE, such that after including these FE $\{\varepsilon_i,\eta_i\}_{i=1}^n$ in (ref) is an independent sequence, we can readily apply chao2023jackknife's (chao2023jackknife) FE JIVE.

The assumption that the clustering in every dimension can be fully captured by FE, is not always plausible. In such cases, the FE CJIVE generalisation of the FE JIVE from ligtenberg2023inference may be applicable. This generalisation allows for more general clustering in one dimension. We repeat the FE CJIVE in slightly adapted form below, for which we use the following notation. For any matrix $\bs A$, let $\operatorname{vecb}(\bs A)$ the vectorisation of the elements in the $\bs A_{[g,g]}$, $g=1,\dots,G$ and $\operatorname{vecb}^{-1}(\cdot)$ the operator that constructs out of a vector a matrix such that $\operatorname{vecb}^{-1}(\operatorname{vecb}(\bs A))=\bs S\odot\bs A$, where $\bs S$ is a selection matrix as before, but with the clustering dimension suppressed. Finally, let $*$ be the Kathri-Rao or blockwise Kronecker product, where we take the blocks to be defined by the clusters.

Assume that there is a single clustering dimension and that the observations are ordered per cluster. Then $\bs P_{M_WZ}-\bs M_{W,Z}\bs H\bs M_{W,Z}$ with $\bs H=\operatorname{vecb}^{-1}((\bs M_{W,Z}*\bs M_{W,Z})^{-1}\operatorname{vecb}(\bs P_{M_{W}Z}))$ has a zero block diagonal and therefore can be used in a FE CJIVE. That is $\hat{\beta}_{FECJ}=(\bs X'(\bs P_{M_WZ}-\bs M_{W,Z}\bs H\bs M_{W,Z})\bs X)^{-1}\bs X'(\bs P_{M_WZ}-\bs M_{W,Z}\bs H\bs M_{W,Z})\bs y$ has no many instrument bias.

In (ref) we show consistency of the FE CJIVE under assumptions similar to chao2023jackknife. We also provide a variance estimator for $\hat{\beta}_{FECJ}-\beta$ in (ref), but again we make no claims about asymptotic normality.

Comparison of MD CJIVE and FE CJIVE

In this section we compare the MD CJIVE and FE CJIVE. We first show why the FE CJIVE may work well even if there is general clustering that cannot be captured by a FE. Next, we discuss a case in which the MD CJIVE does not identify the parameter of interest, whereas a FE CJIVE does.

Modeling general clustering with a fixed effect

In some cases including cluster FE will eliminate most or even all of the many instrument bias, also when there is more general clustering in the data. Namely, including a cluster FE will set the average covariance within a cluster to zero. Consequently, if the elements of the jackknifed projection matrix that weigh these covariances in the bias do not vary too much or not systematically with the covariances, the weighted average of the covariances will also be close to zero, causing no or little bias.

To formalise this argument, consider the FE JIVE with clustering on a single dimension which is modeled with a FE. As before, let $\operatorname{E}(\bs\eta\bs\varepsilon'|\bs Z)=\bs\Xi$. Also, for notational convenience write $\tilde{\bs P}_{Z}$ for the FE JIVE projection matrix. The term causing the many instrument bias can be written as

equation[equation omitted — 492 chars of source]

where we used that $\tilde{\bs P}_{Z}$ has a zero diagonal. Then, $\bs M_{W}\dot{\bs\Xi}'\bs M_{W}=\dot{\bs\Xi}'-\bs P_{W}\dot{\bs\Xi}'-\dot{\bs\Xi}'\bs P_{W}+\bs P_{W}\dot{\bs\Xi}'\bs P_{W}$ is the matrix of covariances minus the cluster row and cluster column averages, corrected for over counting by adding the cluster row-column averages, which makes the cluster averages of the elements in $\bs M_{W}\dot{\bs\Xi}'\bs M_{W}$ zero. Consequently, if the $\tilde{P}_{Z,ij}$ do not differ systematically within a cluster, the bias will be zero.

An example when this happens is when the elements of $\tilde{\bs P}_{Z}$ are independent of those in $\dot{\bs\Xi}$ and have constant within cluster unconditional expectation $\mu_{g}$. Then we can write

equation[equation omitted — 539 chars of source]

This result is similar to the one in chetverikov2023standard, which states that with an exogenously assigned covariate of interest, it is often not necessary to control for clustering if the model includes an intercept. The argument is that after projecting out the intercept, the exogenous regressor is mean zero. Then, given that the regressor is independent from the errors, the expectation of the product of the errors and the regressor can be separated making the entire expectation zero, which removes the need to account for clustering.

The argument here is similar in the sense that we can split the expectation because one factor in the product is independent form the other. Another similarity is that one of these expectations is zero, which simplifies the expression. The argument is different in the sense that we consider a bias stemming from covariances, rather than a variance estimator itself and, more importantly, the inclusion of an intercept or FE does not make the exogenously assigned part, the $\tilde{\bs P}_{Z}$, mean zero in general. Rather, the covariances are made zero by the FE.

When FE CJIVE works, but MD CJIVE does not

The MD CJIVE can handle more general types of clustering than FE CJIVE. This robustness comes at a price. There are cases in which the MD CJIVE cannot identify the parameter of interest, whereas the FE CJIVE can. The following example illustrates our point.

Suppose that, similar to the judge instruments, the data can be clustered and there is one instrument per cluster. That is, $\bs Z_i$ is a zero vector except for the entry corresponding to the cluster that observation $i$ is in. Unlike the judge instrument however, we assume that there is variation of the instrument within the cluster. Furthermore, we assume that the errors consist of a cluster common component and an idiosyncratic component, such that all clustering can be modelled with a FE.

Since there is variation of the instrument within clusters, a cluster FE does not take away all instrument variation. Nevertheless, by the assumptions on the errors, it adequately models the clustering. A FE CJIVE therefore can identify the parameter of interest.

A MD CJIVE on the other hand, does not work. This is because the instrument vector for each individual contains only one nonzero element corresponding to the entry of the cluster it is in. Assuming the observations are sorted on clusters, this structure makes $\bs P_{Z}$ a block diagonal matrix, with the blocks corresponding to the clusters. A MD CJIVE takes sets all these blocks to zero, such that $\dddot{\bs P}_{Z}$ is a zero matrix, leading to no identification of the parameter of interest.

Clustering on the judge

Before applying the MD CJIVE and FE CJIVE to simulated and real data in the next sections, we want to add a note on clustering on a particular dimension: the judge level. We observe that many judge designs cluster their standard errors at the judge level, to allow for covariance in the errors of cases handled by the same judge. For illustration, out of the 65 judge designs we reviewed 30 clustered the standard errors in the main specification on the judge level. Another 8 papers did so in one of the robustness checks.

From (ref) we know that dependencies in the errors can bias the estimates. For clustering on the judge level this bias is especially severe, as per definition the cases with possible covariance due to this clustering dimension are handled by the same judge, and therefore all covariances are positively weighted.

Controlling for clustering at the judge level via the MD CJIVE or by including a FE for it is often not possible, however. The MD CJIVE would namely remove from $\bs P_Z$ all elements corresponding to cases handled by the same judge. In case there are no exogenous covariates these are the only non-zero elements of $\bs P_Z$, hence $\dddot{\bs P}_{Z}$ is a zero matrix making $\hat{\beta}_{MDCJ}$ undefined. Similarly, adding a FE for the judges and using the FE CJIVE is not possible, as the FE would be multicollinear with the instruments, such that $\hat{\beta}_{FECJ}$ is also undefined.

Monte-Carlo

In this section we study the finite sample performance of the MD CJIVE and FE CJIVE in a Monte-Carlo simulation. We generate data from (ref) with the data clustered on two dimensions. We first group $n=500$ observations in $k=30$ groups and assign each group a judge. Next, we further group the observations in two times $G^{(1)}=G^{(2)}=30$ groups for the clustering in the two dimensions. The groups are constructed in the same way as in ligtenberg2023inference. We first determine the group sizes and next randomly divide the observations over the clusters. In each clustering dimension $c=1,2$ we first set $n_g^{(c)}=\max\{1, n\exp(\gamma^{(c)} g/G^{(c)})/[\sum_{g=1}^{G^{(c)}-1}\exp(\gamma^{(c)} g/G^{(c)})+1]\}$ for $g=1,\dots G^{(c)}-1$ and $n^{(c)}_{G^{(c)}}=\max\{1,n-\sum_{g=1}^{G^{(c)}-1}n^{(c)}_g\}$. Next, all cluster sizes are ensured to be integer by rounding them down. Finally, to get the correct sample size, the first $n-\sum_{g=1}^{G^{(c)}}n^{(c)}_g$ clusters are increased by one. The unbalancedness of the clusters are determined by $\gamma^{(1)}=\gamma^{(2)}=2$, such that the smallest and largest clusters are 5 and 38. The groups for the judges are created in the same fashion with the same unbalancedness parameter.

Next, we generate the first stage errors $\eta_i=w_{1}\eta_{1,i}+w_{2}\eta_{2}+(1-w_{1}-w_{2})\eta_{3,i}$, where $\eta_{1,i}$ and $\eta_{2,i}$ are clustered components of the errors, one for each dimension, and $\eta_{3,i}$ is an idiosyncratic component. These components are weighted with $w_{1}=w_{2}=1/3$. We want to capture both the case in which FE can model all clustering and the more general case in which they cannot. Therefore we generate $\eta_{1,i}$ and $\eta_{2,i}$ from the factor model discussed, among others, in mackinnon2023cluster. That is, for $c=1,2$ and $g=1,\dots,G^{(c)}$, $\bs\eta_{c,[g]^{(c)}}=(\sqrt{1-(\omega^{(c)})^2} u^{(c)}_{g}+\omega^{(c)}\bs e^{(c)}_{[g]^{(c)}})f_{g}$. In case $\omega^{(c)}=0$ all clustering can be captured by cluster FE. We draw $u^{(c)}\simN(0,1)$ and $f^{(c)}\simN(0,9)$. Since we showed in (ref) that the covariances of the errors must vary systematically with the elements of the projection matrix to generate a bias, we set $\bs e^{(c)}_{[g]^{(c)}}\simN(\bs 0,\bs\Sigma^{(c)}_g)$, with $\bs\Sigma^{(c)}_g=\bs D_{A_g^{(c)}}^{-1/2}\bs A_g^{(c)}\bs D_{A_g^{(c)}}^{-1/2}$, where $\bs A_g^{(c)}=(\bs P_{Z,[g,g]^{(c)}}+\bs I_{n^{(c)}_g}0.01)$. $\bs\Sigma^{(c)}_{g}$ is therefore a correlation matrix constructed from $\bs P_{Z,[g,g]^{(c)}}$ plus a constant that is added to ensure invertibility. The second stage errors are constructed from the first as $\varepsilon_{i}=\rho\eta_{i}+\sqrt{1-\rho^2}v_{i}$ for $\rho=0.5$ and $v_{i}\simN(0,1)$.

Finally, we set $\beta=0$ and $\bs\Pi$ equal to a draw from $N(\bs 0,\bs I_{k})$, which we keep constant over the simulations.

(ref) shows the boxplots for different estimators when estimated for $10\,000$ data sets from above DGP with $\omega^{(1)}=\omega^{(2)}=0$, such the clustering in both clustering dimensions can be fully captured by FE. In particular, we show 2SLS; JIVE; CJIVE that clusters only on the first clustering dimension; FE JIVE that includes FE for both clustering dimensions; FE CJIVE that models the first clustering dimension with FE and the second dimension with more general clustering; and the MD CJIVE that clusters on both dimensions. In case we were unable to calculate one of the estimators due to for example a singular matrix, we exclude that estimate from the analysis.

As expected, we observe a bias for the three estimators that do not model all clustering dimensions. For 2SLS the bias is the largest and it decreases for JIVE and CJIVE when more and more terms terms causing a bias are removed. For CJIVE the bias is even close to zero.

In this DGP all clustering can be captured by FE, we therefore see that the FE JIVE and FE CJIVE perform well, as they are unbiased. The MD CJIVE on the other hand seems to yield estimates that are on average and in median too low. This relatively poor behaviour of the MD CJIVE may be attributable to the small sample size.

From this figure it also becomes clear that in this DGP it is more efficient to model the clustering dimensions by FE than by leaving them completely free, as the FE JIVE is less dispersed than the FE CJIVE , which in turn is less dispersed than the MD CJIVE.

Next, we turn to (ref) which again shows the boxplot for the same estimators, but now for a DGP in which $\omega^{(1)}=0$ and $\omega^{(2)}=1$ such that only the clustering in the first dimension can be captured by FE. We see that again 2SLS, JIVE and CJIVE are biased. CJIVE is also more biased than in the first figure. More importantly, since the FE JIVE can no longer correctly model all clustering, we observe that also this estimator is biased. Also note that the bias of the MD CJIVE shrunk compared to the previous DGP.

Finally, in (ref) we show the same results for a DGP in which $\omega^{(1)}=\omega^{(2)}=1$ such that neither clustering dimension can be modelled by FE. In this case the bias of the MD CJIVE is close to zero, showing its robustness to general clustering. All other estimators are biased, however. Note in particular that the bias of the FE JIVE and FE CJIVE is similar in magnitude or bigger than the bias of the CJIVE estimator, which ignores one clustering dimension.

figure[figure omitted — 663 chars of source]
figure[figure omitted — 738 chars of source]
figure[figure omitted — 742 chars of source]

Empirical illustrations

In this section we revisit two judge designs and show how the original results change when we account for clustering in multiple dimensions.

\texorpdfstring{di2013criminal}{Di Tella and Schargrodsky (2013)}

The first study that we re-investigate is the one by di2013criminal. The authors estimate the effect of assigning a criminal to electronic monitoring (EM) as alternative to incarceration, on recidivism. For this they use that in the province of Buenos Aires, Argentina, alleged offenders are randomly assigned to judges. In addition, the judges' assignment rates of EM differ widely, which makes that a judge design can be used to estimate the effect of EM on recidivism.

To estimate this effect, di2013criminal use a leave-out mean conviction rate, $\bs L$, as instrument. The model is augmented with exogenous covariates. That is

equation[equation omitted — 194 chars of source]

Here, $\bs y$ are dummy variables indicating recidivism, $\bs X$ are dummy variables indicating whether EM was assigned, $\bs L$ is the leave-out average EM assignment rate per judge, where individuals are left out and $\bs W$ are exogenous covariates. The exogenous covariates, of which there are 41 in total, include indicators for the most serious crime committed divided over 11 categories, year and judicial district fixed effects.

The model is estimated using data on 1503 cases of people that were released from the penal system between the first of January 1998 and the twenty-third of October 2007, in the Province of Buenos Aires, Argentina. In 386 of these cases EM was assigned to the defendant. The sample is constructed such that each of the 171 judges assigns EM at least one time and has handled 10 cases or more in a larger data set.

By using the a leave-out EM assignment rate, di2013criminal implicitly cluster at the individual level and remove the many instrument bias stemming from this clustering dimension. The reported standard errors are furthermore clustered at judicial districts, but this is not taken into account for the coefficient estimates. Additional clustering dimensions may emerge when the FE are projected out.

In (ref) we re-estimated $\beta$ accounting for the different clustering dimensions using the MD CJIVE and FE CJIVE, for which we adopt the model as in (ref). For both estimators we started with the model that naively projects out the controls and clusters on the individual level as defined by the leave-out mean instrument in the original paper. Next, we keep adding clustering dimensions implied either by the standard errors or by the FE. The MD CJIVE takes care of the clustering by removing the terms from the estimator. The FE CJIVE adds FE for the clustering dimensions implied by the standard errors and correctly removes these. When the clustering dimensions implied by the FE are added to the FE CJIVE we also correctly remove these FE. Finally, for this estimator we give the estimates obtained by removing all control variables via the FE CJIVE.

We start with the estimates obtained taking the clustering into account via the MD CJIVE as given in the second column of (ref). By leaving out the terms corresponding to the judicial district and the most serious crime, the estimates are almost one and a half times larger in magnitude. If we also model the last clustering dimension implied by the FE by leave-out the estimate jumps back close to the original level, but stays slightly more negative.

The estimates from the FE CJIVE, as given in the third column of (ref), are unstable. The first estimate, which only clusters at the individual level, is the same as the one by the MD CJIVE, since it does not remove any control variables yet. Throughout we use the individual level as the level that is modelled via leave-out. Adding FE for the judicial district and correctly removing these makes the estimates more negative. Perhaps even unrealistically so, since it would imply that criminals on the margin have more than 60% lower recidivism rates if they receive EM instead of a prison sentence.

Also correctly removing the FE for most serious crime makes the estimates even more negative. If we remove the year FE with the FE CJIVE on the other hand the estimates turn positive. Finally, in the last line we give the estimate of the FE CJIVE with all control variables. This estimate is negative and large in magnitude, but does not have a sensible interpretation, however, as it is far outside the zero and one range.

table[table omitted — 1,155 chars of source]

\texorpdfstring{agan2023misdemeanor}{Agan et al. (2023)}

The second study we revisit is the one by agan2023misdemeanor. This study estimates the effect of not prosecuting someone who committed a nonviolent misdemeanour offence on subsequent criminal behaviour. The authors use that in Suffolk County, Massachusetts, cases of nonviolent misdemeanours are randomly assigned to judges, conditional on the time, the day of the week and the location. The varying strictness of the judges can therefore be used to identify the effect of nonprosecution.

The model to estimate this effect is similar to (ref) from the previous section. Now $\bs y$ indicates whether the defendant faced a new criminal charge within two years after having committed the nonviolent misdemeanour, $\bs X$ is an indicator variable for non-prosecution and $\bs W$ capture individual characteristics and court and time FE. The instrument is constructed differently from before, however. Since the types of misdemeanours vary over time, the day of the week and the location, agan2023misdemeanor residualise the prosecution decisions with respect to FE for these variables before taking a leave-out average. $\bs L$ is therefore a residualised leave-out average prosecution rate per judge, where individuals are left out.

The model is estimated using data on misdemeanour criminal complaints in Suffolk County, Massachusetts, between the first of January, 2004, and the first of January, 2020. In total this data set contains 67\,060 cases involving 53\,358 individuals. The cases are handled by 315 judges. In addition to the prosecution decision, whether the defendant faced new charges within two years and the ADA identity, the data set also contains characteristics of the defendant. agan2023misdemeanor use 16 of these as controls. Together with an intercept and the court-day-of-week and court-month fixed effects, this brings the dimension of $\bs W_{i}$, after removing redundant controls, to 1276.

In their main analysis agan2023misdemeanor cluster the standard errors at the individual and judge level. The many instrument bias due to clustering on the individual level is correctly removed by using the leave-out average prosecution rate. The many instrument bias from clustering on the judge level, on the other hand, is not addressed and can also not be taken care of using the methods developed in this paper. Other relevant clustering may be introduced by projecting out the court-time FE.

In (ref) we show again estimates accounting for different clustering dimensions similar to the previous subsection. Note that, we adopt the model as in (ref) and hence in the baseline model we project out the exogenous control variables differently than in the original study. Contrary to the results in the previous subsection, we see that the MD CJIVE estimates become less negative when taking more clustering dimensions into consideration, which implies a weaker effect of nonprosecution on recidivism. In particular the final estimate is much smaller magnitude than the previous two. This may be due to the court-month clusters being relatively large, which means that many terms need to be removed from the estimator, which can change its value also relatively much.

The FE CJIVE shows again less stable behaviour than the MD CJIVE. The first two and final estimates are all close to each other, but the third estimate deviates. Also, as opposed to the estimates from the MD CJIVE, we observe accounting for more clustering makes the FE CJIVE more negative.

table[table omitted — 1,034 chars of source]

Concluding remarks

In this paper we argued that most judge designs imply that their data are clustered in multiple dimensions due to how they specify their estimator, their standard errors and control variables. Nevertheless, only a single clustering dimension is taken into account by the 2SLS estimator with a leave-out instrument. Consequently and as we show in Monte-Carlo simulations, a many instrument bias may persist.

Conventional estimators cannot handle multidimensional clustered data. We therefore proposed the MD CJIVE and FE CJIVE that either remove from the conventional estimators the terms associated with the bias or eliminate the clustering by including fixed effects. In Monte-Carlo simulations we found that these estimators are unbiased as long as the modelling assumption underlying them are satisfied. Furthermore, we show that modelling the clustering with FE might work well even if the real data clustering is more complex. However, these two new estimators are in general not a solution to a bias that may emerge due to clustering on the judge level, which is an especially popular clustering level.

Next, we showed the empirical relevance of clustering in judge designs by revisiting the studies from di2013criminal and agan2023misdemeanor. We found that taking the clustering into account via the MD CJIVE usually yields more stable results than via the FE CJIVE and that these estimates can differ from those that ignore the clustering.

Finally, for future research we note that the adaptations we made to the projection matrix in the 2SLS estimator to obtain the MD CJIVE and FE CJIVE might also work in other methods for IV models. For example, our techniques can potentially generalise the HLIM and HFUL estimators hausman2012instrumental or the many instrument and identification robust tests crudu2021inference,mikusheva2021inference,ligtenberg2023inference,matsushita2020jackknife to multidimensional clustered data.