EconBase
← Back to paper

Was Javert right to be suspicious? Marginal Treatment Effects with Duration Outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,696 characters · 17 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Was Javert right to be suspicious? Marginal Treatment Effects with Duration Outcomes,

\thispagestyle{empty}

abstractWe identify the distributional and quantile marginal treatment effect functions when the outcome is right-censored. Our method requires a conditionally exogenous instrument and random censoring. We propose asymptotically consistent semi-parametric estimators and valid inferential procedures for the target functions. To illustrate, we evaluate the effect of alternative sentences (fines and community service vs. no punishment) on recidivism in Brazil. Our results highlight substantial treatment effect heterogeneity: we find that people whom most judges would punish take longer to recidivate, while people who would be punished only by strict judges recidivate at an earlier date than if they were not punished. \scriptsize{Keywords: Duration Outcomes, Instrumental Variable, Alternative Sentences, Recidivism} \scriptsize{JEL Codes: C24, C31, C36, C41, K42}

\setcounter{page}{1}

quoteTo owe his life to a malefactor, to accept that debt and to repay it; to be, in spite of himself, on a level with a fugitive from justice, and to repay his service with another service; to allow it to be said to him, “Go,” and to say to the latter in his turn: “Be free”; [...]---this was what overwhelmed him. \begin{flushright} Les Misérables by Victor Hugo \end{flushright}

Introduction

Many relevant applications in the causal inference literature face two simultaneous identification challenges: endogenous selection into treatment and right-censored data. For example, in crime economics, when analyzing the effect of a punishment on defendants' time-to-recidivism, a researcher has to consider that judges observe more information than the econometrician when making their decisions and that time-to-recidivism is a right-censored variable Huttunen2020,Giles2021,Possebom2022,Lieberman2023. A similar problem arises when analyzing the effect of rehabilitation programs on recidivism Alsan2024. In labor economics, the same identification challenges appear when analyzing the effect of receiving unemployment benefits on unemployment spells Delgado2022. In the health sciences, when studying the effect of a drug on survival time, a researcher has to address both identification problems too Sullivan1993.

In this article, we are interested in addressing these two problems and, simultaneously, unpacking treatment effect heterogeneity with respect to unobserved, individual-specific resistance to treatment. Towards this end, we study the identification, estimation, and inference for the distributional marginal treatment effect ($DMTE$) and the quantile marginal treatment effect ($QMTE$) functions when the outcome variable is right-censored. In the crime economics example, these functions capture the distributional effects of receiving a punishment on time-to-recidivism for the defendant who is at the margin of being fined, conditional on the amount of evidence in her case.

By analyzing these distributional treatment effect parameters at different values of judge leniency, a researcher uncovers a detailed picture of how punishments heterogeneously affect recidivism. This picture can be used to design better sentencing criteria and/or train judges to follow a specific protocol. For example, suppose that one finds negatively sloped $QMTE$ functions with some positive and negative effects. This finding would suggest that defendants who would be punished even by very lenient judges---i.e., defendants with low unobserved punishment resistance---would take more time to recidivate as a result of the punishment (punishment is working as intended). On the other hand, defendants who would be punished only by very strict judges---i.e., defendants with high unobserved punishment resistance---would recidivate sooner than if they were not punished (punishment is not effective, perhaps because of scaring effects of a criminal record). Such degree of heterogeneity is usually washed out when using more standard summaries of treatment effects such as local average treatment effect (LATE) Imbens1994 or the IV estimand.\footnote{To be fair, we stress that LATE was not meant to highlight this degree of heterogeneity and that it has the advantage of only requiring binary instruments. However, when the empirical setting presents a continuous instrument, the definition of a “complier” is less clear than in the binary instrument case, potentially making the LATE results more challenging to interpret formally.} Thus, $DMTE$ and $QMTE$ parameters can be used to unpack treatment effect heterogeneity and provide additional economic insights.

Our methodological results highlight that extending the marginal treatment effect ($MTE$) framework of Heckman2006 to deal with a duration variable subject to right-censoring introduces some interesting challenges depending on the censoring mechanism. For instance, if censoring is independent of potential outcomes, we can point-identify the distributional marginal treatment effect and quantile marginal treatment effect functions for some, but not necessarily all, support points and quantiles. Nonparametrically identifying the entire $DMTE$ and $QMTE$ functions is only possible if the support of the censoring variable is at least as large as the support of the duration outcome. Thus, in practice, in the presence of censoring, the conditions for point-identification of (average) $MTE$ parameters are arguably too restrictive in many applications.\footnote{When this support restriction is not satisfied, one can nonetheless nonparametrically point-identify truncated $MTE$ functions, which are still well-defined causal parameters. See Appendix (ref) for more details.} As we discuss in detail, working with $QMTE$ and $DMTE$ side-step these issues. We propose semiparametric estimators and inference procedures for the $DMTE$ and $QMTE$ functions and establish their large sample properties.

We also discuss two potential avenues to handle cases where censoring possibly depends on the potential outcomes. First, we leverage a negative regression dependence between potential outcomes and censoring variables, which can be justified when defendants commit fewer crimes over time. Second, we discuss how one can continuously relax the independent censoring assumption. Both strategies, which do not impose exogenous censoring, lead to partial identification of the causal parameters and are discussed in Appendix (ref).

As highlighted above, explicitly handling right-censored outcomes introduces new econometric challenges. One question that naturally arises is whether empirical researchers can bypass (or ignore) these challenges and use standard causal inference tools. In the crime economics example, one common way to avoid such challenges is to restrict sample construction and focus on recidivism within a given time frame, say two years. This essentially changes the outcome of interest from time-to-recidivism to whether one recidivates within 2 years. The censoring issue is avoided if one follows all defendants for at least two years. Although this is convenient and generically valid, this procedure has drawbacks: (a) it changes the question of interest by changing the outcome variable, and (b) the choice of the cutoff (two years in this example) is arbitrary. For instance, it may be that punishment has no effect on recidivism within two years but then has an effect within three years (or one year).

These concerns are minor if one is interested only in recidivism within a known time frame. One potential way to assess whether this is the case is to ask ourselves: if censoring were not a concern, would we use time-to-recidivism or recidivism within two years as the outcome? If the answer is only the latter, then standard practice is justified. If not, we caution researchers that the conclusions of their analysis may be sensitive to the cutoff used for “binarization” of the time-to-recidivism outcome, and that directly using the time-to-recidivism outcome may yield different conclusions. See Appendix (ref) for a simple example of this.

If the conclusions of a study might be sensitive to the cutoff used for binarization, a natural next step would be for empirical researchers to consider a handful of cutoffs and demonstrate that the results are “robust” to the cutoff. We note that this happens often in practice. However, when doing this, researchers often ensure that the sample used for the entire analysis is not contaminated by compositional changes, which can lead to a loss of power at shorter horizons. To be more concrete, suppose that researchers considered three thresholds for the “binarization”: two, three, and five years. To ensure that the same defendants are analyzed across all transformed outcomes, researchers commonly restrict the sample to those they have observed in the data for at least 5 years. As a result, they drop available observations for shorter horizons, leading to a loss of statistical power. In contrast, our proposed method enables researchers to maximize the use of their data by leveraging longer observation periods for earlier-observed individuals and shorter observation periods for later-observed individuals. This advantage increases as we include more later-observed individuals in our sample, since it allows us to use more observations.

This discussion leads to the next natural question: what if we considered a large set of cutoffs and did not restrict the sample across cutoffs? Heuristically, one can interpret our $DMTE$ results as doing precisely that: avoiding choosing arbitrary cutoffs and considering recidivism within $y$ periods for a continuum of $y\in \mathbb{R}_+$. Our $QMTE$ results “transform” our $DMTE$ results so the underlying treatment effects are expressed in the same units as the time-to-recidivism outcome, leading to additional insights. Here, though, it is worth stressing that we explicitly tackle the censoring problem when considering the continuum of cutoffs by adapting the Frandsen2015's (Frandsen2015) reweighing approach to our context. Erroneously ignoring the censoring problem can indeed lead to misleading conclusions.

We show the appeal of our causal inference tools by evaluating the effect of fines and community service sentences (possibly combined with a “conviction marker”) as a form of punishment on time-to-recidivism in the State of São Paulo, Brazil, between 2010 and 2019.\footnote{São Paulo is the largest state in Brazil, with a population above 44 million people according to the Brazilian Census in 2022. Moreover, analyzing the impact of judicial policies on criminal behavior in this state is relevant due to its relatively high criminality. For example, according to São Paulo Public Safety Secretary, there were 12.56 murders, 1261.95 thefts, and 498.82 robberies per 100,000 inhabitants in 2023. Importantly, theft is one of the most common crimes in our sample.} Our treated group (punished group) contains the defendants who were fined or sentenced to community services due to either a conviction or a non-prosecution agreement, and our untreated group (unpunished group) contains defendants who were acquitted or whose cases were dismissed.\footnote{Our sample only contains cases whose punishment must be a fine or community service sentence. In 1998, a new law established that criminal charges whose maximum prison sentence is less than four years in the 1940 Criminal Law Code must be punished with a fine or a community service sentence from that point onwards. As we are particularly interested in the effect of fines and community service sentences, we focus on these specific criminal cases and define them as misdemeanor offenses.} To measure recidivism, we check whether the defendant's name appears in any criminal case within the sample period after the final sentence's date. More precisely, our outcome variable is the time between the final sentence and a subsequent criminal case. Since the sampling period is finite, the outcome variable is right-censored.

To deploy our proposed methodology, we need a continuous instrumental variable since we face endogenous selection into punishment.\footnote{{ If the researcher is comfortable with parametric assumptions, it is possible to use a discrete instrument as suggested by Brinch2017.}} We use the trial judge's leave-one-out rate of punishment (or “leniency rate”) as an instrument for the trial judge's decision Bhuller2019,Agan2021. Importantly, this instrumental variable is continuous with large support and is independent of the defendant's counterfactual criminal behavior because judges are randomly assigned to cases conditional on court districts according to state law in São Paulo. Our outcome data --- time-to-recidivism --- is right-censored by construction, requiring a methodology that accounts for this identification challenge.

Empirically, we find that the cross-district average QMTE functions for $.10, .15, .25, .40, .50$, and $.75$ quantiles are heterogeneous with respect to unobserved punishment resistance. The treatment effects are estimated to be sometimes positive and sometimes negative. More precisely, we find that people who would be punished by most judges (those with low punishment resistance) take longer to recidivate as a consequence of the punishment, while people who would be punished only by strict judges (high punishment resistance) recidivate at an earlier date than if they were not punished. This result suggests that designing sentencing guidelines that encourage strict judges to become more lenient could increase time-to-recidivism.

Lastly, we compare our results with methods that ignore that time-to-recidivism is right-censored. If one used the typical judge-fixed-effect regression, one would find that treatment increases the likelihood of short-term recidivism. If one used IV quantile regressions (ignoring censoring), one would find that treatment effects are slightly negative. In either case, the researcher would not be able to highlight heterogeneity as in the $DMTE$ and $QMTE$ functions, which show that treatment benefits some defendants while harming others. These differences highlight that our tools can indeed bring new insights to policy discussions.

Related literature: This article contributes to different branches of literature. Concerning its theoretical contribution, our work contributes to the literature on MTE by extending the MTE framework of Heckman2006 and carneirolee2009 to a setting with right-censored data. We also contribute to the literature on duration outcomes; see, e.g., Frandsen2015, tchetgen2015instrumental, Santanna2021, Delgado2022. None of these papers consider MTE-type parameters as we do. Among these, the closest work to ours is Frandsen2015, which considers the case where the censoring variable is observed and shows how one can identify distributional and quantile local treatment effects, assuming that censoring is exogenous. Our results can be interpreted as an extension of Frandsen2015 to the MTE framework, possibly allowing for endogenous censoring.

Concerning its empirical contribution, our work is inserted in the literature about the effect of fines and community service sentences on future criminal behavior; see, e.g., Huttunen2020, Giles2021, Possebom2022, and Lieberman2023.\footnote{This literature focuses on non-incarceration punishments for individuals who are already being prosecuted. The reader who is interested in the effect of alternative solutions to prosecution may check the recent work developed by Agan2021 and ShemTov2024. Readers interested in incarceration's effect may check the recent work of Rose2021, Humphries2023 and Kamat2023.} They all focus on binary variables indicating recidivism within a pre-specified period. Within these, as we build on his dataset, Possebom2022 is the closest to ours. However, his focus differs greatly from ours, and he does not handle duration outcomes as we do.

Organization of the paper: The rest of the paper is organized as follows. Section (ref) defines the causal parameters of interest. Section (ref) presents our model, discusses our identifying assumptions, and provides our identification results with a right-censored outcome variable. Section (ref) explains how to semiparametrically estimate the objects necessary to implement the identification strategy described in the previous section. Furthermore, Section (ref) discusses the empirical context, the data, and our results. Lastly, Section (ref) concludes. This paper also contains an online supporting appendix.

Causal Questions of Interest

In this section, we define our causal questions of interest in terms of treatment effect parameters. To make them concrete and intuitive, we use our empirical application as a running example. In this applied exercise, we study the effect of alternative sentences in the form of fines and community service on time-to-recidivism in the state of São Paulo, Brazil.

For each observation $i$, let $Y_i^*(1)$ be the potential outcome if treated and let $Y_i^*(0)$ be the potential outcome if untreated. In our empirical example, these unobservable variables capture, respectively, the potential time-to-recidivism if defendant $i$ were punished with a fine or community service either due to a conviction or a non-prosecution agreement, and the potential time-to-recidivism if defendant $i$ were not punished with a fine or community service either due to acquitall or dismissal.\footnote{Alternatively, a researcher might be interested in analyzing the effect of alternative sentences on total crime counts or severity-weighted offenses. We focus on time-to-recidivism because this variable facilitates the understanding of the censoring issues that are the focus of our proposed identification method. Importantly, total crime counts or severity-weighted offenses must be measured within a sample period (e.g., total crime counts within 2 or 3 years) and, consequently, also suffer from censoring. Adapting our censoring-focused methods to encompass these more complex outcome variables is an interesting future line of research. Additionally, a researcher might be interested in “true criminal behavior (even if not observed by the police)” as an outcome variable instead of “criminal behavior observed by the court system”. We discuss this case in Appendix (ref).} Defendant $i$'s treatment effect is therefore $\theta_i = Y_i^*(1) - Y^*_i(0)$. Ideally, we would like to learn $\theta_i$ for all defendants. However, that is very challenging (if not impossible) when we allow for (a) heterogeneous treatment effects across defendants and (b) whether a defendant is punished or not being related to $\theta_i$ (“essential heterogeneity” as defined by Heckman2006).

Due to these challenges, it is common for researchers to focus on aggregated summary measures of $\theta_i$, such as the average treatment effect among “compliers” Imbens1994, defined as $LATE = \mathbb{E}\left[Y^*(1) - Y^*(0)|\text{Compliers}\right]$ (see, e.g., Agan2021,Bhuller2019,Huttunen2020).\footnote{As discussed before, these papers use a different outcome of interest $Y^*$ that bypass the censoring issues we face. However, we can ignore these censoring issues while discussing our causal questions of interest (as this does not play a prominent role in it).} Although interesting and policy-oriented, such aggregated measures of causal effects are unsuitable for highlighting important types of treatment effect heterogeneity. In particular, these parameters cannot answer whether defendants with high punishment resistance (i.e., defendants who would only be punished by very strict judges) would, on average, take longer to recidivate if they were punished. The same goes for defendants with lower punishment resistance. These are the exact types of causal questions that interest us in this paper. We want to go beyond LATE-type parameters and provide a more detailed picture of how alternative punishments heterogeneously affect recidivism with respect to the defendant's (unobserved) punishment resistance, which we denote by $V$.\footnote{In general, our punishment resistance variable, $V$, is commonly called unobserved treatment resistance or latent cost of treatment.} Here, punishment resistance may capture the evidence gathered against the defendant and additional defendant-specific characteristics.

{

One can measure the causal effect of fines and community services on time-to-recidivism for defendants with a given punishment resistance using the notion of distribution and quantile treatment effects.\footnote{In Appendix (ref), we also explore how to define and identify types of average treatment effects. Importantly, censoring causes technical issues when defining parameters based on average effects, and we propose two alternative solutions that overcome those challenges.} All these causal parameters build on the Marginal Treatment Effects framework of Heckman2006 and can be used to answer complementary policy-relevant questions (Heckman2006, carneirolee2009). We now carefully define and interpret them.

For treatment status $d \in \{0, 1 \}$, let the distributional and quantile marginal treatment response functions be defined as

eqnarray[eqnarray omitted — 296 chars of source]

where $y \in \mathbb{R}_+$, $\tau \in (0,1)$ and $v \in \left[0, 1\right]$. All these counterfactual parameters have a clear interpretation. For instance, $DMTR_{d}\left(y, v\right)$ gives the proportion of defendants with punishment resistance $v$ who would have already recidivated after $y$ periods since the court's final ruling if they were punished ($d=1$) or not ($d=0$). Analogously, $QMTR_{d}\left(\tau, v\right)$ provides the $\tau$-th quantile of the time-to-recidivism under treatment $d$, among defendants with punishment resistance $v$.

Based on these counterfactual objects, it is straightforward to define the Distributional and Quantile Marginal Treatment Effect functions:

eqnarray[eqnarray omitted — 235 chars of source]

Positive values of the QMTE function indicate that punishment by fines and community services increases the defendant's time-to-recidivism compared to no punishment (so treatment is working as intended). On the other hand, positive values of the DMTE function indicate that punishment by fines and community services increases the proportion of defendants who recidivate by time $y$ compared to no punishment (so treatment is not working as intended). For policy effectiveness in our context, positive values of QMTE are “good”, while negative values of DMTE are “good”.

remarkWhen analyzing the impact of judicial decisions on recidivism, many papers focus on distributional marginal treatment effects for a small set of values of $y$.\footnote{See, e.g., Agan2021, Bhuller2019, Giles2021, Huttunen2020, and Possebom2022.} Here, we entertain the possibility of moving beyond small set of horizons by considering a continuous set of cutoffs $y$'s or by focusing on different target parameters, such as the quantile marginal treatment effects on time-to-recidivism. See Appendix (ref) for a simple illustration of the appeal of our approach compared to the traditional “small-set-of-horizons” approach.

}

Econometric framework and identification results

We face some challenges in identifying the causal parameters of interest described in the previous section. As usual, potential outcomes are only (potentially) observed under one treatment status, i.e., $Y^{*}_i = Y^{*}_i(1) \cdot D_i + Y^{*}_i(0) \cdot \left(1 - D_i\right)$, where $D_i$ is the treatment indicator variable. Furthermore, our outcome of interest is subject to right-censoring, implying that we do not always observe $Y^{*}$ but rather observe $Y_i = \min \{Y^*_i, C_i\}$, where $C$ is the censoring variable. In the case of draws, we assume that $Y^*_i$ happens before $C_i$, as is customary in survival analysis. Finally, we also expect that treatment statuses are related to the potential outcomes and possibly related to the censoring variable.

To tackle all these issues, we build on the MTE framework of Heckman2006 and extend it to handle duration outcomes. Toward this end, we consider a threshold-crossing treatment selection model

align[align omitted — 103 chars of source]

where $Z$ is an observed instrumental variable (with support $\mathcal{Z} \subset \mathbb{R}$), $C$ is an observed censoring variable (with support $\mathcal{C}\subset \mathbb{R}_+$), and $V$ is a latent heterogeneity variable that captures the unobserved treatment resistance. The function $P: \mathcal{Z} \times \mathcal{C} \rightarrow \mathcal{P}$ is unknown and captures the willingness to take the treatment for each value of $Z$ and $C$. Importantly, our treatment selection model (ref) imposes monotonicity Vytlacil2002.

To better understand Equation (ref), let us go back to our empirical context and explain each component of it. Our instrumental variable $Z$ is a measure of the trial judge's leniency, which arguably does not affect time-to-recidivism other than through the judge's decision to punish or not. Our censoring variable $C$ captures the time between the defendant's sentence date and the end of our sampling period, and it is observed for all defendants. $C$ can also capture seasonality patterns, as it is fully determined depending on the sentence's date. The function $P$ captures the trial judge's punishment criteria, and it allows trial judges to update their punishment criteria over time Bhuller2022, as it includes $C$ as an argument.\footnote{Since the end of the sample period is the same for all defendants, we can equivalently express the decision rule for $D$ in terms of the sentencing date $S_{i}$, where the sentencing date equals the fixed calendar date when the sample period ends $\left(\overline{T}\right)$ minus the censoring variable $C_{i}$, i.e., $S_{i} = \overline{T} - C_{i}$. This equality implies that the propensity score can be equivalently expressed as a function of the sentencing date since $P(Z_i,C_i)=P(Z_i,T - S_i) \equiv \tilde{P}(Z_i,S_i)$. This way of writing the propensity score function highlights that trial judges may update their punishment criteria over time in our setting.}\textsuperscript{,}\footnote{We emphasize that we allow function $P$ to depend on the censoring variable, but we do not impose that the censoring variable has an impact on this function. This type of connection is testable through the first-stage regression.} Finally, the variable $V$ can be interpreted as unobserved punishment resistance, and it captures, among other things, the amount of criminal evidence in the defendant's favor. The higher the $V$, the less likely the defendant will be punished, all else equal. As already discussed, $Y^{*}$ captures the length of time between the defendant's sentence date and her next criminal case's starting date, and $Y$ is the minimum of $Y^*$ and time from the sentence's date to the end of our sampling period, $C$.

Assumptions

In our setup, the available data for the researcher are $\left\{Y_i,C_i,D_i,Z_i\right\}_{i=1}^{n}$, while $Y^{*}_i\left(0\right)$, $Y^{*}_i\left(1\right)$, $Y^{*}_i$ and $V_i$ are latent variables. Henceforth, we assume that $\left\{Y_i,C_i,D_i,Z_i\right\}_{i=1}^{n}$ are independently and identically distributed as $\left(Y,C,D,Z\right)$. For simplicity, we drop exogenous covariates from the model and focus on the case with a single instrument. All results derived in the paper hold conditionally on covariates and can be extended to the case with multiple instruments.\footnote{In our empirical application, the conditioning covariates are a full set of court district indicators.} Since we deal with a time-to-event outcome, $Y$ is non-negative by construction.

{ In what follows, we present a set of five assumptions (Assumptions (ref)-(ref)) that will allow us to point-identify the DMTE and the QMTE functions across a range of thresholds and quantile points. These assumptions are related to those imposed by Heckman2006 and Frandsen2015 and involve assuming that censoring is not related to the potential outcomes $Y^*(d)$.

We now state our five assumptions and contextualize each of them to our empirical setup.

}

assumption[Random Assignment] Conditional on C, the potential outcomes $Y^{*}\left(0\right)$, $Y^{*}\left(1\right)$ and $V$ are independent of the instrument $Z$, i.e., $\left. Z \protect\mathpalette{\protect\independenT}{\perp} \left(Y^{*}\left(0\right), Y^{*}\left(1\right), V\right) \right\vert C.$

Assumption (ref) is an exogeneity assumption and is common in the literature about instrumental variables with censored outcomes Frandsen2015. In our empirical application, this assumption holds conditional on the court district because trial judges are randomly assigned to cases within each court district. { Importantly, judges are assigned to criminal cases based on a computer algorithm that creates a lottery of judges and there is no suspicion in the press that this software is manipulated.}

Note also that Assumption (ref) allows the instrument to depend on the censoring variable. In our empirical application, this flexibility is useful because the trial judge's punishment rate may depend on the case's sentence date if judges who entered the Judiciary more recently are more lenient than judges who retired at the beginning of our sampling period, for example.

assumption[Propensity Score is Continuous] Conditional on $C$, $P(z,c)$ is a nontrivial function of $z$ and the random variable $\left. P\left(Z,c\right) \right\vert C = c$ is absolutely continuous in $Z$, with support given by an interval $\mathcal{P} \coloneqq \left[\underline{p}, \overline{p}\right]$ for any $c \in \mathcal{C}$.\footnote{The assumption that $\mathcal{P}$ is an interval is made for notational simplicity. All the proofs can be easily extended to the case where $\mathcal{P}$ is a set with a non-empty interior.}

Assumption (ref) is a rank condition, intuitively imposing that the instrument is locally relevant. In addition, we implicitly assume that the support of the propensity score does not vary with the value of $C$. In our application, this implicit assumption is plausible because the judges are mostly the same over time. Furthermore, the judge's lenience rate has a good amount of variation { since we observe 525 judges who use many different punishment criteria (Figure (ref))}.\footnote{{ Since our main identification result is fully nonparametric, Assumption (ref) requires sufficient (continuous) variation in the instrument, which induces continuous variation in the propensity score, which generates the variation we need on the outcomes to identify the effects for every value of $V$. Nevertheless, our proposed estimation procedure is semi-parametric and relies on series approximations for the propensity score. If one treats the series as fixed (meaning that the first stage is parametric), our approach allows for discrete instruments Brinch2017. The extrapolation and interpolation induced by the parametric second stage permit $Z$ (and $P$) to vary discretely while still being used to identify the effects of interest.}}

assumption[$V$ is continuous] The distribution of the latent heterogeneity variable $V$ conditional on $C$ is absolutely continuous with respect to the Lebesgue measure.

Assumption (ref) is a regularity condition that allows us to normalize the marginal distribution of $\left. V \right\vert C$ to be the standard uniform. Consequently, we can write $P\left(z, c\right) = \mathbb{P}\left[\left. D = 1 \right\vert Z = z, C = c\right]$ for any $z \in \mathcal{Z}$ and $c \in \mathcal{C}$.

assumption[Overlap] Conditional on $C$, all treatment groups exist, i.e., $\mathbb{P}\left[\left. D = d \right\vert C = c \right] \in \left(0, 1\right)$ for any $d \in \left\lbrace 0, 1 \right\rbrace$ and any $c \in \mathcal{C}$.

Assumption (ref) is a regularity condition about overlap. It imposes that there is strictly positive mass in both treatment groups for every value of the censoring variable.

{

assumption[Random Censoring] The censoring variable is independent of the uncensored potential outcomes given the latent heterogeneity $V$, i.e., $\left. C \protect\mathpalette{\protect\independenT}{\perp} \left(Y^{*}\left(0\right), Y^{*}\left(1\right)\right) \right\vert V.$

Assumption (ref) is an exogeneity assumption and is common in the literature about duration outcomes Frandsen2015,Delgado2022. When combined with Assumption (ref), Assumption (ref) implies that $C$ is unconditionally independent of the uncensored potential outcomes, i.e., $C \protect\mathpalette{\protect\independenT}{\perp} \left(Y^{*}\left(0\right), Y^{*}\left(1\right)\right)$. In our empirical application, this restriction imposes that the case's sentence date is independent of the defendant's decision to commit another crime in the future, { i.e., potential recidivism is stationary.}

Importantly, Assumption (ref) requires that controlling for $V$ accounts for all sources of endogeneity coming through the censoring variable. This assumption can be restrictive, as endogeneity may persist even after controlling for latent heterogeneity. For example, in our empirical setting, potential recidivism may not be stationary if legislation or inputs to the production of Justice change over time. These inputs may include the number of police officers, judges, or public defenders. Appendix (ref) has a detailed discussion about these inputs in our empirical application.

If the researcher believes that Assumption (ref) is too strong in a particular application, she can use alternative assumptions that are sufficient to partially identify the distributional marginal treatment effect and some quantile marginal treatment effects when the outcome variable is right-censored. We discuss two alternative partial identification strategies in Appendix (ref).\footnote{Moreover, Appendix (ref) discusses the costs of not imposing restrictions on the censoring mechanism, and Appendix (ref) cautions about simply treating the censoring variable as an additional covariate.}

}

Identification

We present our point-identification results that rely on Assumptions (ref)-(ref).

First, define $\gamma_{d}(y,v,c) = \dfrac{d}{d v}\mathbb{P}\left[\left. Y\leq y, D = d \right\vert P\left(Z,C\right) = v, C = c\right]. $ We now state our main identification result: point-identification of the $DMTR$ functions.

propositionSuppose that Assumptions (ref)-(ref) hold. Then, for any $d \in \left\lbrace 0, 1 \right\rbrace$, $y < \gamma_{C}$ and $v \in \mathcal{P}$, $DMTR_{d}\left(y, v\right) = \left(2 d - 1\right) \cdot \mathbb{E}\left[\gamma_{d}(y,p,C) | P(Z,C) = v, C > y\right],$ { where $\gamma_{C} \coloneqq \inf\left\lbrace c \in \overline{\mathbb{R}} \colon \mathbb{P}\left[C \leq c\right] = 1 \right\rbrace$ is the upper bound of the support of the censoring variable $C$.}
proofSee Appendix (ref).

The above proposition shows how we can point-identify the distributional marginal treatment response for a given unobserved treatment resistance $v$. It involves first taking the derivative of the conditional joint distribution of the realized outcome $Y$ and treatment status $D$ given the propensity score $P=v$ and the censoring variable being above $y$ ($C = y + \delta$ for $\delta > 0$) with respect to $v$, and then integrating over all values $\delta > 0$ such that the $y+\delta$ remains in the support of the censoring variable $C$. Differently from the results in carneirolee2009, we need to tackle the right-censoring problem, which manifests in our results by having to condition on $C=y+\delta$, so that $C > y$, and then integrating over $\delta$.

Furthermore, our results are specific to the $DMTR_d(y,v)$ function, and not for a generic transformation of $Y^*(d)$, say $G(Y^*(d))$ as in carneirolee2009. This follows from the fact that we may not be able to identify the $DMTR_d(y,v)$ over all values of $y$ in the support of $Y^*(d)$, as a consequence of the censoring problem. Having said that, there are several functions that we can actually nonparametrically point-identify without additional restrictions and under standard regularity conditions, including the QMTE functions for a range of quantiles. We state these results{, which use only Assumptions (ref)-(ref) and a regularity condition, as a corollary for convenience. The proof is a direct consequence of the previous proposition and the definition of quantiles.

corollarySuppose that Assumptions (ref)-(ref) and Assumption (ref) listed in the Appendix (ref) hold. Then, the $QMTE\left(\tau,v\right)$ function (Equation (ref)) is point-identified for any $v \in \mathcal{P}$ and $\tau \in \left(0, \overline{\tau}\left(v\right)\right)$, where $\overline{\tau}\left(v\right) \coloneqq \min\left\lbrace \overline{\tau}_{0}\left(v\right), \overline{\tau}_{1}\left(v\right) \right\rbrace$ and $\overline{\tau}_{d}\left(v\right) \coloneqq DMTR_{d}\left(\gamma_{C}, v\right)$ for any $d \in \left\lbrace 0, 1 \right\rbrace$.

} Notice that the right-tail of the $DMTR_d(\cdot, v)$ may be differentially affected by the censoring problem, implying that $\overline{\tau}_{1}\left(v\right)$ may be different from $\overline{\tau}_{0}\left(v\right)$. As a consequence, we can only identify the $QMTE\left(\tau,v\right)$ over the common range of identified quantiles among treated and untreated units.

Estimation and inference

In this section, we provide an algorithm on how to semiparametrically estimate the DMTE and QMTE functions based on the identification results described in Proposition (ref) and Corollary (ref). Section (ref) first discusses that, in our empirical setting, judges are randomly assigned within court districts and then explains the consequences of this mechanism for our identification and estimation procedures. Afterwards, we provide practical and formally justified estimation and inference procedures based on a semi-parametric class of estimators for the nuisance functions (Section (ref)).\footnote{We refer the reader to Appendix (ref) for a generic procedure to estimate the marginal treatment effect functionals while remaining agnostic about the type of estimators used to estimate the nuisance functions.} Lastly, Section (ref) discusses how to aggregate our target parameters across covariate values.

Conditioning on Covariates: Judges are randomly assigned within court districts

All the results derived in Section (ref) could be interpreted as conditional on covariates. This possibility is particularly relevant in our application since random judge assignment is legally guaranteed within the court district where the crime occurred. Thus, conditioning on a full set of dummy variables indicating each one of the court districts in São Paulo is fundamental in our empirical setting.

Consequently, we want to estimate court-district-specific $DMTE$ and $QMTE$ functions capturing the effect of alternative sentences on time-to-recidivism. Since non-parametric estimation in the presence of covariates can be demanding due to the curse of dimensionality, we propose an easy-to-implement semiparametric estimator in Section (ref).

However, in our empirical setting, there are 193 court districts in the State of São Paulo. Consequently, our semi-parametric procedure estimates 193 $DMTE$ functions and 193 $QMTE$ functions. To facilitate interpretation of these functions, we propose summary functions that aggregate across covariate values. We do so using the proportion of cases per court district as weights and explain the details of this procedure as well as its asymptotic validity in Section (ref).

Semiparametric estimation and inference procedures

This section provides a specific procedure for semiparametrically estimating marginal treatment effect functionals. We have data on $\left\lbrace Y_{i}, C_{i}, D_{i}, Z_{i}, X_{i}' \right\rbrace_{i = 1}^{n}$.\footnote{Appendix (ref) explains the technical details behind our estimation procedure.} In the context of our application, $X$ is a set of 193 court district indicators.

Similar to carneirolee2009, we model $P(Z,C,X) \coloneqq \mathbb{E}\left[D|Z,C,X\right]$ using an additive partially linear series regression

equation[equation omitted — 111 chars of source]

where $(\alpha_0, \alpha_X', \alpha_C)'$ are unknown finite-dimensional parameters, and $\varphi$ is an unknown (infinite-dimension) function. In our context, the partially linear additive specification (Equation (ref)) allows one to pool information from different court districts and run a single propensity score model for all courts.\footnote{Alternatively, one could use a semiparametric logit model. In Appendix (ref), we show that our regularity conditions hold for this model as well. The same is true for a fully nonparametric series estimator.}

We pick a polynomial basis function to approximate $\varphi(\cdot)$, $\psi^L(z) = \left(z, z^2, \dots, z^L\right)'$, though other options, such as B-splines, are also possible. Note that all the series coefficients can be estimated via ordinary least squares to obtain an estimate $(\widehat{\alpha}_0, \widehat{\alpha}_X', \widehat{\alpha}_C, \widehat{\alpha}_Z))'$ and compute:\footnote{See Appendix (ref) for details.}

equation[equation omitted — 197 chars of source]

Since, in finite samples, $\widetilde{P}_i$ might be negative or larger than one, we follow carneirolee2009 and use the trimmed version of $\widetilde{P}_i$ as our estimator,

equation[equation omitted — 218 chars of source]

for a sufficiently small $\epsilon$.\footnote{ Our application uses $\epsilon = 0.01$, though this is not material as only one observation is trimmed.}

Next, we move into the estimation of the conditional distribution function of $Y \cdot \mathbf{1}\left\lbrace D = d \right\rbrace$ given $P, C, X$ for $d \in \{0,1\}$. Here, we impose the following model:\footnote{See Appendix (ref) for additional details.}

eqnarray[eqnarray omitted — 298 chars of source]

where $\theta_0(\cdot, \cdot) = (\beta_0 \left(\cdot,\cdot\right), \beta_X \left(\cdot,\cdot\right)', \beta_C \left(\cdot,\cdot\right), \beta_P \left(\cdot,\cdot\right))' \mapsto \Theta \subseteq \mathbb{R}^{3+k_X}$ is a vector of nonparametric functions, $k_X$ is the dimension of $X$, and $\Lambda$ is a known link function.\footnote{This class of distribution regression models nests and extends many traditional duration models such as the proportional hazard model and the accelerated time model. See Delgado2022 for a discussion.} For concreteness, we focus on a logistic link function, $\Lambda(\cdot) = \exp(\cdot)\big{/}(1 + \exp(\cdot))$.

To estimate these unknown functions, we first need to acknowledge that the propensity score $P_i$ is not observed. However, we can use the “generated regressor” $\widehat{P}_i$ from Equation (ref). Once we replace $P_i$ with $\widehat{P}_i$, we can then leverage the insights of Foresi1995 and Chernozhukov2013a by noticing that, for a fixed $y$ and $d$, Equation (ref) is a binary regression problem. Consequently, we can pointwise estimate these parameters by maximizing the (feasible) conditional likelihood function and obtain a distribution regression estimate of $\theta_{0}(y,d)$ denoted by $\hat{\theta}\left( y,d\right)$.\footnote{ Computing the distribution regression estimators for many $(y,d)$ points only requires running a sequence of binary regressions.}

Next, note that the derivative of $\Gamma$ (Equation (ref)) with respect to $P$ can computed in closed-form for each $(y,d)$,

eqnarray[eqnarray omitted — 167 chars of source]

where we explored that $\Lambda$ is the logistic link function. Denote the estimated fitted values of $\gamma_{d}(y,v,c,x)$ by

eqnarray[eqnarray omitted — 172 chars of source]

where the distribution regression coefficients are the components of $\hat{\theta}\left( y,d\right)$ , and

eqnarray*[eqnarray* omitted — 347 chars of source]

Using Equation (ref), we can estimate $DMTR_d (y,v,x) \coloneqq \mathbb{P}\left[\left.Y^*(d) \leq y \right\vert V = v, X=x\right].$ To do so, let $n_{d,x,y} = \sum_{i=1}^n \mathbf{1}\{D_i=d, X_i=x, C_{i} > y \}$ denote the sample size with treatment status $d$, covariate value $x$, and censoring variable above $y$. Our proposed estimator for $DMTR_d (y,v,x)$ is given by\footnote{Since $\widehat{DMTR}_d (y,v,x)$ is an estimator for a conditional distribution, it needs to be non-decreasing in $y$ for all $(d,v,x) \in \{0,1\} \times \mathcal{P} \times \mathcal{X}$. However, this may not be the case in finite samples. We recommend using the rearrangement procedure of Chernozhukov2009b.}

eqnarray[eqnarray omitted — 184 chars of source]

Based on Equation (ref), we can then estimate $DMTE(y,v,x) \coloneqq DMTR_1(y,v,x) - DMTR_0(y,v,x)$ using

eqnarray[eqnarray omitted — 117 chars of source]

Analogously, one can estimate $QMTE(\tau,v,x)$ functionals using

eqnarray[eqnarray omitted — 120 chars of source]

where $\widehat{QMTR}_d(\tau,v,x) = \inf\{y\in \mathbb{R}_+ \colon \widehat{DMTR}_d(y,v,x)\geq \tau \}$.

We summarize all these estimation steps in Algorithm (ref) in Appendix (ref).

The next theorem establishes the large-sample properties of our proposed estimators. We defer all the regularity conditions to the appendix to streamline the presentation. Let $\overline{\tau}\left(v,x\right) \coloneqq \min\left\lbrace \overline{\tau}_{0}\left(v,x\right), \overline{\tau}_{1}\left(v,x\right) \right\rbrace$ and $\overline{\tau}_{d}\left(v,x\right) \coloneqq DMTR_{d}\left(\gamma_{C}, v,x\right)$ for any $d \in \left\lbrace 0, 1 \right\rbrace$.

theoremSuppose that Assumptions (ref)-(ref) and Assumptions (ref)-(ref) listed in Appendices (ref) and (ref) hold. Then, as $n \rightarrow \infty$, \begin{enumerate} • for each fixed $y < \gamma_{C}$, $v \in \mathcal{P}$, and $x \in \mathcal{X}$, $\sqrt{n}\left( \widehat{DMTE}(y,v,x) - {DMTE}(y,v,x) \right)$ ${\overset{d}{\rightarrow}} N(0,V_{y,v,x}^{dmte}),$ with $\widehat{DMTE}(y,v,x)$ in Equation (ref) and $V_{y,v,x}^{dmte}$ in Appendix (ref). • for each fixed $ \tau \in (0,\overline{\tau}\left(v,x\right))$, $v \in \mathcal{P}$, and $x \in \mathcal{X}$, $\sqrt{n}\left( \widehat{QMTE}(\tau,v,x) - {QMTE}(\tau,v,x) \right)$ ${\overset{d}{\rightarrow}} N(0,V_{\tau,v,x}^{qmte}),$ with $\widehat{QMTE}(\tau,v,x)$ in Equation (ref) and $V_{\tau,v,x}^{qmte}$ in Appendix (ref). \end{enumerate}

Theorem (ref) follows from first deriving the influence function of the $DMTR$ functions (Equation (ref)), paying particular attention to quantifying the estimation effect arising from replacing the true propensity score with the estimated one. After this step, all the results follow from the functional delta method and the continuous mapping theorem. The proof strategy is similar to Rothe2009.

Although Theorem (ref) indicates that one can potentially conduct inference using plug-in estimates of the variance, this procedure would involve estimating additional nuisance functions and could be cumbersome in practice. To avoid this issue, we propose using a weighted bootstrap procedure as in Ma2005 and Chen2009. This bootstrap procedure is very straightforward to implement and the details can be found in Algorithm (ref) in Appendix (ref).

Covariate Aggregation

In our empirical setting, there are 193 court districts in the State of São Paulo. Consequently, our semi-parametric procedure estimates 193 $DMTE$ functions and 193 $QMTE$ functions. To facilitate interpretation of these functions, we propose summary functions that aggregate across covariate values.

To do so, we aggregate the court-district-specific DMTE functionals across court districts using the proportion of cases per court district as weights.\footnote{Another possible covariate aggregation and its limitations are discussed in Appendix (ref).} Let $w_{x} = \mathbb{P}(X=x)$ denote the probability of a covariate $X$ taking the value $x$, which, in our case, denotes the true proportion of cases assigned to a court district $x$. Let $\widehat{w}_{x} = n^{-1}\sum_{i=1}^n \mathbf{1}\{X_i=x\}$ be the plug-in estimator of $w_{x}$.

For each $d \in \{0,1\}$, $y \in \mathcal{Y}$ and $v\in \mathcal{P}$, let $$DMTR_d^{avg}(y,v) = \mathbb{E}\left[DMTR_d (y,v,X)\right] = \sum_{x\in \mathcal{X}}w_x~DMTR_d (y,v,x).$$

Analogously, let $DMTE^{avg}(y,v) = DMTR_1^{avg}(y,v) - DMTR_0^{avg}(y,v)$ and $QMTE^{avg}(\tau,v) = QMTR_1^{avg}(\tau,v) - QMTR_0^{avg}(\tau,v)$, where $ QMTR_{d}^{avg}\left(\tau, v\right) \coloneqq \inf\{y\in \mathbb{R}_+ \colon DMTR_d^{avg}(y,v)\geq \tau \}. $ All these functionals can be straightforwardly estimated using functionals of $ \widehat{DMTR}_d^{avg}(y,v) = \sum_{x\in \mathcal{X}}\widehat{w}_x~\widehat{DMTR}_d (y,v,x),$ with $\widehat{DMTR}_d (y,v,x)$ as in Equation (ref), just like in Equations (ref)-(ref). Their large-sample properties follow from the delta method and are summarized in the following corollary.

corollarySuppose that Assumptions (ref)-(ref) and Assumptions (ref)-(ref) listed in Appendices (ref) and (ref) hold. Then, as $n \rightarrow \infty$, \begin{enumerate} • for each fixed $y < \gamma_{C}$, and $v \in \mathcal{P}$, $\sqrt{n}\left( \widehat{DMTE}^{avg}(y,v) - {DMTE}^{avg}(y,v) \right) {\overset{d}{\rightarrow}} N(0,V_{y,v}^{dmte, {avg}}).$ • for each fixed $ \tau \in (0,\overline{\tau}\left(v,x\right))$, and $v \in \mathcal{P}$, $\sqrt{n}\left( \widehat{QMTE}^{avg}(\tau,v) - {QMTE}^{avg}(\tau,v) \right)$ ${\overset{d}{\rightarrow}} N(0,V_{\tau,v}^{qmte, avg}).$ \end{enumerate}

It is also straightforward to construct a weighted-bootstrap confidence interval for these functionals by using $\widehat{w}^*_{x} = n^{-1}\sum_{i=1}^n \omega_i~\mathbf{1}\{X_i=x\}$ as weights for the MTE functionals, where $\omega_{i}$ is defined in Algorithm (ref). We omit a detailed description to avoid repetition.

Effect of alternative sentences on time-to-recidivism

Our empirical application answers the question: “Do alternative sentences such as fines and community service impact time-to-recidivism in São Paulo, Brazil?”. To do so, we start by explaining our empirical context and data in Subsection (ref). Then, Subsection (ref) defines the variables of interest and argues that long-run recidivism is a relevant problem in Brazil. Lastly, Subsection (ref) presents the results of our empirical analysis using our proposed tools. Subsection (ref) also compares our methods against more traditional methods, highlighting the ability of our proposed tools to unpack treatment effect heterogeneity. For the interested reader, we assess the plausibility of our identifying assumptions in Appendix (ref), and provide robustness checks in Appendices (ref) and (ref).

Empirical Context and Data

To study the effect of alternative sentences in the form of fines and community service on time-to-recidivism, we collect data from all criminal cases brought to the Justice Court System in the State of São Paulo, Brazil, between January 4\textsuperscript{th}, 2010, and December 3\textsuperscript{rd}, 2019.\footnote{See Appendix (ref) for an overview of the data-construction.} According to a Brazilian law from 1998, criminal charges whose maximum prison sentence is less than four years in the 1940 Criminal Law Code must, from that year onwards, be punished with a fine or a community service sentence if the defendant is found guilty. As we are particularly interested in the effect of these alternative sentences, we focus on these specific criminal cases { and define them as misdemeanor offenses}. We also restrict our sample to cases that started between 2010 and 2017. Based on these restrictions, the most common types of crime in our sample are theft and domestic violence.

There are 332 court districts in the state of São Paulo. Criminal complaints are analyzed by a trial judge working at the court with geographic jurisdiction over the location of the alleged offense.\footnote{{ In Brazil, instead of being elected, judges are appointed for life based on their performance in a civil service exam and frequently serve as judges until retirement Laneuville2024breaking. Furthermore, they make decisions about conviction, sentence type and sentence intensity in all cases but “crimes against life” (murder, attempted murder, manslaughter, incentivizing or assisting suicide, and abortion). Importantly, our sample does not contain “crime against life” cases. Therefore, all our cases are entirely decided by appointed judges.}} Moreover, there are 862 trial judges during our sample period. We keep 642 judges who analyzed more than 20 cases out of these. In court districts that have more than one judge, the case is randomly allocated to one of the judges { according to a computer algorithm that creates a lottery of judges}. Of the 332 court districts in our sample, 193 have more than one judge who analyzed more than 20 cases. Given that our econometric procedure explores the random allocation of judges to criminal cases and their different leniency levels, we restrict our attention to court districts with more than two judges who analyzed 20 cases or more. After imposing these two restrictions, our sample has 525 trial judges from 193 different court districts, handling 43,468 cases in total.\footnote{{ We treat each case-defendant pair as a separate observation. Consequently, defendants involved with more than one case will appear more than once in our dataset. Appendix (ref) analyzes the question of repeated offenders in detail.}} Appendix (ref) discusses the distribution of judges per court district and cases per judge.

Defining the Variables of Interest

In our dataset, we observe the defendant's full name, the defendant's court district, the case's starting date, the assigned trial judge's full name, the case's final ruling, and the case's final ruling's date. All our variables of interest will be constructed from these pieces of information. Henceforth, let $X$ denote the full set of court district dummies, which will play the role of covariates in our analysis.

Let us start with our treatment variable, $D$, which denotes the final ruling in the case. Defendants who were fined or sentenced to community services because they were either convicted or signed a non-prosecution agreement according to the final ruling in their case belong to our treatment group, $D=1$. Defendants who were acquitted or their cases were dismissed according to the final ruling in their case belong to our comparison group, $D=0$.

Our outcome of interest, $Y^*$, is the “time-to-recidivism”, i.e., the number of days it takes for a defendant to appear in court once again after the case's final ruling's date. Here, note that our outcome of interest is a duration variable and that some defendants may not recidivate by the end of our sampling period, though they may recidivate later. Putting it simply, we do not always observe $Y^*$, but rather observe a right-censored version of $Y^*$: $Y=min(Y^*, C)$, where $C$ is a right-censoring variable.\footnote{Appendix (ref) presents summary statistics about our outcome variable.} In our context, $C$ is the follow-up period for each defendant, i.e., the number of days from their case's final ruling date to December 3\textsuperscript{rd}, 2019.

Besides the censoring problem, it is important to be explicit about how we define recidivism. In this paper, a defendant $i$ in a case $j$ recidivated by the end of our sample if and only if defendant $i$'s full name appears in a case $\bar{j}$ whose starting date is after case $j$'s final sentence's date.\footnote{To match defendants' names across cases, we follow the same procedure as in Possebom2022 and define a fuzzy match if the similarity between full names in two different cases is greater than or equal to 0.95 using the Jaro–Winkler similarity metric. Appendix (ref) discusses the number of words per defendant's name.}\textsuperscript{,}\footnote{In this manuscript, our definition of recidivism is being prosecuted for another crime. In Appendix (ref), we discuss how to combine the methods proposed here with the methods proposed by Bartalotti2021 to identify the effect of judicial decisions on committing a crime even if the police are not able to observe all crimes. A similar issue arises when defendants move to other states and commit crimes in different jurisdictions. Even though these out-of-state recidivism events are not captured by our outcome variable, we do not believe they represent a relevant concern because out-migration in São Paulo is low, accounting for less than 2% of the state's population according to data from the 2022 Census.} Then, we measure our outcome variable as the number of days between case $j$'s final ruling's date and case $\bar{j}$'s starting date.\footnote{Case $\bar{j}$ can be about any type of crime, including more severe crimes whose maximum sentence is over four years, while case $j$ has to be about a crime whose maximum sentence is at most four years.} If defendant $i$ did not recidivate by the end of the sampling period, then $Y = C$.

At this stage, it is important to stress that we are not adopting a more restrictive notion of “short-run” recidivism based on a fixed period, say two years, which could potentially allow us to “ignore” the censoring problem. Instead, we focus on time-to-recidivism directly, which, in our view, entails some important advantages. For instance, we do not need to pick a threshold to define (short-run) recidivism arbitrarily. Doing so may lead to potentially sensitive conclusions, as illustrated in an example in Appendix (ref).

However, if almost all defendants who recidivate do it in the short run, then focusing on short-run measures would be sufficient. But this is an empirical matter and should be handled as such. To assess if this is the case in our data, Figure (ref) displays estimates of the right tail of the probability distribution function (PDF) of the uncensored potential outcome ($Y^{*}$) among cohorts defined based on the censoring variable. These descriptive results reveal that, in the case of the state of São Paulo, a non-negligible share of defendants have their first recidivism event in their fifth, sixth, or seventh year after their sentence's date, implying that analyzing long-term recidivism is practically relevant. Consequently, we must tackle the censoring problem directly. See Appendix (ref) for additional motivations for leveraging time-to-recidivism as a key outcome of interest from a welfare maximizer decision-maker perspective.

figure[figure omitted — 1,504 chars of source]

As it will be clear in the next sections, our causal inference procedures leverage the availability of an instrumental variable $Z$ with large support. In our context, the instrument $Z$ is the trial judge's leniency rate. This variable equals the leave-one-out rate of punishment for each trial judge, where the defendant's own decision is excluded from this average.\footnote{{ Similarly to DiTella2013 and Bhuller2019, we use the simple leave-one-out rate of punishment for each trial judge as our instrumental variable. Alternatively, we could have used the residualized leave-one-out rate of punishment as done by Agan2021, who remove court-district averages before computing each decision maker's rate. We choose to use the simple leave-one-out rate because we already include court-district fixed effects in our regression specifications, and each judge analyzes many cases as shown in Figure (ref).}} We ensure that the minimum and maximum values of the $Z$ are the same across both treatment arms to enforce better overlap properties.

Empirical results

In this section, we present our empirical results.\footnote{Appendix (ref) contains information about the first stage of our estimation procedure (Equation (ref)), which relates how censoring, court district dummies, and the judge's leniency rate affect the final ruling of the case.} To estimate the DMTE and QMTE functions in our empirical application, we flexibly account for court district fixed effects. More precisely, we estimate 193 district-specific functions for each of our treatment effect parameters (Theorem (ref)). Although very flexible, this strategy makes it challenging to concisely report summary results. The way we proceeded was to average these district-specific functions over court districts using the proportion of cases per court district as weights, as in Corollary (ref). We report the average DMTE function in Section (ref) and the average QMTE function in Section (ref).\footnote{{ Moreover, Appendix (ref) discusses whether our main results are robust to including case-processing time as an additional covariate.}} Moreover, we compare our proposed methods against standard methods in the literature in Section (ref).

Estimated DMTE function

Figures (ref) and (ref) shows the estimated average $DMTE\left(y,\cdot\right)$ functions for $y \in \left\lbrace 1, 2, \ldots, 8 \right\rbrace$, where instead of measuring time-to-recidivism in days we measured it in years (to enhance readability).\footnote{In our data, we observe time-to-recidivism in days. To illustrate the readability improvements of writing the $DMTE$ function in years instead of days, we focus on one value of the time-to-recidivism variable. The $DMTE\left(y,\cdot\right)$ function when $y = 2$ shows the distributional marginal treatment effect given by $\mathbb{P}\left[\left.Y^*(1) \leq 2 \cdot 365 \text{ days } \right\vert V = v\right] - \mathbb{P}\left[\left.Y^*(0) \leq 2 \cdot 365 \text{ days } \right\vert V = v\right]$. } These point estimates show relevant heterogeneity with respect to the treatment resistance (horizontal axis denotes values of $V$) and with respect to the recidivism horizon (different colors denote different values of $y$).

figure[figure omitted — 2,076 chars of source]

First, the $DMTE\left(y,\cdot\right)$ functions are increasing. This functional behavior indicates that defendants whom almost all judges would punish are less likely to recidivate, while defendants who would be punished only by tough judges are more likely to recidivate compared to situations in which they would not be punished.\footnote{Appendix (ref) shows that these results are robust to violations of Assumption (ref).} This conclusion is supported by our 90%-confidence intervals (Figures (ref)-(ref)). { Importantly, these point-wise confidence intervals suggest that constant treatment effects are implausible in our empirical context. Hence, they highlight the importance of taking idiosyncratic latent heterogeneity seriously through an analysis of “MTE-like” parameters.}

Second, the $DMTE\left(y,\cdot\right)$ functions are steeper for $y \in \left\lbrace 3, 4, 5, 6\right\rbrace$ than for $y \in \left\lbrace 1, 2, 7, 8\right\rbrace$. This functional behavior indicates that the effect of alternative sentences on recidivism is more intense for extreme cases (small or large punishment resistance levels) in the mid-run than it is in the short or long-run horizon.

This rich heterogeneity underscores the importance of accounting for different levels of treatment resistance. Our point estimates suggest that designing sentencing guidelines that encourage strict judges to become more lenient could increase time to recidivism. However, the $DMTE$ functions do not allow us to quantify this impact directly.

For this reason, $DMTE$ functions may not be the ideal way to convey the main takeaway of the application, even though they answer well-posed and policy-relevant questions. In what follows, we show that this limitation can be minimized by focusing on other functionals of interest, such as the QMTE, which are measured in days instead of percentage points.

Estimated QMTE function

To better understand the time trade-offs associated with the effect of punishment on time to recidivism, we now focus on the average quantile marginal treatment effect functions. These functionals are easier to interpret than the DMTE functions because they express the underlying treatment effects in the same units as the time-to-recidivism outcomes, i.e., days before the first recidivism event.

Figures (ref) and (ref) show the estimated average QMTE$\left(\tau,\cdot\right)$ functions for $\tau \in \left\lbrace .10, .15, .25, .30, .40, .50, .75 \right\rbrace$. Once more, these point estimates show relevant heterogeneity with respect to the punishment resistance (horizontal axis denotes values of $V$).

Although the level of the estimated $QMTE\left(\tau,\cdot\right)$ functions depends on the quantile, all functions are decreasing in the unobserved resistance to punishment. These point estimates suggest that defendants whom almost all judges would punish would take longer to recidivate when punished. In contrast, defendants who would be punished only by tough judges would recidivate faster compared to situations in which they would not be punished. This result is statistically significant for $\tau \in \left\lbrace .10, .15, .25, .30, .40, .50, .75 \right\rbrace$ at the 10% significance level, according to Figures (ref) and (ref)-(ref) in the Appendix. { Interestingly, these point-wise confidence intervals suggest that constant treatment effects are likely invalid in our empirical setting. Hence, they emphasize the importance of taking essential heterogeneity seriously through an analysis of “MTE-like” parameters.}

We reach a similar conclusion when we analyze the $QMTE\left(\cdot, v\right)$ as a function of the quantiles for specific values of unobserved resistance to treatment. Figure (ref) in the Appendix shows the average $QMTE\left(\cdot, v\right)$ for $v \in \left\lbrace .3, .4, .5, .6, .7 \right\rbrace$. Our results suggest that this function is always positive for small values of the unobserved resistance to punishment, while it is always negative for large values of $v$.

Overall, our QMTE point estimates suggest that designing sentencing guidelines that encourage strict judges to become more lenient could increase time-to-recidivism.

Comparison with other available methods

{

Here, we compare our proposed methods against other available methods in the literature. First, we compare our DMTE methods against the typical judge-fixed-effect regressions that drop many observations to avoid handling censored outcomes directly. Second, we compare our QMTE methods against alternative methods.

When analyzing distributional effects, we focus on recidivism within 2 and 3 years in Figures (ref) and (ref). Our proposed methods are illustrated by the purple lines. We have the average $DMTE\left(2,\cdot \right)$ function in Figure (ref) and the average $DMTE\left(3,\cdot \right)$ function in Figure (ref) (Corollary (ref)). The orange lines are the treatment coefficients of two-stage least squares regressions that use each judge's punishment rate to instrument for the defendant's final punishment and include a full set of court district fixed effects. In Figure (ref), the outcome variable is an indicator equal to 1 if the defendant recidivated within 2 years, and the regression includes 43,468 observations. In Figure (ref), the outcome variable is an indicator equal to 1 if the defendant recidivated within 3 years, and the regression includes only 35,405 observations because it has to remove 8,063 defendants who are observed for less than 3 years. The light blue lines show the treatment coefficients from regressions that include the censoring variable as an additional control in the specification of the orange lines.

figure[figure omitted — 2,738 chars of source]

The comparisons in this figure illustrate the drawbacks of using typical methods instead of embracing essential heterogeneity and addressing censoring directly as we advocate. First, the regressions in the orange and blue lines in Figure (ref) have 8,063 fewer observations than the regressions in the orange and blue lines in Figure (ref). This reduction in sample size shows that the typical approach either risks compositional changes (i.e., different samples for each time horizon) or must lose statistical power for shorter time horizons. Second, the typical judge-fixed effect regression does not capture the rich heterogeneity behind the treatment effects of fines and community service. In particular, the typical regression estimates suggest a positive effect, ignoring that the treatment decreases the probability of recidivism for some defendant types. Importantly, these regression estimates do not lie entirely within the 90%-confidence intervals of the correctly estimated $DMTE\left(2,\cdot\right)$ and $DMTE\left(3,\cdot\right)$ functions (Figures (ref) and (ref)).

Now, we compare our QMTE methods against other available methods in the literature. Differently from our approach, these estimates ignore that the outcome variable is right-censored and provide different conclusions when compared against our proposed estimator.} For brevity, we focus our attention on the effects on the 25th and 50th percentiles ($QMTE\left(.25, \cdot \right)$ and $QMTE\left(.50, \cdot \right)$ functions) in Figures (ref) and (ref). Our proposed methods are illustrated by the purple lines. We have the cross-district average $QMTE\left(.25,\cdot \right)$ function in Figure (ref) and the cross-district average $QMTE\left(.50,\cdot \right)$ function in Figure (ref) (Corollary (ref)). The light blue lines denote a “naive” version of our estimators that follows the same steps as described in Section (ref), but do not condition on the censoring variable, i.e., it removes the terms associated with $C$ from Equations (ref)-(ref). The orange lines denote the standard method in the IV literature that accounts for endogenous selection into treatment but ignores (or aggregates) treatment effect heterogeneity with respect to unobserved resistance to treatment. The orange line in Figure (ref) is the treatment coefficient of an IV quantile regression Kaplan2017 for the 25th percentile, while the orange line in Figure (ref) is the treatment coefficient of an IV quantile regression Kaplan2017 for the 50th percentile. Both IV quantile regressions use the censored outcome variable as the left-hand side variable, control for court district fixed effects, and use the judge's punishment rate as the instrument for the defendant being punished.

Analyzing Figures (ref) and (ref), we find that the IV quantile regression does not capture the rich heterogeneity behind the treatment effects of fines and community service. In particular, the IV quantile regression estimates suggest a negative effect, ignoring that the treatment increases time-to-recidivism for some defendant types. Importantly, the IV quantile regression estimates do not lie entirely within the 90%-confidence intervals of the correctly estimated $QMTE\left(.25,\cdot\right)$ and $QMTE\left(.50,\cdot\right)$ functions (Figures (ref) and (ref)).

Moreover, in Figure (ref), we observe that our proposed estimator (purple line) and its naive version (light blue line) reach similar point estimates. This finding is unsurprising because the estimated $QMTR_{d}\left(.25,\cdot\right)$ functions are always smaller than 2.5 years, and all defendants are observed for at least 2 years. Consequently, the censoring problem is not binding for low percentiles.

However, the censoring problem is binding for higher percentiles. In Figure (ref), we focus on the $QMTR_{d}\left(.50,\cdot\right)$ function and find that our proposed estimator (purple line) and its naive version (light blue line) differ in relevant ways. For example, the naive estimator finds a $QMTE$ function that is less steep, implying smaller treatment effects for extreme values of punishment resistance. Importantly, the naive estimates do not lie entirely within the 90%-confidence intervals of the correctly estimated $QMTE\left(.50,\cdot\right)$ function (Figure (ref)).

All in all, our results indicate that our proposed tools can provide detailed measures of treatment effect heterogeneity of punishing misdemeanor offenses on time-to-recidivism that other methods are not meant to capture.

Conclusion

In this paper, we identify the distributional marginal treatment effect ($DMTE$) and the quantile marginal treatment effect ($QMTE$) functions when the outcome variable is right-censored. To do so, we extend the MTE framework Heckman2006,carneirolee2009 to scenarios with duration outcomes. In this section, we deepen our empirical discussion.

Concerning its empirical contribution, our work is inserted in the literature about the effect of fines and community service sentences on future criminal behavior. Four recent papers in this field were written by Huttunen2020, Giles2021, Possebom2022, and Lieberman2023. They all focus on binary variables indicating recidivism within a pre-specified period. Huttunen2020 and Giles2021 find that this type of punishment increases the probability of recidivism in Finland and Milwaukee (a city in the State of Wisconsin in the U.S.), respectively. Possebom2022 finds that this type of punishment has a small and statistically insignificant effect on the probability of recidivism in São Paulo, Brazil. Finally, Lieberman2023 analyzes five American states and finds that court fees do not impact recidivism.

Unlike these four papers, our outcome variable is time-to-recidivism. Using a continuous outcome instead of binary indicators allows for a finer analysis of the heterogeneous effects of fines and community service sentences on future criminal behavior, and may reconcile the conflicting results in the previous literature. For example, we find that this type of punishment increases time-to-recidivism for some individuals while decreasing it for other individuals. If the first type of individual is more common in the states analyzed by Lieberman2023 than in Milwaukee and Finland, our focus on essential heterogeneity may explain these results.

\singlespace

\pagenumbering{arabic}

\setcounter{table}{0}

\setcounter{figure}{0}

\setcounter{equation}{0}