Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,989 characters · 13 sections · 66 citation commands
Abstract: I partially identify the marginal treatment effect (MTE) when the treatment is misclassified. I explore two restrictions, allowing for dependence between the instrument and the misclassification decision. If the signs of the derivatives of the propensity scores are equal, I identify the MTE sign. If those derivatives are similar, I bound the MTE. To illustrate, I analyze the impact of alternative sentences (fines and community service v. no punishment) on recidivism in Brazil, where Appeals processes generate misclassification. The estimated misclassification bias may be as large as 10% of the largest possible MTE, and the bounds contain the correctly estimated MTE.
Keywords: Misclassification, Instrumental Variable, Partial Identification, Alternative Sentences, Recidivism. JEL Codes: C31, C36, K42.
Evaluating a policy with a misclassified treatment variable is theoretically challenging Ura2018,Calvi2019,Acerenza2021. However, this type of problem is widespread in empirical economics. For example, when analyzing the effect of incarceration or alternative sentences in Crime Economics, the treatment variable will be misclassified if the researcher has information only about the trial judge's ruling. In this context, measurement error is created by the appeal process because Appeals Court judges may reverse the trial judge's ruling Green2010. Furthermore, education attainment Black2003 and welfare program participation Hernandez2007 are likely to be mismeasured. Last but not least, with the increasing availability of large data sets, prediction methods are being used to infer the treatment status in a variety of empirical questions ArellanoBover2020. Since no prediction algorithm perfectly classifies the treatment status, the observed treatment variable is misclassified.
In this paper, I provide easy-to-compute bounds around the marginal treatment effect (MTE) function when the treatment variable is misclassified. To do so, I propose two partial identification strategies under increasingly restrictive sets of assumptions, extending the MTE framework Heckman2006 to scenarios with a misclassified treatment variable. Importantly, I allow the instrument to depend on the potential misclassified treatment variables and on the misclassification decision.
The MTE is a function that captures the effect of a treatment for the individual who is indifferent between taking the treatment or not. In the Crime Economics example, the MTE function captures the effect of being punished with an alternative sentence (fines or community service) on recidivism for the defendant who is at the margin of being punished or not conditional on her judge's leniency level. By analyzing this treatment effect parameter at different margins of judge's leniency, we may find that its heterogeneity is correlated with judge's leniency. Therefore, understanding the unobserved heterogeneity of alternative sentences' impact is key to understanding its benefits and costs.
My partial identification strategy for the MTE function with a misclassified treatment variable offers a menu of estimates based on two sets of assumptions. These assumptions simultaneously address endogenous selection into treatment and non-classical measurement error. Since these sets gradually add stronger assumptions to tighten the bounds around the MTE function, the researcher can transparently analyze the informational content of each assumption Tamer2010.
Four assumptions are common to all sets of assumptions used in this paper. First, the instrument is independent of the potential outcomes and the latent heterogeneity variable that defines the true treatment. Second, the instrument is relevant, impacting the correctly measured and mismeasured probabilities of receiving the treatment. Third, the latent heterogeneity variable that defines the treatment is continuous. Finally, the potential outcomes' first moments are finite. These assumptions are standard in the literature about instrumental variables Heckman2006 and address endogenous selection into treatment.
Under these four assumptions, I show that the Local Instrumental Variable (LIV) estimand is biased relative to the MTE function when the treatment variable is misclassified. This bias depends on the instrument's value and may move the LIV estimand in any direction. Consequently, even with expert knowledge, predicting the direction of the misclassification bias is challenging. Moreover, when using a misclassified treatment variable, the standard instrument validity tests (Fradsen2019, and Heckman2006) may fail, and the IV weights may not integrate to one even when all weights are positive.
To account for misclassification of the treatment variable, I add two increasingly strong assumptions to those four common assumptions.
First, I impose that the instrument's impacts on the correctly measured probability of treatment and on the mismeasured probability of treatment have the same sign. With this assumption, I identify the sign of the MTE function at any point in the instrument's support.
Second, I impose that the instrument's impacts on the mismeasured probability of treatment and on the correctly measured probability of treatment are similar. With this assumption, I uniformly bound the magnitude of the MTE function at any point and analyze the bounds' sensitivity to the degree of misclassification.
To estimate the bounds around the MTE function, I suggest a parametric model. I impose a polynomial model for the correctly measured propensity score and for the conditional expectation of the treatment effect as a function of the latent resistance to treatment, implying that the outcome equation's reduced-form model is a polynomial function too. By combining this object with a reduced-form polynomial model for the misclassified treatment variable, I estimate the mismeasured LIV estimand, and the MTE function's sign and bounds.
To exemplify the identification tools proposed in this work, I evaluate the effect of alternative sentences (fines and community service in comparison to no punishment) on recidivism in the State of São Paulo, Brazil, between 2010 and 2019.\footnote{São Paulo is the largest state in Brazil, with a population above 44.4 million people according to the Brazilian Census in 2022.} To do so, I observe trial judge's full sentences and divide them into two groups: (i) punished (treated group), containing defendants who were fined or sentenced to community services, and (ii) not punished (untreated group), containing defendants who were acquitted or whose cases were dismissed. To measure recidivism, I check whether the defendant's name appears in any criminal case within two years after the final sentence date.
This context illustrates my framework in two dimensions. First, we need an instrumental variable because the econometrician does not observe all the variables that influence the defendant's future criminal behavior and are used by the trial judge to decide the defendant's punishment. To address endogenous selection into punishment, I use the trial judge's leave-one-out rate of punishment (or “leniency rate”) as an instrument for the trial judge's decision Bhuller2019. Second, in Brazil, defendants will only fulfill their sentences after their judicial case is closed, implying that they will only be punished after the Appeals process or after they inform the Court System that they will not appeal. Consequently, using only trial judges' rulings to define which defendants were punished with an alternative sentence introduces a natural misclassification problem.\footnote{Within the empirical judge fixed effect literature, some authors construct their treatment variables based on trial judges' rulings --- e.g., Kling2006; Green2010,Bhuller2019 --- while others construct their treatment variables based on the final ruling or on the actual sentence served by the defendant --- e.g., Kling2006; Arteaga2019. Although the first group of authors is careful when interpreting their results as the impact of trial judge's decisions on future criminal and labor outcomes, we may affirm that there would be a misclassification problem if their focus had been on identifying the effect of final rulings.}
To better understand my methods' ability to identify the correctly measured MTE function, I also collect data on Appeals Court's decisions. By doing so, I can use the results based on the correctly measured punishment decision (each case's final ruling) to evaluate the bias in an analysis that ignores the misclassification problem. Moreover, I can compare these results with the results derived from the proposed set identification methods.
I find that the misclassification bias can be relatively large and complex. For example, I estimate that this type of bias can be as large as 10% of the largest possible treatment effect. Furthermore, the misclassification bias can be either positive or negative depending on the instrument's value. Consequently, even an expert may have difficulties theorizing about the misclassification bias' direction and magnitude. For this reason, adopting methods that account for a possibly misclassified treatment variable is useful.
I also find that the proposed partial identification strategies work in practice. When bounding the MTE function, the estimated sets cover the estimated MTE function entirely in every court district.
Finally, using each case's final rulings as my treatment variable, I find that the effect of alternative sentences on recidivism is likely small. Additionally, the observable geographic heterogeneity across court districts appears to matter more than the unobservable heterogeneity across defendants' resistance to treatment.
Concerning its theoretical contribution, my work is inserted in the literature about identifying treatment effect parameters with measurement error. As illustrated by Hu2017, this literature is vast and growing. Similarly to my work, four recent papers focused on identifying treatment effect parameters with heterogeneous effects, endogenous selection into treatment and misclassification of the treatment variable: Ura2018, Calvi2019, Tommasi2020 and Acerenza2021.\footnote{Another important characteristic in this literature is whether the measurement error is differential or not. Appendix (ref) provides a detailed discussion on this topic and illustrates that my framework allows for differential measurement error.}\textsuperscript{,}\footnote{Focusing on the difference-in-differences framework instead of the instrumental variable model, Denteh2022 and Negi2022 also contribute to the literature about misclassified treatment variables.}
These papers differ, for example, with respect to their target parameters. While the first three papers focus on the Local Average Treatment Effect (LATE) parameter, Acerenza2021 and I focus on the MTE function.
Moreover, unlike these papers, I do not assume the instrument is independent of the potential misclassified treatment variables or the misclassification decision. This flexibility is relevant in a variety of applied examples Bound2001,Haider2020.
Allowing for co-dependence between the instrument and the misclassification decision is particularly important in my empirical application. Since sentences by extreme trial judges may be more frequently reversed than sentences by median trial judges, my instrument may be correlated with the misclassification decision. In fact, I find that stricter trial judges are more likely to have their sentences reversed in my empirical application (Subsection (ref)).
Even when Ura2018 and Acerenza2021 extend their main bounds to allow for dependence between the instrument and the potential misclassified treatment variables, our strategies still complement each other. While I provide easy-to-derive bounds that can be used in a sensitivity analysis with respect to the degree of misclassification, they focus on worst-case bounds.\footnote{Appendix (ref) provides a detailed comparison between the bounds provided in Section (ref) and the bounds proposed by Acerenza2021. In particular, I show that, under some conditions, the bounds proposed by Acerenza2021 contain zero while the bounds provided in Section (ref) exclude this value.}
Concerning its empirical contribution, my work is inserted in the literature about the effect of alternative sentences on future criminal behavior. Three recent papers in this field were written by Huttunen2020, Giles2021 and Klaassen2021. While the first group of authors uses data from Finland, the second uses data from Milwaukee (a city in the State of Wisconsin in the U.S.). Both sets of authors find that alternative sentences increase recidivism. Differently from them, Klaassen2021 finds that alternative sentences decrease recidivism in North Carolina (a state in the U.S.). Unlike these previous studies, my estimated treatment effect parameters are small and rarely statistically different from zero. Given the large amount of observable geographic heterogeneity in my estimated results, the difference between the recent literature and my findings may be due to different geographic contexts. A deeper understanding of the mechanisms behind these differences is beyond the scope of this work, even though they deserve further investigation.
This paper is organized as follows. Section (ref) presents the structural model and the misclassification mechanism and discusses the identifying assumptions. In Section (ref), I provide the identification results for the MTE bounds with a misclassified treatment variable under each set of assumptions. Moreover, Section (ref) briefly explains how to estimate the objects that are necessary to implement the identification strategy described in the previous section. Finally, Section (ref) describes the data and discusses the empirical results, while Section (ref) concludes. This paper also contains an online supporting appendix.
In this section, I assume that I do not observe the final ruling in each case and develop an econometric model that simultaneously addresses endogenous selection-into-treatment and misclassification. To analyze the Marginal Treatment Effect (MTE) when the treatment variable is mismeasured, I start with the standard generalized selection model Heckman2006, described in the potential outcome framework:
where $Z$ is an observable continuous instrumental variable (trial judge's leniency rate) with support given by a set $\mathcal{Z} \subset \mathbb{R}$, $P_{D} \colon \mathcal{Z} \rightarrow \mathbb{R}$ is an unknown function, $U$ is a latent heterogeneity variable (defendant's resistance to treatment or the amount of criminal evidence in her favor) and $D$ is the correctly classified treatment status (indicator that the defendant received some type of punishment --- non-prosecution agreement or conviction --- in her case's final ruling).\footnote{My framework can be adapted to binary instruments. This case, which bounds the LATE parameter, is detailed in Appendix (ref) and complements the work developed by Ura2018 and Calvi2019.} Equation (ref) models how the agent self-selects into treatment.\footnote{Note that, according to Vytlacil2002, this threshold-crossing model is equivalent to the monotonicity assumption imposed in the LATE framework Imbens1994.} Variable $Y$ is the realized outcome variable (recidivism indicator), while $Y_{0}$ and $Y_{1}$ are the potential outcomes when the agent is untreated (not punished) and treated (punished), respectively.
I augment this model with the possibly misclassified treatment status indicator $T$ (trial judge's sentence). Note that the binary nature of the treatment variable implies that the measurement error $\left(T - D\right)$ is non-classical, i.e., $Cov\left(T - D, D\right) < 0$. The misclassified treatment is relevant because the researcher observes only the vector $\left(Y, T, Z\right)$, while $Y_{1}$, $Y_{0}$, $D$ and $U$ are latent variables.\footnote{For brevity, I drop exogenous covariates from the model even though all results derived in the paper hold conditionally on covariates. Moreover, I assume that there is only one instrument even though all results hold with partial derivatives when there is a vector of continuous instrumental variables (Appendix (ref)). This appendix also discusses a model with more than one misclassified treatment variable. Such an extension is useful to empirical applications whose treatment variable is defined based on a prediction algorithm ArellanoBover2020.}
Following Heckman2006, I impose four assumptions.
Assumption (ref) is an exogeneity assumption and is common in the literature about instrumental variables. In my empirical application, this assumption holds conditional on the court district because, in the State of São Paulo, Brazil, trial judges are randomly assigned to cases within each court district. Furthermore, this type of assumption is common in the judge fixed-effect literature Bhuller2019, Education Cornelissen2018 and many other applied fields.
Assumption (ref) is a rank condition, intuitively imposing that the instrument is locally relevant everywhere.\footnote{When combined with the continuous differentiability of $P_{D}$ and $P_{T}$ as implicitly imposed by my estimation method (Section (ref)), Assumption (ref) implies that $P_{D}$ and $P_{T}$ are strictly monotone. Alternatively, I could impose that the set $\mathcal{Z}_{0} \coloneqq \left\lbrace z \in \mathcal{Z} \colon \dfrac{dP_{D}\left(z\right)}{dz} = 0 \text{ or } \dfrac{dP_{T}\left(z\right)}{dz} = 0 \right\rbrace$ has measure zero. In this case, $P_{D}$ and $P_{T}$ could be non-monotone. Moreover, Propositions (ref)-(ref) and Corollaries (ref), (ref) and (ref) would still hold for any $z \in \mathcal{Z} \setminus \mathcal{Z}_{0}$ if I defined the domain of all relevant functions as $\mathcal{Z} \setminus \mathcal{Z}_{0}$ instead of $\mathcal{Z}$. I use Assumption (ref) instead of its weaker version for notational ease.} Note also that Assumption (ref) is stronger than the rank condition usually imposed in the literature about marginal treatment effects Heckman2006. In particular, since the correctly measured treatment variable is not observed, it is impossible to directly test that $\dfrac{dP_{D}\left(z\right)}{dz} \neq 0$ for every $z \in \mathcal{Z}$. However, since the mismeasured propensity score is trivially identified, it is possible to test that $\dfrac{dP_{T}\left(z\right)}{dz} \neq 0$ for every $z \in \mathcal{Z}$. Observe also that Assumption (ref) implies that $0 < \mathbb{P}\left[D = 1\right] < 1$, a support condition that is required for any evaluation estimator.
Assumption (ref) is a regularity condition that allows me to normalize the marginal distribution of $U$ to be the standard uniform. Consequently, I have that the true propensity score $\mathbb{P}\left[\left. D = 1 \right\vert Z = z\right]$ is equal to $P_{D}\left(z\right)$ for any $z \in \mathcal{Z}$. However, due to misclassification of the treatment variable, the correctly measured propensity score is not identified.
Assumption (ref) is a regularity condition that allows me to apply standard integration theorems and ensures that average treatment effects are well-defined.
Due to misclassification of the treatment variable, Assumptions (ref)-(ref) are not sufficient to identify the marginal treatment effect function (Proposition (ref)). To address this problem, I gradually impose two increasingly strong assumptions that allow me to derive increasingly strong identification results in Section (ref). Consequently, I offer a menu of estimates whose credibility can be assessed by each reader based on the plausibility of each assumption Tamer2010. Intuitively, these assumptions restrict the amount of measurement error by constraining the relationship between $D$, $T$ and $Z$.
To derive my first partial identification result (Corollary (ref)), I require a weak sign restriction on the impact of $Z$ on the true treatment variable $D$ and on the misclassified treatment variable $T$.\footnote{Sign restrictions have been used previously in the misclassification literature to identify the treatment effect parameter in a linear model with homogeneous treatment effects Haider2020.} Under Assumption (ref), it is possible to identify the sign of the marginal treatment effect.
In my empirical application, Assumption (ref) imposes that tougher trial judges also increase the probability of receiving some type of punishment according to each case's final ruling. Intuitively, this assumption holds if tougher trial judges write more compelling rulings that are more likely to be affirmed by the Appeals Court. Alternatively, if only a small share of defendants appeal, this assumption may hold even if tougher trial judges' rulings are more likely to be reversed. Moreover, as discussed in Example (ref) (Appendix (ref)), this assumption holds if trial judges and Appeals judges have the same sentencing criteria, but face different information sets.\footnote{For another example about investment decisions, see Appendix (ref). Additionally, in Appendix (ref), I impose restrictions on the model's primitives that imply Assumptions (ref) and (ref). The sufficient conditions associated with Assumption (ref) are connected with the framework proposed by Tommasi2020.}
To derive a stronger partial identification result (Proposition (ref)), I impose not only that the impact of $Z$ on $D$ and $T$ have the same sign, but also that those impacts are not arbitrarily different from each other. Under Assumption (ref), it is possible to uniformly bound the marginal treatment effect function.
Note that using a smaller $c$ imposes a stronger restriction on the relationship between $D$, $T$ and $Z$.\footnote{I impose that the function $\dfrac{\sfrac{dP_{D}\left(\cdot\right)}{dz}}{\sfrac{dP_{T}\left(\cdot\right)}{dz}}$ is bounded by a constant for simplicity. Alternatively, I could impose that $\dfrac{\sfrac{dP_{D}\left(\cdot\right)}{dz}}{\sfrac{dP_{T}\left(\cdot\right)}{dz}}$ is bounded by $c\colon\mathcal{Z}\rightarrow\left[1,+\infty\right)$, where $c\left(\cdot\right)$ is a nontrivial function of the instrument. However, knowing an entire bounding function in any concrete empirical context may be difficult.} In practice, the researcher may be uncertain about the value of $c$. After deriving Proposition (ref), I explain how to choose the largest value of $c$ that is compatible with the data and the other model assumptions. Alternatively, the researcher can follow a sensitivity analysis strategy Cinelli2019 and present results for different values of $c$.\footnote{If the researcher observes the correctly classified and the misclassified treatment variable for a subsample of her sample, she can use this subsample to estimate $c$ and use the estimated $c$ in Assumption (ref) to partially identify the MTE function in her entire sample.}
In my empirical application, Assumption (ref) imposes that decreasing the trial judge's leniency increases the punishment probabilities according to the trial judge's ruling and according to the final ruling by similar amounts. As explained in Section (ref), the minimum valid $c$ in my empirical application is estimated to equal $1.13$.
To illustrate the theoretical plausibility of my framework in my empirical application, Example (ref) (Appendix (ref)) describes a simple model of Appeals Courts reversing trial judges' rulings.\footnote{Appendix (ref) explains that Assumption (ref) may also be plausible when analyzing returns to education if having a college degree is randomly miscoded in a survey. Moreover, Appendix (ref) illustrates that Assumptions (ref)-(ref) and (ref) are compatible with differential measurement error.}
In Appendix (ref), I impose a stronger assumption that allows me to derive sharp uniform bounds around the MTE function and sharply bound any weighted integral of the MTE function (e.g., Average Treatment Effect (ATE), Average Treatment Effect on the Treated (ATT) and Average Treatment Effect on the Untreated (ATU)). This extra condition imposes a restriction on the functional relationship between the correctly measured propensity score and the mismeasured propensity score in the sense that it connects the level and all the derivatives of those two objects.
My goal is to derive partial identification results for the Marginal Treatment Effect as a function of the value of the instrument (MTE function). Formally, I define this object as $\theta \colon \mathcal{Z} \rightarrow \mathbb{R}$ such that, for any $z \in \mathcal{Z}$,
Intuitively, this definition of the MTE function captures the effect of a treatment for the individual who is indifferent between taking the treatment or not, where the margin of indifference is defined by the value of the individual's instrument.\footnote{Given the equivalence result by Vytlacil2002, it is possible to write the MTE function $\theta$ as depending on counterfactual treatment choices that satisfy the standard LATE monotonicity conditions. For any $z \in \mathcal{Z}$, define the random variable $D^{*}\left(z\right) = \mathbf{1}\left\lbrace U \leq P_{D}\left(z\right) \right\rbrace$. Note that $D^{*}\left(\cdot\right)$ satisfy the standard LATE monotonicity conditions and that the MTE function satisfies $\theta\left(z\right) = \mathbb{E}\left[\left. Y_{1} - Y_{0} \right\vert z = \operatorname*{\arg\!\inf}\limits_{z^{\prime} \colon D^{*}\left(z^{\prime} \right) = 1} P_{D}\left(z^{\prime}\right)\right]$.}
In my empirical application, the MTE function captures the effect of being punished with an alternative sentence on future criminal behavior for the defendant who is at the margin of being found guilty given her judge's leniency levels. Analyzing the MTE function at different margins of judge's leniency is important because punishment may have heterogeneous effects on the defendants and the heterogeneity may be correlated with judge's leniency. Consequently, understanding the heterogeneity of the impact of alternative sentences is key to understanding its benefits and costs.
Note that, while Heckman2006 define the MTE as a function of the latent heterogeneity $U$, I define the MTE as a function of the instrument.\footnote{In my work, I do not use the standard definition of the MTE function and its classic identification result $\left(\mathbb{E}\left[\left. Y_{1} - Y_{0} \right\vert U = p\right] = \dfrac{d \mathbb{E}\left[\left. Y \right\vert P_{D}\left(Z\right) = p\right]}{d p}\right)$ because I cannot point-identify $P_{D}$ due to misclassification of the treatment variable. Consequently, I am unable to associate an instrument value $z$ with a specific margin $u$ of the latent heterogeneity $U$, i.e., I do not know $u = P_{D}\left(z\right)$ even though I know that $\theta\left(z\right) = \mathbb{E}\left[\left. Y_{1} - Y_{0} \right\vert U = P_{D}\left(z\right)\right]$.} Consequently, different instrumental variables are associated with different MTE functions. Although the function $\theta\left(\cdot\right)$ is not policy-invariant according to the definition of Heckman2006, I can still use it to compute interesting policy-relevant treatment effect parameters (PRTE) when the policy-maker can set the level of the instrument. For example, in my empirical application, I can still compute the treatment effect of making all judges as strict or as lenient as the strictest or most lenient judges.\footnote{Similar policy-relevant treatment effect parameters can be defined in any judge's lenience design study, one of the most common applications of the MTE.}
Moreover, in Appendix (ref), I impose Assumptions (ref) and (ref) to derive a one-to-one map between my definition of the MTE function and its standard definition. Consequently, under these restrictive assumptions, I can derive bounds around common treatment effect parameters, such as the ATE, ATT, ATU and PRTE (Corollary (ref)).\footnote{In particular, Assumption (ref) is not valid in my empirical application.}
To partially identify the MTE function (Equation (ref)), I analyze the consequences of a misclassified treatment variable on the Local Instrumental Variable (LIV) estimand and, then, derive increasingly strong identification results based on Assumptions (ref) and (ref). Analyzing the misclassification bias of the LIV estimand is important because this estimand is traditionally used to identify the MTE function in the previous literature.
If the researcher ignores that the treatment variable is misclassified, she can compute the LIV estimand using the misclassified treatment variable $T$ as if it was the actual treatment variable. In this case, the LIV estimand is defined as $f\colon\mathcal{Z}\rightarrow\mathbb{R}$ such that, for any $z \in \mathcal{Z}$,
following Chalak2017 and as $\tilde{f}\colon\mathcal{P}\rightarrow\mathbb{R}$ such that, for any $p$ in the support $\mathcal{P}$ of $P_{T}\left(Z\right)$,
following Heckman2006. The next proposition analyzes which object is identified by both definitions of the LIV estimand and clarifies two negative consequences of ignoring misclassification.
The first negative consequence shown by Proposition (ref) is that, when misclassification is ignored, two different definitions for the LIV estimand do not identify the MTE function. Similarly to the comparison between the LATE and the Wald Estimator when the treatment variable is misclassified Calvi2019,Tommasi2020, there is a scaling factor connecting the LIV estimand and the MTE function.
Differently from the work done by Calvi2019 and Tommasi2020, this scaling factor is a function, suggesting that the bias may be positive, negative or even zero depending on the point where the LIV estimand is evaluated. Interestingly, the LIV estimand's bias attenuates the true MTE function if $\dfrac{\sfrac{dP_{D}\left(z\right)}{dz}}{\sfrac{dP_{T}\left(z\right)}{dz}} < 1$ and enlarges the true MTE function if $\dfrac{\sfrac{dP_{D}\left(z\right)}{dz}}{\sfrac{dP_{T}\left(z\right)}{dz}} > 1$. As illustrated by my empirical application (Section (ref)), this phenomenon complicates an intuitive analysis of the misclassification bias where the researcher tries to guess its direction based on expert knowledge.
The second negative consequence shown by Proposition (ref) is that IV validity tests may fail if misclassification is ignored. For example, when the outcome variable has compact support, Fradsen2019 proposes to test the monotonicity condition (Equation (ref)) by testing whether the function $\dfrac{d\mathbb{E}\left[\left. Y \right\vert P_{T}\left(Z\right) = p \right]}{dp}$ is bounded between the smallest and the largest possible treatment effects. Proposition (ref) shows that this test is not valid when the treatment variable is mismeasured because the scaling factor $\dfrac{\sfrac{dP_{D}\left(P_{T}^{-1}\left(p\right)\right)}{dz}}{\sfrac{dP_{T}\left(P_{T}^{-1}\left(p\right)\right)}{dz}}$ in Equation (ref) is possibly unbounded without extra assumptions. A misclassified treatment variable also renders the monotonicity test proposed by Heckman2005 invalid as detailed in Appendix (ref). Moreover, the mismeasured propensity score may not satisfy index sufficiency as explained in Appendix (ref).
Furthermore, ignoring misclassification makes interpreting the usual IV estimand more difficult. Following Heckman2006, I can show that the naive IV estimand satisfies $$\dfrac{Cov\left(Z, Y\right)}{Cov\left(Z, T\right)} = \int_{0}^{1} \omega\left(u\right) \cdot \mathbb{E}\left[\left. Y_{1} - Y_{0} \right\vert U = u\right] \, du,$$ where $\omega\left(u\right) = \dfrac{\mathbb{E}\left[\left. Z - E\left[Z\right] \right\vert P_{D}\left(Z\right) \geq u\right] \cdot \mathbb{P}\left[P_{D}\left(Z\right) \geq u\right]}{Cov\left(Z,T\right)}$. Unless $Cov\left(Z,T\right) = Cov\left(Z,D\right)$, the weights $\omega\left(\cdot\right)$ do not integrate to one and the IV estimand do not identify a proper weighted average of the MTE even when the weights are positive. Additionally, since $Cov\left(Z,T\right) = Cov\left(Z,D\right)$ is equivalent to $Cov\left(Z,T - D\right) = 0$, this condition intuitively imposes a testable restriction in my empirical application: sentences by extreme trial judges are as likely to be reversed as sentences by median trial judges. Since I find that stricter trial judges are more likely to have their sentences reversed by the Appeals Court (Subsection (ref)), the naive IV estimand will have weighting problems in my empirical application.
Now, to derive increasingly strong identification results for $\theta\left(\cdot\right)$, I add Assumptions (ref) and (ref). The first identification result (Corollary (ref)) shows that I can identify the sign of the MTE function under a weak assumption about the signs of the derivatives of the correctly measured propensity score function and of the mismeasured propensity score function.
Knowing the sign of the MTE function $\theta\left(z\right)$ at a point $z \in \mathcal{Z}$ is important. If the instrument is policy-relevant, this result can be used to ensure that, at every choice margin given by the instrument, the expected benefit is positive. For example, in my empirical application, the policymaker can re-educate judges whose punishment rates are related to a positive effect on recidivism to change their punishment criteria to points $\theta\left(z\right)$ that are associated with a negative effect on average. Even when the instrument is not policy-relevant, knowing whether the MTE function $\theta\left(\cdot\right)$ is mostly positive or negative is useful to evaluate the pros and cons of a treatment.
The second identification result (Proposition (ref)) shows that, under an assumption about the ratio between the derivatives of the correctly measured propensity score function and of the mismeasured propensity score function, I can uniformly bound the MTE function. Moreover, the distance between the true MTE function and any function in this set is bounded above by an identifiable constant under Assumptions (ref)-(ref) and (ref).
Proposition (ref) is strictly stronger than Corollary (ref) in the sense that not only it identifies the sign of $\theta\left(z\right)$ at any point $z \in \mathcal{Z}$, but it also provides the largest possible effect and the smallest possible effect for each value of the instrument. If the instrument is policy-relevant, this result can be used to ensure that, at every choice margin given by the instrument, the expected benefit is larger than the expected treatment cost. For example, in my empirical application, the policymaker can re-educate judges whose punishment rates are related to effects that are not large enough to compensate for punishment costs to change their punishment criteria to points associated with effects that pass, on average, a cost-benefit analysis. Even when the instrument is not policy-relevant, bounding the MTE function $\theta\left(\cdot\right)$ is useful to know whether most points pass, on average, a cost-benefit analysis.
Due to Proposition (ref), I can also adapt the test proposed by Fradsen2019 to test Assumptions (ref)-(ref) and (ref) when the support of $Y_{0}$ and $Y_{1}$ is bounded. To do so, I check whether $\sup_{\tilde{\theta} \in \Theta_{1}}\sup_{z \in \mathcal{Z}}\tilde{\theta}\left(z\right)$ is smaller than the largest possible effect and whether $\inf_{\tilde{\theta} \in \Theta_{1}}\inf_{z \in \mathcal{Z}}\tilde{\theta}\left(z\right)$ is larger than the smallest possible effect. If that is the case, I do not reject Assumptions (ref)-(ref) and (ref). Note that this test can be used to define the largest value of $c \in \left[1, +\infty \right)$ that is plausible according to the data. To implement it, find $\overline{z}$ and $\underline{z}$ that respectively maximizes and minimizes the function $f\left(\cdot\right)$ and, then, find the largest $c$ such that $\left(\dfrac{1}{c} \cdot f\left(\overline{z}\right), c \cdot f\left(\overline{z}\right), \dfrac{1}{c} \cdot f\left(\underline{z}\right), c \cdot f\left(\underline{z}\right) \right) \in \left[\underline{\theta}, \overline{\theta}\right]^{4}$, where $\underline{\theta}$ is the smallest possible effect and $\overline{\theta}$ is the largest possible effect.
In this section, I briefly explain how to estimate the MTE function's sign (Corollary (ref)) and its bounds (Proposition (ref)). To estimate these objects, I need to estimate two objects: the mismeasured propensity score and the outcome equation's reduced-form model. Importantly, in my empirical application, these objects depend not only on the value $z$ of the instrument but also on the value $x$ of the covariates. These extra variables contain a full set of court district dummies and are included because, in São Paulo, trial judges are randomly allocated to criminal cases only after conditioning on the court district.
To estimate the mismeasured propensity score, I treat it as a purely reduced-form object. Consequently, I model the misclassified treatment variable's conditional expectation as a separable function between the instrument and the covariates, depending on a polynomial of the instrument and a full set of court district dummies. Under these parametric assumptions, this model can be estimated by OLS.
To estimate the outcome equation's reduced-form model, I treat it as an object derived from the economic model's primitives. Specifically, the correctly measured propensity score function and the conditional expectation of the treatment effect as a function of the latent resistance to treatment are separable between the instrument and the covariates, depending on a polynomial of the instrument and a full set of court district dummies. Consequently, the outcome equation's reduced-form model is a polynomial function that depends on the interaction between the instrument and the court district dummies. Under these parametric assumptions, this model can be estimated by OLS.
By combining the parameters from these two OLS regressions, I can estimate the mismeasured LIV estimand (Equation (ref)) and use it to estimate the MTE function's sign and bounds.
Moreover, in my empirical application, I also need to estimate the correctly measured MTE function to use it as a benchmark against the analysis that ignores misclassification. To do so, I need to estimate the LIV estimand that uses the correctly measured propensity score in its denominator. This object's estimator is very similar to the mismeasured LIV estimand's estimator. The only difference is that I now use a parametric polynomial approximation for the conditional expectation of the correctly classified treatment variable.
To have a deeper understanding of the estimation methods and their performance in a Monte Carlo exercise, see Appendix (ref).
In my empirical application, I answer the question: “Do alternative sentences impact recidivism?”. To answer this question, I collect data from all criminal cases brought to the Justice Court System in the State of São Paulo, Brazil, from 2010 to 2019. In Subsection (ref), I briefly explain my dataset and provide key descriptive statistics. In Appendix (ref), I provide detailed descriptive statistics. In Appendix (ref), I provide a detailed explanation on how I constructed my dataset. Finally, in Subsection (ref), I describe the results of my empirical analysis.
I collect data from all criminal cases brought to the Justice Court System in the State of São Paulo, Brazil, between January 4\textsuperscript{th}, 2010, and December 3\textsuperscript{rd}, 2019. I restricted my sample to cases that started between 2010 and 2017 because the last two years are used only to define my outcome variable. Moreover, I focus on criminal cases whose maximum prison sentence is less than 4 years because, according to Brazilian Law, these cases must be punished with alternative sentences (fines and community service). Due to this sample restriction, the most common crime types in my sample are theft and domestic violence. After those two restrictions, my dataset contains 51,731 case-defendant pairs.
In my dataset, I observe the defendant's full name, the defendant's court district, the case's starting date, the assigned trial judge's full name, the trial judge's full sentence, the trial judge's sentence's date, whether the case went to the Appeals Court, the Appeals Court's ruling if there is one, and the Appeals Court's ruling's date if there is one. Based on those variables, I define my outcome variable ($Y = $“recidivism within 2 years of the final sentence”), my misclassified treatment variable ($T = $“trial judge's decision”), my correctly classified treatment variable ($D = $“final ruling”), my instrument ($Z = $ “trial judge's leniency rate”) and my covariates (X = “full set of court district dummies”).
My misclassified treatment variable $T$ is based only on trial judge's sentences and divides them into two groups. The first group (treated) receives a punishment, i.e., its defendants were fined or sentenced to community services because they were either convicted or signed a non-prosecution agreement. The second group (control) did not receive a punishment, i.e., its defendants were acquitted or its cases were dismissed.
My correctly classified treatment variable $D$ also divides the case-defendant pairs into the groups “punished” and “not punished”. However, it considers the final ruling in each case. If the case did not go to the Appeals Court, this variable equals the misclassified treatment variable $T$. However, if the case went to the Appeals Court, this variable may differ from $T$ because it also considers the Appeals Court's decision. If the trial sentence was affirmed, then $D$ is equal to $T$. However, if the trial sentence was reversed, then $D$ is equal to $1 - T$.
Importantly, both treatment variables are based on a logistic LASSO that maps rulings into binary variables (Appendix (ref)). Although my prediction algorithms may make classification errors, I ignore this source of misclassification in my empirical analysis. Consequently, I focus exclusively on the misclassification problems generated by the appeals process. This source of misclassification is arguably more interesting because it is embedded in the economic nature of the problem instead of being mechanically created by a prediction algorithm.\footnote{Alternatively, I could focus on the misclassification error generated by the prediction algorithms. The method proposed in this paper partially identifies the $MTE$ function even if the final sentence is possibly misclassified. The drawback of this approach is the impossibility of estimating the misclassification bias because I would not observe the correctly measured final sentence if I took prediction errors into account. If I were to follow this approach, I could define one misclassified treatment variable for each prediction algorithm (e.g., random forest or logistic LASSO) and use the methods proposed in Appendix (ref) for the case with more than one misclassified treatment variable.}
Table (ref) shows the joint distribution of the correctly classified treatment variable $D$ and the misclassified treatment variable $T$. Since most cases (67.3%) do not go to the Appeals Court, most cases are correctly classified (95.6%) as described in the main diagonal. The other two cells describe the cases that are misclassified when I ignore the Appeals Court's decisions. First, I find that 3.5% of the defendants were punished by the Trial Judge and were able to reverse their sentences in the Appeals Court. Moreover, in 0.9% of the cases, the defendant was not punished by the Trial Judge, but the prosecutor was able to appeal and reverse the decision.
My instrument $Z$ is the trial judge's leniency rate. This variable equals the leave-one-out rate of punishment for each trial judge, where the defendant's own decision is excluded from this average. To do so, I only use the 639 judges who analyzed more than 20 cases during my sample period.
Having described the treatment and instrumental variables, I can now discuss their relationship. Figure (ref) shows three conditional probabilities where the conditioning variable is the instrument $Z$. The orange line is the share of defendants who were initially found not guilty by the trial judge and had their sentences reversed by the Appeals Court conditional on the punishment rate of the trial judge. The dark blue line is the share of defendants who were initially found guilty by the trial judge and had their sentences reversed by the Appeals Court conditional on the punishment rate of the trial judge. Finally, the light blue line is the share of defendants who had their sentences reversed by the Appeals Court conditional on the punishment rate of the trial judge. The dotted lines are robust bias-corrected 95%-confidence intervals Calonico2019.
Figure (ref) illustrates the importance of having more than one alternative set of assumptions when addressing misclassification issues. Although the orange line suggests that there is no clear dependency between being incorrectly not punished (i.e., $T = 0, D = 1$) and the trial judge's leniency rate, the dark blue and light blue lines suggest that being incorrectly punished (i.e., $T = 1, D = 0$) and having a misclassified sentence (i.e., $T \neq D$) depend positively on the trial judge's punishment rate. Consequently, the method proposed by Acerenza2021 is not appropriate to analyze this empirical problem because it imposes that having a misclassified sentence and the instrument are independent. Since my framework (Sections (ref) and (ref)) relies on alternative restrictions on the relationship between $T$, $D$ and $Z$, my method may be appropriate to analyze this empirical question. Importantly, Assumptions (ref) and (ref) seem valid in this application as discussed in Subsection (ref).
I, now, describe how I define my outcome variable ($Y = $“recidivism within 2 years of the final sentence”). A defendant $i$ in a case $j$ recidivated ($Y_{ij} = 1$) if and only if defendant $i$'s full name appears in a case $\bar{j}$ whose starting date is within 2 years after case $j$'s final sentence's date. Importantly, case $\bar{j}$ can be about any type of crime, including more severe crimes whose maximum sentence is greater than four years, while case $j$ has to be about a crime whose maximum sentence is at most 4 years. To match defendants' names across cases, I use the Jaro–Winkler similarity metric and define a match if the similarity between full names in two different cases is greater than or equal to 0.95.\footnote{Abramitzky2019 match full names in historical Censuses in the U.S. and Norway. They define a match between two individuals if the Jaro–Winkler similarity between their names is greater than or equal to 0.90 and if their dates of birth match exactly. Since I do not observe defendants' dates of birth, I adopt a stricter Jaro-Winkler similarity threshold to define a match in my dataset.}
Even though this fuzzy matching algorithm may misclassify the outcome variable of some case-defendant pairs, I assume that this type of error is negligible and do not account for it in my empirical analysis. This assumption is plausible because Brazilian names are frequently long, containing four or more words. For instance, in my dataset, 59.4% of the case-defendant pairs have names with four or more words, and only 6.7% of them have names with only two words.
My covariates contain a full set of court district dummies. Since my identification strategy leverages the random allocation of judges to criminal cases, I only use districts with two or more judges.
In Section (ref), I estimate my results using the entire sample, but I focus my discussion on one court district: Ribeirão Preto. This district has the second largest number of judges in my sample (9 judges) and illustrates most of the issues caused by a misclassified treatment variable when our target parameter is the MTE function. Moreover, Ribeirão Preto is the seventh largest city in the State of São Paulo with more than 700 thousand inhabitants.
At the end, I impose one final restriction in my dataset: common support between the treatment and control groups. To do so, I impose that the minimum and maximum values of the instrument $Z$ are the same across both treatment arms. My final sample has 43,468 case-defendant pairs when I use the correctly classified treatment variable $D$ and 43,461 case-defendant pairs when I use the misclassified treatment variable $T$.
In this subsection, I describe the results of my empirical analysis in four parts. I start by describing the first-stage results in Part (ref). Then, I report the results for the correctly estimated MTE function in Part (ref). Furthermore, I analyze the misclassification bias in my empirical application in Part (ref). Lastly, I discuss the estimated bounds around the MTE function in Part (ref).
I start by presenting the results of the first stage regression in my empirical analysis. In my model, the correctly classified treatment variable $D$ (“final ruling”) is a function of instrument $Z$ (“trial judge's punishment rate”) and court district fixed effects. Following Section (ref) and Appendix (ref), I use a polynomial series to approximate the correctly measured propensity score and report the estimated coefficients of linear and quadratic models in Columns (1) and (2) in Table (ref). Even though the quadratic coefficient is not statistically significant, I use a quadratic propensity score in my analysis because this is the most parsimonious parametric model that allows for a non-constant scaling factor in Equation (ref). Moreover, in Appendix (ref), I estimate the correctly measured propensity score semiparametrically and find that this function is well-behaved, implying that a quadratic model is a good approximation for the correctly measured propensity score.
Furthermore, estimating the correctly measured propensity score allows me to verify the validity of Assumption (ref). First, my instrument is relevant according to the F-statistic of the first-stage regression. Second, Subfigure (ref) shows that the derivative of the correctly measured propensity score (orange line) is around 0.75. Both results imply that the first part of Assumption (ref) is valid under the assumption that the correctly measured propensity score is quadratic, i.e., $\dfrac{dP_{D}\left(z,x\right)}{dz} \neq 0$ for every value $z$ of the instrument and every value $x$ of the covariates.\footnote{In Appendix (ref), I estimate $P_{D}\left(\cdot,\cdot\right)$ semiparametrically. The steep inclines of the estimated functions also suggest that Assumption (ref) holds.}
Additionally, when I assume that $D$ is not observable, my partial identification strategy (Section (ref)) relies on the mismeasured propensity score $\left(P_{T}\left(z,x\right) = \mathbb{E}\left[\left. T \right\vert Z = z, X = x \right]\right)$ to capture features of the MTE function $\theta\left(z,x\right)$. Following Section (ref) and Appendix (ref), I use parametric models to approximate the mismeasured propensity score and report the estimated coefficients in Columns (3) and (4) in Table (ref).
Similarly to the correctly measured propensity score, I focus on the quadratic model for the mismeasured propensity score. First, note that my instrument is relevant according to the F-statistic of the first-stage regression. Second, Subfigure (ref) shows that the derivative of the mismeasured propensity score (purple line) is around 0.75. Both results imply that the second part of Assumption (ref) is valid under the assumption that the mismeasured propensity score is quadratic, i.e., $\dfrac{dP_{T}\left(z,x\right)}{dz} \neq 0$ for every value $z$ of the instrument and every value $x$ of the covariates.\footnote{In Appendix (ref), I estimate $P_{T}\left(\cdot,\cdot\right)$ semiparametrically. The steep inclines of the estimated functions also suggest that Assumption (ref) holds.}
More importantly, Subfigure (ref) shows that Assumptions (ref) and (ref) are valid. Since the ratio between the derivatives of the propensity score functions $\left(\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}}\right)$ is always positive, I know that these derivatives have the same sign. Note also that this ratio is bounded above by 1.13. For this reason and to be conservative, I impose that Assumption (ref) holds with $c = 1.2$ in my empirical analysis (Subfigure (ref)).\footnote{When I semiparametrically estimate the correctly measured propensity score and the mismeasured propensity score (Appendix (ref)), I find that Assumption (ref) holds with $c = 1.19$. To be cautious about the correct value of $c$, my sensitivity analysis results impose a larger value ($c = 1.5$ in Figure (ref)). This difference illustrates the importance of conducting a sensitivity analysis Cinelli2019 where we gradually increase $c$ to understand the impact of allowing for a more intense misclassification problem.}
Subfigure (ref) also illustrates a key consequence of Proposition (ref). The ratio between the propensity score functions can be greater or smaller than one depending on the instrument's value. Consequently, the LIV estimand's bias will attenuate or enlarge the true MTE function depending on the instrument's value.
Now, I report, in Figure (ref) the results associated with Ribeirão Preto's court district. In Subfigure (ref), the orange line is the correctly estimated MTE function $\theta\left(\cdot, \cdot\right)$ (Equation (ref)), the dark blue line is the estimated misclassified LIV estimand (Equations (ref) and (ref)), and the light blue lines are the estimated upper and lower bounds of the set $\Theta_{1}$ (Proposition (ref) and Equations (ref) and (ref)).\footnote{For a discussion about the estimated sign of the MTE function (Corollary (ref)) for all court districts, see Appendix (ref).}
Focusing on the correctly estimated MTE function (orange line), I find that the marginal treatment effect depends on the value of the instrument. For example, my results suggest that, in Ribeirão Preto, alternative sentences (fines and community service) have almost no impact on agents who would be punished by most judges while they increase recidivism for agents who would be punished only by stricter judges (around a punishment rate equal to 0.6). However, this result should be interpreted cautiously because some of the point-estimates are larger than one, the largest possible effect given the support of the outcome variable. Importantly, when I compute bootstrapped 90%-confidence bands (Subfigure (ref)), I find that the confidence bands contain, at least partially, the zero function. This finding indicates that the effect of alternative sentences on recidivism is likely small.\footnote{A small effect of alternative sentences on recidivism is also supported by standard two-stage least squares (2SLS) regressions (Appendix (ref)) and by the standard analysis of the function $\mathbb{E}\left[\left. Y_{1} - Y_{0} \right\vert U = u, X = x \right]$ (Appendix (ref)).}
Figure (ref) also illustrates the danger of ignoring misclassification of the treatment variable. By comparing the estimated MTE function (orange line) and the estimated misclassified LIV estimand (dark blue lines), I can estimate the bias that is generated by the misclassified treatment variable (Subfigure (ref)).
Note that this bias can be negative or positive depending on the sign of the true MTE function and on the size of the ratio between the derivatives of the propensity score functions $\left(\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}}\right)$. First, when the MTE function is negative and $\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}} < 1$ (around a punishment rate equal to 0.35), the bias is positive, implying that there is an attenuation bias. Second, when the MTE function is positive and $\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}} < 1$ (around a punishment rate equal to 0.5), the bias is negative, implying that there is an attenuation bias. Third, when the MTE function is positive and $\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}} > 1$ (around a punishment rate equal to 0.65), the bias is positive, implying that there is an exploding bias. Finally, when the MTE function is negative and $\dfrac{\sfrac{dP_{D}\left(z\right)}{d z}}{\sfrac{dP_{T}\left(z\right)}{d z}} > 1$ (around a punishment rate equal to 0.73), the bias is negative, implying that there is an exploding bias.
This result highlights the complexity of the misclassification bias when estimating entire functions. Note that its sign may change and it may move the estimates away from zero. Consequently, even an expert may have difficulties theorizing about its direction.
To also understand the magnitude of the misclassification bias, I report summary statistics for the LIV estimand's bias in Table (ref). Panel A reports the mean, the standard deviation, the minimum and the maximum of the raw differences between the estimated LIV estimand and the estimated MTE function for five values of the instrument $\left(z \in \left\lbrace .3, .4, .5, .6, .7 \right\rbrace\right)$ across 192 court districts. Note that the average bias varies for each value of the instrument and it achieves values almost as large as 3. More interestingly, the minimum bias across districts is always negative while the maximum bias is always positive.
Although Panel A shows that the estimated bias can be large, these concerning values may be due to the fact that polynomial models do not take into account that the treatment effect parameters must be between -1 and 1 in this empirical application. For this reason, Panel B reports the same summary statistics as Panel A, but, before taking the difference between the estimated LIV estimand and the estimated MTE function, it trims these estimates to be between the minimum possible treatment effect (-1) and the maximum possible treatment effect (1).
Observe that, despite a much smaller average bias, the minimum and maximum biases across districts can be as large as -.11 and .09, respectively. This result implies that the misclassification bias can be relatively large, reaching around 10% of the maximum possible effect. Since it is a priori unknown whether any given empirical application resembles the small bias context of most court districts or the large bias context of some districts, this result illustrates the usefulness of adopting methods that account for misclassification bias when the treatment variable may be misclassified.
As argued in the last subsection, it is important to use methods that account for misclassification bias. When the target parameter is the MTE function, one of these methods is proposed in Sections (ref) and (ref) and illustrated by the light blue lines in Subfigure (ref). Note that the identified set safely contains the estimated MTE function (orange line), exemplifying that the proposed method works in a specific real-world example.\footnote{Subfigure (ref) shows bootstrapped 90%-confidence bands around the identified set.} Moreover, the estimated sets cover the estimated MTE functions entirely in every court district in my sample.
Although successful in my empirical example, the proposed partial identification method may have worked only because my unique dataset allows me to estimate the correctly measured propensity score $P_{D}$ and find a constant $c$ that conservatively satisfies Assumption (ref). Since most real-world applications do not have access to the correctly measured treatment variable $D$, this approach to choosing $c$ is frequently unfeasible.
For this reason, I propose two alternative ways to approximate the constant $c$ using data. The first one is described in Section (ref) and illustrated in Figure (ref). It consists simply of choosing different values of $c$ to understand the impact of allowing for a more intense misclassification problem. The second way to approximate the constant $c$ is described in Section (ref) and consists of choosing a $c$ that is associated with the most intense misclassification problem (largest plausible $c$) in applications whose set $\Theta_{1}$ is still contained within the set of possible treatment effects.\footnote{This method of choosing $c$ is a straightforward adaptation of the test proposed by Fradsen2019.} Unfortunately, it is not possible to illustrate this method of choosing $c$ using Ribeirão Preto's results, because its LIV estimand is already greater than the largest possible treatment effect for some values of the instrument.
Finally, Subfigure (ref) suggests that Assumption (ref) is not valid. For this reason, I do not estimate the sharp bounds of set $\Theta_{2}$ in Proposition (ref) nor the bounds around standard treatment effect parameters (Corollary (ref)).
In this paper, I address a widespread empirical challenge: policy evaluation with a misclassified treatment variable. I propose a novel partial identification strategy to identify the MTE function with a misclassified treatment. This method explores restrictions on the relationship between the instrument, the misclassified treatment and the correctly measured treatment, allowing for dependence between the instrument and the misreporting decision.
As an illustration, I analyze whether alternative sentences (fines and community service) affect recidivism in Brazil. I find that the misclassification bias is empirically relevant, reaching 10% of the largest possible treatment effect. I also find that the estimated bounds contain the correctly estimated MTE function entirely for every court district. Lastly, I find that the effect of alternative sentences on recidivism is likely small even though the point estimates present a large amount of observable geographic heterogeneity.
This result contrasts with recent findings in the empirical literature Huttunen2020,Giles2021,Klaassen2021. While they find significant effects in different directions, my estimated parameters are rarely statistically different from zero. These differences may be due to different geographic contexts and deserve further investigation in future work.
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{equation}{0}