EconBase
← Back to paper

Partial Identification of Marginal Treatment Effects with discrete instruments and misreported treatment

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,930 characters · 10 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification of Marginal Treatment Effects With Discrete Instruments and Misreported Treatment

\setstretch{1}

tabular[tabular omitted — 206 chars of source]

} \newsavebox{\tablebox} \newlength{\tableboxwidth}

center[center omitted — 540 chars of source]

\ Keywords: Treatment effects, instrumental variables, measurement error, partial identification. JEL Codes: C21, C26. Word Count: 11688. \doublespacing

Introduction

This paper provides partial identification results for the Marginal Treatment Effect ($MTE$) in the presence of measurement error in the treatment variable when only a discrete instrument is available. The discrete instrument case is relevant as many applications in the literature rely on these type of instruments. See for example AKK, AKK2, AKK3 and AKK4. The discrete nature of the instrument requires identification strategies to recover the $MTE$ that differ from those explored in the previous literature with continuous instruments. The results of this paper are relevant since it is often true that researchers have access to an instrument with discrete variation (for example, assignment to treatment via an institutional rule), and it is also true that misreporting is a common problem in survey data which is one of the main sources of empirical research. In a more general way, our results can serve as a sensitivity analysis tool for when researchers are interested in recovering the $MTE$ in the presence of a discrete instrument and suspect measurement error and have doubts about their parametric assumptions. Researchers mostly work with self-reported data from surveys; such data systematically present reporting problems that lead to measurement error of the treatment status and, consequently, to bias in the treatment effect of interest. The combination of measurement error with discrete instruments has not been explored in the literature, and it is a fairly common situation to encounter. The results in this paper are useful for identifying MTE (which can be used to recover average effects or policy-relevant effects) in the presence of the two previously mentioned problems for identification. In most cases, researchers observe a discrete (often binary) instrument such as assignment to treatment. In these cases, point identification of the $MTE$ (even without measurement error) is not possible, relying only on the standard assumptions of instrument exogeneity and relevance (See for example BMW). In this paper, under a set of restrictions on the severity of measurement error and shape restrictions, we provide partial identification results for the $MTE$ in the presence of measurement error when a discrete instrument is available. The $MTE$ can help reveal the heterogeneity in the treatment effect. The $MTE$ is relevant in recovering Policy Relevant Treatment Effect parameters ($PRTE$s), Average Treatment Effect ($ATE$), Average Treatment on the Treated ($ATT$), Average Treatment on the Untreated ($ATU$), Local Average Treatment Effects ($LATE$), etc.\footnote{See HV2, HUV, who show the link between the $MTE$ and those parameters via properly weighting the $MTE$.} To achieve partial identification, we introduce smoothness conditions on the marginal treatment responses ($E[Y_d|V=v]$). To deal with the misreporting of the binary treatment, the analysis relies on treating the unconditional probability of misreporting as given.\footnote{One could alternatively take the results from this paper and assume a known upper bound of this probability and take the union of the bounds derived here.} This can be either interpreted as the researcher having prior knowledge on the possible value of the misclassification rates or as a sensitivity analysis tool where the researcher allows for the possibility of misclassification up to a certain level. Relevance and independence of the instrument is required. Although partial identification of the $MTE$ will do not imply in general sharp bounds on the $ATE$. It is still a useful tool to move from local effects and generate bounds on an aggregate relevant effect. Empirical research usually combines a measurement error problem with endogeneity and heterogeneity. U documents in his work, as an example of this, that there is a substantial measurement error in educational attainments in the 1990 U.S. Census. At the same time, educational attainments are endogenous as treatment variables in return to schooling analyses because, among other possibilities, unobserved individual ability affects both schooling decisions and wages. Labor supply response to welfare program participation, in which the outcome is employment status, and the treatment is welfare program participation is subject to similar issues. Self-reported program participation in survey datasets can be misreported as stated by HP. The psychological cost of welfare program participation affects job search behavior and welfare program participation simultaneously.

Related literature

This subsection lists some relevant papers related to the current research paper based on their connections to different aspects of the problem. Namely, misreporting and partial identification of marginal treatment effects. Partial identification of $LATE, MTE$ and $ATE$ with endogenous misreported binary treatments and heterogeneous effects U using a binary instrumental variable, derives bounds for $LATE$ with a binary misreported treatment when an instrument is available, and monotonicity of the true (not observed) treatment in the instrument holds. Identification is achieved by exploiting the relationship between the probability of being a complier and the total variation distance\footnote{The total variation distance between two probability measures $P$ and $Q$ on a sigma-algebra $\mathcal {F}$ of subsets of the sample space $\Omega$ is defined via $\delta (P,Q)=\sup _{A\in {\mathcal {F}}}\left|P(A)-Q(A)\right|$. It can alternatively be defined for probability measures that have densities to be $\frac{1}{2} \int |p-q|d\nu$ where $\nu$ is a measure dominating both probability measures. In the context of U paper the total variation distance calculated is $TV_{Y,D}=\frac{1}{2}\int\Big(\sum_{d=0,1}|f_{y,d|Z=1}(y,d)-f_{y,d|Z=0}(y,d)|\Big)d\nu$ where $f_{y,d|z}$ is the joint density of the observed outcome variable and the observed treatment variable conditional on the value of the instrument.} between people assigned to treatment and the ones that are not. The under-identification for $LATE$ is a consequence of the under-identification for the size of compliers; with no measurement error, one could compute the size of compliers based on the measured treatment and, therefore, $LATE$ would be the Wald estimand. The total variation distance plays a key role in determining the sharp identified set in U. First, it measures the strength of the instrumental variable; when the total variation distance is positive, the identified set of $LATE$ is a strict subset of the whole parameter space, which implies that $Z$ has some identifying power. Secondly, as shown in U lemma 3, the total variation distance is a lower bound for the proportion of compliers which is the under-identified element in the presence of measurement error. CLT, TZ extend U's results for the case where the instrument can take multiple discrete values. AK focuses on bounding the marginal treatment effects when there is a continuous instrument. KPGJ using auxiliary information about the possibility of misreporting and under different combinations of the outcome, treatment, and instrumental monotonicity bounds the $ATE$ for a binary outcome. VP focuses on partially identifying the $MTE$ with a continuous instrument and imposing sign and functional relationships between the derivatives of the true propensity score and the observed one with respect to the continuous instrument. This current paper complements the previously mentioned papers. Fundamentally this paper focuses on identifying the $MTE$ when discrete instruments are available. Such a task requires a different set of assumptions than the ones used to recover directly $ATE$, $LATE$, or $MTE$ with continuous instruments. We complement U, KPGJ and TZ because we are interested in identifying $MTE$ (which can then be used to achieve identification of $LATE$ and $ATE$) instead of the $LATE$ and $ATE$. It is also complementing AK since their analysis relies on the continuity of the instrument. It is worth noticing that it is more common to observe discrete (mostly binary) instruments such as random selection to receive treatment like in medical studies or random selection to receive a treatment conditional on covariates in social sciences ( e.g., Supplemental Nutrition Assistance Program, SNAP). We complement VP since we provide an alternative set of assumptions to identify the $MTE$, and also, we are focusing on a discrete instrument. Identifying marginal treatment effects with discrete instruments BMW show how a discrete instrument can be used to identify the marginal treatment effects under a functional structure that allows for treatment heterogeneity among individuals with the same observed characteristics and self-selection based on the unobserved gain from treatment. This paper builds upon BMW results by considering the case with (endogenous) misreporting and more flexible restrictions (such as shape restrictions instead of parametric assumptions) at the cost of losing point identification. The second one is MST which using the observed instrumental variables estimates, develops a linear programming approach to recover policy-relevant treatment effects such as the $MTE$. This paper differs from it by finding analytical bounds under the different smoothness and shape restrictions. Such bounds permit one to have a first-hand insight into how the assumptions are aiding identification. Estimation of the analytical bounds is simple since it can be performed using their respective sample analogs. Additionally, MST does not allow for the possibility of the treatment to be misreported while here is allowed. In the presence of misreporting, the results from MST do not apply directly while the ones derived here do. In the case of no misreporting, our bounds remain valid; in that sense, our results complement the ones from MST and BMW.

Outline of the paper

The rest of the paper is organized as follows, section (ref) introduces the main framework and assumptions. Section (ref) shows the main identification results and illustrates them. Section (ref) has an application of the identification results to KPGJ. Section (ref) concludes. Additional results are collected in the online appendix. Non analytical results on partial identification without additional shape restrictions extending MST are included the online appendix section (ref). Sections (ref) and (ref) of the appendix focuses on inference for the $ATE$. Section (ref) discusses how to choose the tuning parameter $b$. Section (ref) illustrates the bounds for the $ATE$. Section (ref) illustrates the analytical results on a $DGP$. Section (ref) extends the results with additional monotonicity assumptions. Sections (ref) and (ref) derives the results for the case when the instrument takes more than $2$ values.\footnote{The case of an instrument taking more than two values can also be seen as the generalization to the case of multiple discrete instruments. This is the case because multiple discrete instruments can be combined in one single multi-valued discrete instrument.} Finally, section (ref) collects all the figures from the document.

Analytical Framework

Consider the following framework (AK, HUV and HV):

eqnarray[eqnarray omitted — 193 chars of source]

Where $Y$ is an outcome variable that can be discrete, continuous, or mixed, the potential outcomes are denoted by $Y_d$, which is the outcome realization for when treatment $D=d$, $D=\{0,1\}$ is a binary unobserved endogenous treatment. Let $Z\in \mathcal Z=\{z_0,z_1,...,z_k\}$ be a discrete instrument,\footnote{The results will focus on the binary $z$ case but the generalization is natural for more than two values of $z$.} $V$ is a latent scalar random variable normalized to be uniformly distributed between $(0,1)$. $D^*$ is a misreported binary proxy of $D$, the true unobserved treatment status. $\varepsilon \in \{0,1\}$ is a random variable indicating the presence of misreporting or not. The vector $(Y,D^*,Z)$ is the observed data while $(Y_1, Y_0, D, \varepsilon, V)$ are latent (unobserved). In the rest of the document, small case letters denote realizations of the respective random variables. Object of interest: In this paper, we care about identifying the $MTE(v^*)$ which is the marginal treatment effect at a particular level $V=v^*$, more precisely, it is defined as $E[Y_1-Y_0|V=v^*]$. To identify the $MTE$ in this context, we introduce baseline assumptions that additional assumptions will aid. The baseline assumptions are:

assumption[Random Assignment and Absolute Continuity] The following two conditions hold: \begin{enumerate} • $Z$ is independent of $(Y_d,V, \varepsilon)$ for all $d=(0,1)$. • The distribution of $V$ is absolutely continuous. \end{enumerate}

The previous assumption and the model structure makes innocuous to say that $V$ is uniform between $[0,1]$ and that $p(Z)=P(D=1|Z)$.

assumption[Relevance] Let $Z$ be such that for any $z \in \mathcal{Z}$: \begin{enumerate} • $1>p(z)>0$. • $p(z)\neq p(z\prime)$ for any $z,z\prime \in \mathcal{Z}$. • For any $z,z\prime \in \mathcal{Z}$, we can determine if either $p(z)\leq p(z\prime)$ or $p(z)\geq p(z\prime)$. \end{enumerate}

Assumption (ref)-(ref) include the instrument independence and validity assumption as in HV, HUV among others. Assumption (ref) requires that $Z$ be a valid instrument, in the sense that it is statistically independent of the unobservables in the selection equation and the outcome equation. This assumption was used in AK. This assumption does not require the measurement error to be non-differential.\footnote{Non-differential measurement error is that conditional on the unobserved heterogeneity that drives the selection into treatment, misreporting is independent of the potential outcomes.} Non-differential measurement error combined with assumption (ref) implies that misreporting is independent of the outcome conditional on the true treatment, which is in general restrictive. Note that the measurement error can still depend on $z$, but this is through the true treatment since $D^*=D+(1-2D)\varepsilon$. Assumption (ref) is restricting the indicator of the existence of measurement error to be independent of $z$ but not the measurement error itself. Assumption (ref) requires the existence of an instrument that shifts the probability of selection into treatment. In addition, (ref) says that, even though the propensity scores cannot be recovered from the observed data (because $D$ is unobserved in practice), the ascending order of them in $Z$ is still known. This can be seen as a structural restriction imposed on the true treatment $D$. See TZ.\footnote {An example of sufficient condition is a constant-coefficient latent-index model. That is, suppose the treatment is generated by $D = 1(bZ > e)$, where $b$ is a parameter and $e$ is an error term independent of $Z$. Then, the order of $p(z)$, is determined by the sign of $b$. It is plausible in many applications that the sign of $b$ can be retrieved from economic theory. For example, in the study of the returns to schooling, distance to college is often used as an instrument for completed college education. In this specific example, the parameter $b$ is negative.}. Under assumption (ref), the sign of $p(z)-p(z\prime)$ for any two $z,z\prime$ is known. Under Assumption (ref),we are imposing that

eqnarray*[eqnarray* omitted — 273 chars of source]

With the false-positive probability $P(\varepsilon=1|D=0, Z=z)$ and the false-negative probability $P(\varepsilon=1|D=1, Z=z)$. We introduce the working example that will help interpret the assumptions and results through the rest of the document.

exampleThe researcher is interested in measuring marginal returns of recieving the Supplemental Nutrition Assistance Program (SNAP) on food security. It is well documented that underreporting of SNAP exists. In this case, the variable $Y$ is a binary outcome of being food secure, and $D$ is the true indicator for being a SNAP recipient. The variable $Z$ is the indicator of having certain assets in the household or having cars exempt from an asset test that recipients have to complete (see KPGJ and references therein). The latent variable $V$ could be interpreted as the stigma cost of SNAP as in Welfare2. As stated by Welfare1 stigma is acknowledged as one of the determinants of welfare participation, and there is wide evidence that it negatively affects take-up rates. Let $Y_1$ be the potential food security status for someone on SNAP, and $Y_0$ when the same individual does not receive it. $Y_d$ can be correlated with the stigma cost $V$. As noted by Food1, internalized stigma may lead to food insecurity if it causes or intensifies isolation from social support systems that would allow access to food. Additionally, as stated by Food2 stigma manifestations lead to food inequities through a series of mediating mechanisms experienced and enacted by targets of the stigma that undermine healthy food consumption, contribute to food insecurity, and ultimately impact diet quality. In that sense, psycho-social processes represent how individuals respond to stigma, which ultimately shapes their food selection, purchasing, and consumption behaviors. Enacted and anticipated stigma are characterized as significant stressors, and individuals may cope with these stressors through unhealthy eating behaviors or irrational choices that increase the likelihood of food insecurity. This is then implicitly saying that stigma could be correlated with the potential outcomes. The variable $D^*$ is the individual’s reported (observed) indicator for SNAP recipiency. In this context, the last part of assumption (ref) is consistent with saying that the stigma cost is also determining the misreporting behavior of the individual says $\varepsilon= 1\{f(V)\geq e\}$, if the function of the stigma cost is big enough to pass some threshold $e$ the individual chooses to misreport consistent with HP. Additionally, the assumption is consistent with random misreporting; one could think that individuals make errors when answering the survey question about SNAP recipiency with no intention. In such case $\varepsilon= 1\{f(\eta)\geq 0\}$ where $\eta$ is independent of $V, Y_d$.
remarkImposing that the instrument is independent from the misclassification decision may not be appropriate in many empirical contexts although we claim it is valid here. To illustrate when it is not valid using the current example for instance, if in the SNAP example, the instrument $Z$ is a result of the political forces that regulate SNAP implementation in each state the assumption would not hold. These political forces may influence how people perceive the benefits and costs associated with welfare participation. If those perceived costs are associated with individual willingness to lie about SNAP participation, then $Z$ is not independent of the decision to misreport, implying that the assumption does not hold in this empirical example.

Besides the previously mentioned baseline assumptions, the following assumption is introduced.

assumption[Smoothness] There exists known constants, $b$ such that for any pairs $v_1\neq v_2$ in the support of $V$: \begin{align} \begin{array}{lcl} -b|v_1-v_2|\leq E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b|v_1-v_2| \end{array} \end{align}

KKKL introduces smoothness conditions for $E[Y_d|D=d]$ to bound the $ATE$ without an instrument and treatment exogeneity; this approach has the same spirit. In this case, we can build on their insight to provide bounds for the $MTE$ using similar smoothness conditions. The previous assumption states the degree of smoothness of the marginal treatment responses ($E[Y_d|V=v]$). Generally speaking, we may interpret our identification analysis in this section as a conditional one indexed by $b$. Furthermore, we may conduct a sensitivity analysis by looking at different values of $b$. The parameter $b$ is the Lipschitz constant which serves as a measure of smoothness. In this case we are assuming a maximum level of smoothness $b$. Assumption (ref) is restricting the functional form for the marginal responses, but considering all possible functionals in the lipschitz family with smoothness parameter $b$ or smaller instead of a particular parametric family (like for example linear functions). It is stating the degree of smoothness of the potential responses without assuming a particular functional form of it. In this sense $E[Y_d|V=v]$ could be for example linear $E[Y_d|V=v]=\mu_d+a_dv$ (in which case $b=a_d$) or quadratic $E[Y_d|V=v]=\mu_d+a_dv+c_dv^2$ (in which case $b=a_d+2c_d$) among different possibilities. This assumption introduces constraints in the underlying selection mechanism since is imposing restrictions on how the potential outcomes behave in relationship to the underlying cost of selecting into treatment. The smoothness assumption also relies on the choice of the Lipschitz constant which makes the result sensitive to the choice. This later point is discussed in the online appendix. Assumption (ref) might more appropriately be called something like bounded slope, bounded rate of change, or Lipschitz continuity of the $MTR$ functions but they are directly impacting the degree of parsimony of the functions, so we call it smoothness.\footnote{It is worth noticing that in the standard analysis on the $MTE$, the normalization of $V$ does not change any content of the model, but it does change the interpretation of the Lipschitz condition. This is because it is not the same to impose a Lipschitz condition on the conditional mean of $Y_d$ on $E$ where $E$ has a normal distribution, than to put it on $V=F_E(E)$ which is uniform.} The following remark adapted from KKKL is relevant to understand what this type of assumption is imposing on the marginal treatment responses.

remarkAn alternative way of bounding the rate of change in the marginal treatment responses is to impose further global restrictions in addition to monotonicity such as concavity. The approach used in this paper imposes restrictions directly on the rate of change in its nature, whereas the combination of concavity and monotonicity restricts the rate of change indirectly. There is no clear dominance between each of these ways of imposing restrictions except the belief the researcher has on the behaviour of the marginal treatment responses.

More generally one could say $b^{\prime}|v_1-v_2|\leq E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b|v_1-v_2|$ as stated by KKKL, furthermore, letting $b^{\prime}=0$ and saying $E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b|v_1-v_2|$ for $v_1>v_2$ would be combining monotonicity of the treatment responses with assumption (ref). More precisely:

assumptionThere exists known constants, $b_1,b_0>0$ such that for any pairs $v_1\geq v_2$ in the support of $v$: \begin{align} \begin{array}{lcl} 0\leq E[Y_1|V=v_1]-E[Y_1|V=v_2]\leq b_1(v_1-v_2) \end{array} \end{align} \begin{align} \begin{array}{lcl} 0\leq E[Y_0|V=v_1]-E[Y_0|V=v_2]\leq b_0(v_1-v_2) \end{array} \end{align} Or more generally: \begin{align} \begin{array}{lcl} 0\leq E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b(v_1-v_2) \end{array} \end{align} Where $b=\max\{b_0,b_1\}$
example[Continued] In the context of SNAP, a binary treatment, and food security, a binary outcome, one could model the relationship using a bivariate probit model. Nevertheless, this can be restrictive since it implies a known joint distribution of the unobservables and a parametric index structure. Alternatively, one could choose to allow for all the models with $b \leq 0.5$. This is consistent with the bivariate probit models and allows for more generality by relaxing the normality assumption.

Identification breakdown

Note that following HUV and their standard assumptions ((ref)-(ref) above), without further restrictions the $MTE$ at the level of heterogeneity $v^*$ ($E[Y_1-Y_0|V=v^*]$) is not identified in this setting with discrete instruments and a misreported treatment. From standard results, we get:

eqnarray*[eqnarray* omitted — 158 chars of source]

The second equation takes the difference of the first equation for any two values of the instrument connecting the observed shift in $Y$ caused by changes in $z$ and the underlying treatment effect for all the individuals affected by such a change of the instrument. For any given $z$, say $z\prime$ we can get:

align[align omitted — 180 chars of source]

Where the first equality is because we are conditioning on $D^*=D, D^* \neq D$ and applying the properties of probabilities. This last equation reflects that the observed propensity score for the proxy of the true treatment variable conditional on $z\prime$ equals the share of treated individuals who at that particular $z\prime$ report treatment status correctly multiplied by the probability of reporting correctly, plus the share of not treated individuals who at that particular $z\prime$ report treatment status incorrectly multiplied by the probability of reporting incorrectly. The previous expressions depend on unobserved components. While $ E[Y|Z=z\prime], P[D^*=1|Z=z\prime]$ are observed, $P[D^*=D|Z=z\prime],P[D=1|D^*=D,Z=z\prime]$ and $p(Z)$ are not, which without further assumptions do not allow for identification of the true propensity score and also of the $MTE$. If $p(z)$ was observed and $z$ was continuous, then the $MTE(v^*)$ would be identified as $\frac{\partial E[Y|P(Z)=v^*] }{\partial v^*}=\frac{\int_0^{v^*}E[Y_1|V=v]dv + \int_{v^*}^1E[Y_0|V=v]}{\partial v^*}=E[Y_1|V=v^*]-E[Y_0|V=v^*]$. So this displays the two main identification challenges, the non-continuity of $z$ and the fact that $p(z)$ is not observed.

Before proceeding to the identification results, it is worth showing the main elements of the current work and how they differentiate from previous identification results of $MTE$ with discrete instruments. It is also relevant to show the role of misreporting. From the observed data if there is no misreporting from assumptions (ref)-(ref) one can identify:

eqnarray*[eqnarray* omitted — 147 chars of source]

The first equality comes from the definition of the model, the second one from the laws of probability, the third one by the independence of $Z$ from $Y_1, V$ and the last one from the properties of conditional expectations and the normalization that $V$ is marginally uniform. Similarly, we have:

eqnarray*[eqnarray* omitted — 60 chars of source]

This then implies the equality expressed at the beginning of this subsection:

eqnarray[eqnarray omitted — 95 chars of source]

Without misreporting the propensity score $p(z)$ is identified and, given assumptions (ref)-(ref) index sufficiency holds and thus $E[YD|Z=z]=E[YD|p(Z)=p]$. So we can rewrite the previous equalities as functions of $p$ instead of $z$. Where $p \equiv p(z)$. In this context without differentiability of $p$ the key insight from BMW is to introduce parametric restrictions that for example say that $E[Y_d|V=v]=\mu_d +a_d v$ and thus $E[Y_1-Y_0|V=v]=\mu+a v$, where $\mu\equiv \mu_1-\mu_0$ and $a \equiv a_1-a_0$. Additionally define $c=\mu_0+\frac{a_0}{2}$. In this case

eqnarray*[eqnarray* omitted — 151 chars of source]

Then from $E[YD|p(Z)=p]$ for different values of $p$ (at least two which is enough with a binary instrument) we can solve for $\mu_1,a_1$. Similarly for $a_0,\mu_0$ from $E[Y(1-D)|p(Z)=p]$. Note that then given the marginal treatment responses, $E[Y_d|V=v]$ is identified, so it is the $MTE$ as their difference. One might not be willing to assume particular parametric specifications for the conditional expectations of the potential outcomes since they are restrictive. One of the contributions of the current work is relaxing such restrictions and still recovering analytically tractable expression for the bounds of the $MTE(v^*)$. The current work relates MST in the following way. MST relies on recovering the set of marginal treatment responses consistent with observed $IV$-like estimands. In this setting, such strategy would rely on finding all the candidates $E[Y_d|V=v]$ functions consistent with:

eqnarray*[eqnarray* omitted — 427 chars of source]

Their strategy relies on the fact that $w_1 w_0$ are known (or identified). In the case of misreporting, where we do not know exactly the rate of false positives and false negatives for every value of $z$, we have that $p(z)$, and thus, the weights are not identified. This makes the current work to differ from the existing literature since developed computational methods rely on the weights being known or identified. In this context one of the main contributions of the current paper is working in the context where the weights are not identified but actually can be partially identified. In section (ref) we start from the same insight as MST, but instead of solving a linear problem, we aid identification with shape restrictions to get analytical bounds on the $MTE$. In the online appendix an extension of MST is discussed without aiding identification with shape restrictions by solving the same linear problem as in MST but for different values of $p(z)$ score in the identified set. As the identified set of $p(z)$ is not finite, the solution can only be approximated.

Identification results

In subsection (ref) identification without misreporting will be discussed. Subsection (ref) incorporates misreporting.

Identification without misreporting

Note than since

eqnarray*[eqnarray* omitted — 84 chars of source]

For any values $z$ and $z\prime$ with $p(z) \geq p(z\prime)$:

eqnarray*[eqnarray* omitted — 383 chars of source]

Where we are adding and subtracting the marginal treatment responses ($E[Y_d|V=v^*]$) related to the marginal treatment effect at the $v^*$ of interest ($E[Y_1|V=v^*]-E[Y_0|V=v^*]$), then using $E[Y_d|V=v]-E[Y_d|V=v^*]\leq b|v-v^*|$ twice and also the definition of $MTE$. Similarly we can get:

eqnarray*[eqnarray* omitted — 114 chars of source]

Then:

eqnarray*[eqnarray* omitted — 204 chars of source]

The bounds depend on the propensity score and the difference of the propensity score for different values of $z$.

remarkNote that if one integrates the bounds for the $MTE$ one does not point identify $LATE$. In particular, integrating $v^*$ over $p(z\prime)$ and $p(z)$ yields the following bounds for $LATE$: \begin{eqnarray*} \frac{E[Y|Z=z]-E[Y|Z=z\prime]-\frac{2b}{3}(p(z)-p(z\prime))^3}{p(z)-p(z\prime)} \leq LATE(p(z\prime),p(z)) \leq \frac{E[Y|Z=z]-E[Y|Z=z\prime]+\frac{2b}{3}(p(z)-p(z\prime))^3}{p(z)-p(z\prime)}. \end{eqnarray*} This is because the way the bounds are derived, a quantity that is bigger (or smaller) of the numerator of $LATE$ is central to derive the bounds for the $MTE$. The method to derive bounds for the $MTE$ is not exactly an extrapolation of $LATE$ since there is no unique way to do that given assumption (ref). More precisely, the way the bounds are computed, assumption (ref) is applied at every point $v^*$ without consideration of the joint restrictions for pairs of evaluation points such as $v^*,v^{**}$. Additionally, implications of (ref) are not necessarily fully exploited in the constructive identification approach. In this sense, the bounds are neither functional nor point-wise sharp.
remarkThen the partial identification analysis of the $ATE$ starts from the well-known Manski’s worst-case bound. This formulation of the identification region reveals that the identification power becomes weak when the upper and lower bounds for $Y_d$ are large. An advantage of the proposed method is that it does not require the existence of upper and lower bounds (although it requires a tuning parameter $b$). In this case as shown in the online appendix, for example, the $ATE$ can be bounded above by: \begin{eqnarray*} ATE&\leq&\frac{E[Y|Z=z]-E[Y|Z=z\prime]}{p(z)-p(z\prime)}+\frac{2b}{3}\frac{p(z)^3-p(z\prime)^3}{p(z)-p(z\prime)}-b\frac{p(z)^2-p(z\prime)^2}{p(z)-p(z\prime)}+b \end{eqnarray*} If $Y_d$ is bounded, then the previous bound on the $ATE$ is complemented with the Manski worst case bounds: \begin{eqnarray*} ATE&\leq&\min\{ Y_{u1}-Y_{l0} ,\frac{E[Y|Z=z]-E[Y|Z=z\prime]}{p(z)-p(z\prime)}+\frac{2b}{3}\frac{p(z)^3-p(z\prime)^3}{p(z)-p(z\prime)}-b\frac{p(z)^2-p(z\prime)^2}{p(z)-p(z\prime)}+b \} \end{eqnarray*} Note there is a $b$ such that the proposed bounds are numerically the same as the worst case bounds in the case the outcome variable is bounded.

Identification of the MTE with misreporting

In order to identify the $MTE$ first we need to identify $p(z)$. In AK such identification is discussed. Subsequent subsections, builds upon the results from AK. For clarity, let $\Delta_p \equiv p(z)-p(z'), \Delta_{D^*Z}(z',z) \equiv P(D^*=1|Z=z)-P(D^*=1|Z=z'), \alpha \equiv P(\varepsilon=1) $, then let the bounds derived in AK be:

eqnarray*[eqnarray* omitted — 355 chars of source]

The bounds on the propensity score rely on any given level of unconditional misreporting ($\alpha$). Misreporting enters in two ways. First of all, if the probabilities of misreporting are unknown, $p(z)$ is no longer identified (see AK for more details), which then creates a problem since such quantity appears systematically in the bounding strategies. To solve this problem, we use the previously defined bounds on the misreporting probabilities. The lack of point identification of $p(z)$ affects the strategy using smoothness restrictions. See for example that:

eqnarray*[eqnarray* omitted — 115 chars of source]

The bound of $MTE$ depends on the sign of it since it is not always true that $MTE(v^*)(p(z)-p(z\prime)) \leq MTE(v^*)(p_u(z)-p_l(z\prime))$. These considerations are taken into account in theorem (ref). Smoothness assumptions for marginal treatment responses

The bounds from theorem (ref) builds on the following inequalities due to assumption (ref)

eqnarray[eqnarray omitted — 127 chars of source]
eqnarray[eqnarray omitted — 128 chars of source]

Where the inequalities use the smoothness assumption as in the previous section without misreporting and fact that $\int_{p(z\prime)}^{p(z)}2b|v-v^*|dv \leq \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv$. Note that we have a sufficient condition for identifying the sign of the $MTE(v^*)$. If $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$ then the $MTE(v^*)$ is positive. If $E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\leq 0$ then the $MTE(v^*)$ is negative. If $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$ then from equations (ref) and (ref) combined with the bounds on $p(z)-p(z\prime)$ we get:

eqnarray*[eqnarray* omitted — 223 chars of source]

Note that the form of the bounds depend on $|v-v^*|$. In the integral of the absolute value is where the relative position of $v^*$ with respect of $p_l(z\prime),p_u(z)$ will matter. Note that if the $v^*$ of interest is such that $v^*\leq p_l(z\prime)$, the integral involving $|v-v^*|$ is $ [\frac{p_u(z)^2-p_l(z\prime)^2}{2}+v^*(p_l(z\prime)-p_u(z))]$. If the $v^*$ of interest is such that $v^*\geq p_u(z)$ then $[\frac{-p_u(z)^2+p_l(z\prime)^2}{2}+v^*(-p_l(z\prime)+p_u(z))]$. If the $v^*$ of interest is such that $p_l(z\prime)\leq v^*\leq p_u(z)$ then $[v^{*2}-v^*(p_u(z)+p_l(z\prime))+(\frac{p_u(z)^2+p_l(z\prime)^2}{2})]$

The following theorem summarizes the previous discussion.

theoremIf assumptions (ref)-(ref) and (ref) holds. Then the following bounds are valid: \begin{enumerate} • If $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$: \begin{eqnarray*} MTE^+(v^*)_{lb}=\frac{E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv}{\Delta_{pu}} \\ MTE^+(v^*)_{ub}=\frac{E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv}{\Delta_{pl}} \end{eqnarray*} • If $E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\leq 0$: \begin{eqnarray*} MTE^-(v^*)_{lb}=\frac{E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv}{\Delta_{pl}} \\ MTE^-(v^*)_{ub}=\frac{E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv}{\Delta_{pu}} \end{eqnarray*} • Otherwise: \begin{eqnarray*} MTE(v^*)_{lb}=\max\{MTE^-(v^*)_{lb},MTE^+(v^*)_{lb}\} \\ MTE(v^*)_{ub}=\min\{MTE^-(v^*)_{ub},MTE^+(v^*)_{ub}\} \end{eqnarray*} \end{enumerate}
remarkIn some situations like in the case of SNAP, one could be willing to assume that for every level of heterogeneity $v$, $P(Y_1<Y_0|V=v)=1$ holds which means that receiving SNAP is not making anyone more food insecure. This is the “treatment cannot hurt” assumption or known as the monotone treatment response assumption in the partial identification literature. In such a case, we would be imposing the sign of the $MTE$ even if we cannot extract it from $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$ or $E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\leq 0$
remarkThe previous bounds are not assuming there is a known support for $Y$ if the nature of $Y$ is bounded then the previous bounds change in the following way: \begin{eqnarray*} \Tilde{MTE}(v^*)_{lb}&=&\max\{MTE(v^*)_{lb},Y_{1l}-Y_{0u}\} \\ \Tilde{MTE}(v^*)_{ub}&=&\min\{MTE(v^*)_{ub},Y_{1u}-Y_{0u}\} \end{eqnarray*}

The previous theorem is extended for discrete instruments taking more than 2 values in the online appendix. A researcher might be interested in combining the monotonicity assumption on the treatment responses and the smoothness assumption. So instead of using assumption (ref), the researcher might be willing to use (ref). This result is collected in the online appendix. The choice of $b$ is not arbitrary. The fact that it operates as a tuning parameter might lead to a discretionary use of it to get the desired result; different ways of choosing this parameter are discussed in the online appendix. The previous identification results are illustrated in the appendix.

Application: $MTE$ of SNAP on child health when participation is endogenous and misreported

In this section, the developed methods are applied to get bounds on the $MTE$. We then integrate them over $v^*$ to get bounds on the $ATE$ of receiving SNAP on the outcome of being food insecure. As stated by KPGJ SNAP, formerly known as the Food Stamp Program, is by far the largest food assistance program in the United States and, as such, constitutes a crucial component of the social safety net in the United States. In any given month during 2009, SNAP assisted more than 15 million children, and it is estimated that nearly one in two American children will receive assistance during their childhood. Concluding about the program's impact is complex due to two of the fundamental problems studied in this paper. First, a selection problem arises because the decision to participate in SNAP is unlikely to be exogenous. On the contrary, unobserved factors such as expected future health status, parents’ human capital characteristics, financial stability, and attitudes towards work and family are all thought to be jointly related to participation in the program and health outcomes such as food security. Families may decide to participate precisely because they expect to be food insecure or in poor health. Second, a nonrandom measurement error problem arises because a large fraction of food stamp recipients fails to correctly report their program participation in household surveys. Using administrative data matched with data from the Survey of Income and Program Participation (SIPP), for example, BD find that errors in self-reported receipt of food stamps exceed 12 percent and are related to respondents’ characteristics, including their true participation status, health outcomes, and demographic attributes. MMS provide evidence of extensive underreporting of food stamps in the SIPP, the Current Population Survey (CPS), and the Panel Study of Income Dynamics (PSID). In this context KPGJ studies the average effects using the December Supplement of the 2003 Current Population Survey (CPS). In the data, we can observe a self-reported measure of food stamp receipt over the past year, food insecurity over the past year, and the ratio of income to the poverty line.\footnote{For further details about the data see GK. In there, they state that just over 40 percent of the households report receiving food stamps, and the food insecurity rate among self-reported recipients is 17.9 percentage points higher than among eligible non-recipients (52.3 percent vs. 34.4).} As KPGJ states, the data is rich enough to allow the construction of instrumental variables for SNAP participation used in previous literature. In particular, state identifiers in the CPS apply a more traditional instrumental variable (IV) assumption based on cross-state variation in program eligibility rules.\footnote{In general terms, program eligibility rules are income requirements (most households must meet both gross and net income limits to qualify for SNAP benefits), resource requirements (households must also meet a resource limit in their bank accounts), work requirements (If you are an able-bodied adult without dependents, between the ages of 18 and 49, and able to work but currently unemployed, you may only be eligible for SNAP benefits for three months within a three-year period) and other eligibility requirements (to be eligible for SNAP benefits, households must also, meet other conditions in addition to the income and resource requirements, such as everyone in your household having, or have applied for, a social security number). To establish the income and resource requirements, each state computes asset tests, but there is variation in how these states evaluate the assets of individuals. More specifically, they may or may not include certain assets that will affect the individuals' eligibility.} Merging the Urban Institute’s database of state program rules with the CPS data KPGJ create two instrumental variables: an indicator for whether the state uses a simplified semi-annual reporting requirement for earnings and an indicator for whether cars are exempt from the asset test.\footnote{For more details on the construction see KPGJ.} These instrumental variables if valid and independent allows using the current methods to bound the $MTE$. For example, if one is willing to assume that the state variation in the asset test is exogenous, then instrument independence is satisfied. It is worth noticing that only $LATE$ can be identified from an instrumental variable regression under individual heterogeneity. In this sense, the methods developed here can be used to recover bounds on the average treatment effect, and complement KPGJ results. In this context, the $MTE$ at a particular level of $v^*$ represents the treatment effect receiving SNAP has for a particular level of stigma. Stigma mat is connected with the potential outcomes since, as stated by Food2 stigma manifestations affect health outcomes (and food security as such). More precisely, in this context to illustrate our methods, $Y$ is a food insecurity indicator over the past year, $D^*$ is the self-reported SNAP participation (subject to potential measurement error as stated before), $Z$ is a binary variable for cars exempt from the asset test. Concerning choosing $b$ for the sake of exposition, we present the results here for several potential values of $b$. In the online appendix we discuss different methods on how to choose $b$ which could be used in this context. Intuitively choosing $b$ is restricting the degree of smoothness the $MTE$ would have. One can draw a parallelism between choosing a linear form (a low $b)$ for the $MTE$ versus choosing a high order polynomial (high $b$). The linear form is rather restrictive on the behavior of higher-order derivatives (and thus smoothness) compared with the polynomial. A researcher choosing $b$ for SNAP should consider what he thinks the underlying decision-maker optimization problem looks like. If the expected utility, for example, has a quadratic form, or the optimal expected demand of food security is linear under both receiving and not receiving SNAP, then the expected benefit on the optimal choice of food security both under getting SNAP and not getting it could be considered linear. Consistently with KPGJ a treatment cannot hurt assumption ($P(Y_1<Y_0|V=v)=1$) will be introduced. Making then the conservative upper bounds of the $MTE$ to be $0$. The following table summarizes the data and shows the average and median characteristics in the sample.

table[table omitted — 439 chars of source]

We can see that around forty-two percent of the people are food insecure, while forty-one report receiving SNAP. Thirty percent have their cars excluded from the asset test, which implies that thirty percent of the individuals in the sample live in states where cars are excluded from the asset test. On average (and in the median), individuals in this sample are below the poverty line. Each of the graphs in the following figure is computed in the following way. For any given level of $b$ and $\alpha$, point estimates of the bounds are constructed for the $MTE$ at different values $v^*$ using their sample analogs. To decide which type of bound to use, the relative position of the sample analog estimates of the bounds for the propensity score is calculated. The maximum level of $\alpha$ is twenty percent which is chosen as an arbitrary big upper bound above the existing results from KPGJ. See Figure (ref).

In orange, we have the upper and lower bounds assuming $\alpha=0.2$. In blue, we have the upper and lower bounds assuming $\alpha=0.1$, and in red, we report the upper and lower bounds assuming $\alpha=0$ (no misreporting). The limits of the $Y$-axis are the worst-case upper bound ($0$ under treatment cannot hurt) and the worst-case lower bound ($-1$). The $X$-axis goes from $0$ to $1$, the different potential values of $v^*$. We can see that the bounds have identification power over different regions of $v^*$ support. We can see that when the level of misreporting decreases, the bounds become tighter since there is less lack of identification due to misreporting. Similarly, when $b$ decreases, we also get tighter bounds consistent with reducing the potential functional forms of the marginal treatment responses. These bounds are computed by fixing a level of $\alpha$. If, for example, in the case of $b=0.1$, the researcher is not interested in $\alpha=0.2$ rather in $\alpha \leq 0.2$ then all the region between the upper orange curve and the lower orange curve is the identified set consistent with $b=0.1$ and all the $\alpha$'s less or equal to twenty percent. The length of the identified set becomes tighter in the region between the estimated observed propensity scores ($\widehat{P}[D^*=1|Z=1]=0.49,\widehat{P}[D^*=1|Z=0]=0.38$) since there is more information being used to bound the $MTE$. The previous display also implies a simple way of computing bounds of a parameter of interest such as the $ATE$. It is known that $ATE=\int_{0}^1MTE(v^*)dv^*$. Then we know that $\int_{0}^1MTE_{lb}(v^*)dv^* \leq ATE \leq \int_{0}^1MTE_{ub}(v^*)dv^*$. So then we can approximate bounds for the $ATE$ as:

eqnarray*[eqnarray* omitted — 178 chars of source]

Where $N_{v^*}$ is the number of grid points where $MTE(v^*)$ was evaluated and where $\widehat{MTE}$ are the bounds from the previous graphs. On the online appendix estimates of the bounds on the $ATE$ are computed. In the online appendix, there is also an alternative way of estimating (and doing inference) on the $ATE$ for the case of smoothness restrictions. On the online appendix a method for asymptotic normality for an outer-set of the $ATE$ in the case of no misreporting and with smoothness conditions is developed. On the online appendix a method for asymptotic normality for an outer-set of the $ATE$ in the case of misreporting, treatment cannot hurt assumption, and smoothness conditions can be found. The $ATE$ bounds are easily estimated and used for inference, since as shown in the appendix, each component can be replaced by their sample analogs, which themselves are asymptotically normal, and thus, by the continuous mapping theorem the upper bounds and lower bounds for the $ATE$ are also. Then, an asymptotically valid bootstrap procedure can be used to build confidence intervals for the entire identified set, such as those constructed by MN. The population identification region is an interval $[L, U]$, we can estimate each side of the interval with consistent and asymptotically normal estimators $\widehat{L}, \widehat{U}$ and via this procedure we get confidence interval $[\widehat{L}-z_{\frac{\alpha+1}{2}}\widehat{\sigma}_l/\sqrt{N}, \widehat{U}+z_{\frac{\alpha+1}{2}}\widehat{\sigma}_u/\sqrt{N}] \equiv [\widehat{L}_{\alpha}, \widehat{U}_{\alpha}]$ such that

eqnarray*[eqnarray* omitted — 125 chars of source]

In order to control for covariates $X$ such that independence of $Z$ holds conditional on it in a tractable manner, we can assume $E[Y_d|V=v,X=x]=m_d(v)+\beta X$. Then,

eqnarray*[eqnarray* omitted — 165 chars of source]

Where $P(z,x)=P(D=1|X=x,Z=z)$. We can then any two values of $Z$:

eqnarray*[eqnarray* omitted — 101 chars of source]

We can then follow a similar display is in Section (ref). In this case, we can estimate $E[Y|Z=z,X=x]$ with a partial linear regression while $p(z,x)$ is estimated with a non-parametric regression. So far, we have reported results for the instrument taking only two values. The data set used in this problem counts with two potential instruments. The already used one, and an eligibility criterion specifying if earners report twice a year or not. Based on this, a three-valued instrument can be constructed related to the intensity of the likelihood of receiving SNAP. That is, it takes $0$ if both instruments take the value $0$, takes the value $1$ if either of them takes the value $1$, and it takes the value $2$ if both of them take the value $1$. The details of how the bounds look for more discrete non-binary instruments and the particular case of an instrument taking three values are collected in the online appendix. In such a case, the length of the identified set for the $ATE$ becomes smaller; this is intuitive since we now have a more exogenous variation to exploit. We we illustrate it with the case of $b=0.5, \alpha=0.1$. Additionally, we can see that the form of the identified set of the $MTE$ changes, since now the regions rely on the different exogenous variation in zones where previously, only the shape restrictions could be used. These results are collected in the online appendix.

Conclusions

In this paper, we provided partial identification results for the Marginal Treatment Effect in the presence of measurement error and a discrete instrument building over MST, BMW and AK. To do so, given the discrete nature of the instruments, we introduced smoothness restrictions. Results are illustrated via a numerical example and quantifying the marginal treatment effect of SNAP on food insecurity, a case in which measurement error and endogeneity of treatment are known to be an issue. In a more general way, our results can serve as a sensitivity analysis tool for when researchers are interested in recovering the $MTE$ in the presence of a discrete instrument and suspect measurement error, and have doubts about their parametric assumptions. This sensitivity analysis is executed by varying $\alpha, b$. If no measurement error exists, the results from this paper provided analytical partial identification results of the $MTE$ in the presence of discrete instruments that can serve as a complement for the already existing results. \singlespace

thebibliography{7} \bibitem[Acerenza et al.(2021)]{AK} Acerenza, S., K. Ban and D. K\'edagni. 2021. "Marginal Treatment effects with misclassified treatment." Working paper. \bibitem[Angrist(1990)]{AKK2} Angrist, J. D. 1990. "Lifetime Earnings and the Vietnam Era Draft Lottery: Evidence from Social Security Administrative Records" The American Economic Review 80(3):313-336. \bibitem[Angrist and Evans(1998)]{AKK3} Angrist, J. D. and W. N. Evans. 1998. "Children and Their Parents' Labor Supply: Evidence from Exogenous Variation in Family Size" The American Economic Review 88(3):450-477. \bibitem[Angrist and Krueger(1991)]{AKK} Angrist, J. D. and A. B. Krueger. 1991. "Does Compulsory School Attendance Affect Schooling and Earnings?" The Quarterly Journal of Economics 106(4):979-1014. \bibitem[Armstrong and Koles\'ar(2020)]{AK2} Armstrong, T and M. Koles\'ar. 2020. "Simple and honest confidence intervals in nonparametric regression." Quantitative Economics:1-39. \bibitem[Bollinger and David(1997)]{BD} Bollinger, C., and M. David. 1997. "Modeling Discrete Choice with Response Error: Food Stamp Participation." Journal of the American Statistical Association 92(439): 827-835. \bibitem[Brinch et al.(2017)]{BMW} Brinch, C.N. , M. Mogstad and M. Wiswall. 2017. "Beyond $LATE$ with a Discrete Instrument." \textit{Journal of Political Economy} 125(4): 985-1039. \bibitem[Calvi et al.(2021)]{CLT} Calvi, R. , A. Lewbel and D. Tommasi. 2021. "LATE With Missing or Mismeasured Treatment." \textit{Journal of Business and Economic Statistics}. \bibitem[Contini and Richiardi(2012)]{Welfare1} Contini, D. and A. M. Richiardi. 2012. "Reconsidering the effect of welfare stigma on unemployment." \textit{Journal of Economic Behavior and Organization} 84(2):224-244. \bibitem[Earnshaw and Karpyn(2020)]{Food2} Earnshaw, V. and A. Karpyn. 2020. "Understanding stigma and food inequity: a conceptual framework to inform research, intervention, and policy." \textit{ Translational Behavioral Medicine} 10(6): 1350–1357. \bibitem[Gundersen and Kreider(2008)]{GK} Gundersen, C., and. B. Kreider. 2008. "Food Stamps and Food Insecurity: What Can Be Learned in the Presence of Nonclassical Measurement Error?" \textit{Journal of Human Resources} 43(2): 352-382. \bibitem[Haider and Stephens(2020)]{Hai} Haider, S. and Stephens M. 2020. "Correcting for Misclassified Binary Regressors Using Instrumental Variables." \textit{NBER Working paper series} Working Paper 27797. \bibitem[Hausman et al.(1998)]{HAS} Hausman, J.A., Abrevaya, J. and Scott-Morton, F.M. 1998. "Misclassification of the dependent variable in a discrete-response setting." \textit{Journal of Econometrics} 87:239–269. \bibitem[Heckman and Vytlacil(1999)]{HV} Heckman, J.J. and E. Vytlacil. 1999. "Local Instrumental variables and latent variable models for identifying and bounding treatment effects." \textit{Proceedings of the National Academy of Sciences} 96:4730–4734. \bibitem[Heckman and Vytlacil(2005)]{HV2} Heckman, J.J. and E. Vytlacil. 2005. "Structural Equations, Treatment Effects, and Econometric Policy Evaluation." \textit{Econometrica} 73(3):669–738. \bibitem[Heckman et al.(2006)]{HUV} Heckman, J.J., S. Urzua and E. Vytlacil. 2006. "Understanding Instrumental Variables in models with essential heterogeneity." \textit{The Review of Economics and Statistics} 88(3):389–432. \bibitem[Hernandez and Pudney(2007)]{HP} Hernandez, M. and S. Pudney . 2007. "Measurement error in models of welfare participation." \textit{Journal of Public Economics} 91:327–341. \bibitem[Kim et al.(2018)]{KKKL} Kim, W., K. Kwon, S. Kwon and S. Lee. 2018. "The identification power of smoothness assumptions in models with counterfactual outcomes." \textit{Quantitative Economics} 9, 617–642. \bibitem[Kreider et al.(2012)]{KPGJ} Kreider, B., J.V. Pepper, C. Gundersen and D. Jolliffe. 2012. "Identifying the Effects of SNAP(Food Stamps) on Child Health Outcomes When Participation is Endogenous and Misreported." \textit{Journal of the American Statistical Association} 107:432–441. \bibitem[Krueger(1999)]{AKK4} Krueger, A. B. 1999. "Experimental Estimates of Education Production Functions" \textit{The Quarterly Journal of Economics} 114(2):497-532. \bibitem[Machado et al.(2019)]{MSV} Machado, C., A.M. Shaikh and E.J. Vytlacil. 2019. "Instrumental variables and the sign of the average treatment effect." \textit{Journal of Econometrics} 212(2):522–555. \bibitem[Manski and Nagin(1998)]{MN} Manski, C. F., and D. Nagin. 1998. "Bounding Disagreements About Treatment E§ects: A Case Study of Sentencing and Recidivism." \textit{Sociological Methodology} 28:99–137. \bibitem[Matsen and Poirier(2021)]{MP} Masten, M.A. and Poirier, A. 2021. "Salvaging Falsified Instrumental Variable Models." \textit{Econometrica} 89:1449–1469. \bibitem[Meyer et al.(2009)]{MMS} Meyer, B., D.W. Mok and J.X. Sullivan. 2006. "The Under-Reporting of Transfers in Household Surveys: Its Nature and Consequences." \textit{Working Paper} University of Chicago, Harris School of Public Policy Studies. \bibitem[Moffitt(1983)]{Welfare2} Moffitt, R. 1983. "An Economic Model of Welfare Stigma." \textit{The American Economic Review} 73(5):1023–1035. \bibitem[Mogstad et al.(2018)]{MST} Mogstad, M., A. Santos and A. Torgovitsky. 2018. "Using Instrumental Variables For Inference About Policy Relevant Treatment Parameters." \textit{Econometrica} 86(5):1589–1619. \bibitem[Palar et al.(2018)]{Food1} Palar, K., Frongillo, E. A., Escobar, J., Sheira, L. A., Wilson, T. E., Adedimeji, A., Merenstein, D., Cohen, M. H., Wentz, E. L., Adimora, A. A., Ofotokun, I., Metsch, L., Tien, P. C., Turan, J. M., and Weiser, S. D. 2018. "Food Insecurity, Internalized Stigma, and Depressive Symptoms Among Women Living with HIV in the United States." \textit{AIDS and behavior} 22(12):3869–3878. \bibitem[Possebom(2021)]{VP} Possebom, V. 2021. "Crime and Mismeasured Punishment: Marginal Treatment Effect with Misclassification." \textit{Working Paper}. \bibitem[Ratcliffe and McKernan(2010)]{RM} Ratcliffe, C. and McKernan, S-M. 2010. "How Much Does Snap Reduce Food Insecurity?" \textit{Contractor and Cooperator Report United States Department of Agriculture (USDA)} 60. \bibitem[Tommasi and Zhang(2020)]{TZ} Tommasi, D and L. Zhang. 2020. "Bounding Program Benefits When Participation Is Misreported." \textit{IZA IZA DP No. 13430} \bibitem[Ura(2018)]{U} Ura, T. 2018. "Heterogeneous treatment effects with mismeasured endogenous treatment." \textit{Quantitative Economics} 9(3):1335–1370.

\setcounter{table}{0} \setcounter{figure}{0} \setcounter{equation}{0}