Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,930 characters · 10 sections · 76 citation commands
Partial Identification of Marginal Treatment Effects With Discrete Instruments and Misreported Treatment
\setstretch{1}
} \newsavebox{\tablebox} \newlength{\tableboxwidth}
\ Keywords: Treatment effects, instrumental variables, measurement error, partial identification. JEL Codes: C21, C26. Word Count: 11688. \doublespacing
This paper provides partial identification results for the Marginal Treatment Effect ($MTE$) in the presence of measurement error in the treatment variable when only a discrete instrument is available. The discrete instrument case is relevant as many applications in the literature rely on these type of instruments. See for example AKK, AKK2, AKK3 and AKK4. The discrete nature of the instrument requires identification strategies to recover the $MTE$ that differ from those explored in the previous literature with continuous instruments. The results of this paper are relevant since it is often true that researchers have access to an instrument with discrete variation (for example, assignment to treatment via an institutional rule), and it is also true that misreporting is a common problem in survey data which is one of the main sources of empirical research. In a more general way, our results can serve as a sensitivity analysis tool for when researchers are interested in recovering the $MTE$ in the presence of a discrete instrument and suspect measurement error and have doubts about their parametric assumptions. Researchers mostly work with self-reported data from surveys; such data systematically present reporting problems that lead to measurement error of the treatment status and, consequently, to bias in the treatment effect of interest. The combination of measurement error with discrete instruments has not been explored in the literature, and it is a fairly common situation to encounter. The results in this paper are useful for identifying MTE (which can be used to recover average effects or policy-relevant effects) in the presence of the two previously mentioned problems for identification. In most cases, researchers observe a discrete (often binary) instrument such as assignment to treatment. In these cases, point identification of the $MTE$ (even without measurement error) is not possible, relying only on the standard assumptions of instrument exogeneity and relevance (See for example BMW). In this paper, under a set of restrictions on the severity of measurement error and shape restrictions, we provide partial identification results for the $MTE$ in the presence of measurement error when a discrete instrument is available. The $MTE$ can help reveal the heterogeneity in the treatment effect. The $MTE$ is relevant in recovering Policy Relevant Treatment Effect parameters ($PRTE$s), Average Treatment Effect ($ATE$), Average Treatment on the Treated ($ATT$), Average Treatment on the Untreated ($ATU$), Local Average Treatment Effects ($LATE$), etc.\footnote{See HV2, HUV, who show the link between the $MTE$ and those parameters via properly weighting the $MTE$.} To achieve partial identification, we introduce smoothness conditions on the marginal treatment responses ($E[Y_d|V=v]$). To deal with the misreporting of the binary treatment, the analysis relies on treating the unconditional probability of misreporting as given.\footnote{One could alternatively take the results from this paper and assume a known upper bound of this probability and take the union of the bounds derived here.} This can be either interpreted as the researcher having prior knowledge on the possible value of the misclassification rates or as a sensitivity analysis tool where the researcher allows for the possibility of misclassification up to a certain level. Relevance and independence of the instrument is required. Although partial identification of the $MTE$ will do not imply in general sharp bounds on the $ATE$. It is still a useful tool to move from local effects and generate bounds on an aggregate relevant effect. Empirical research usually combines a measurement error problem with endogeneity and heterogeneity. U documents in his work, as an example of this, that there is a substantial measurement error in educational attainments in the 1990 U.S. Census. At the same time, educational attainments are endogenous as treatment variables in return to schooling analyses because, among other possibilities, unobserved individual ability affects both schooling decisions and wages. Labor supply response to welfare program participation, in which the outcome is employment status, and the treatment is welfare program participation is subject to similar issues. Self-reported program participation in survey datasets can be misreported as stated by HP. The psychological cost of welfare program participation affects job search behavior and welfare program participation simultaneously.
This subsection lists some relevant papers related to the current research paper based on their connections to different aspects of the problem. Namely, misreporting and partial identification of marginal treatment effects. Partial identification of $LATE, MTE$ and $ATE$ with endogenous misreported binary treatments and heterogeneous effects U using a binary instrumental variable, derives bounds for $LATE$ with a binary misreported treatment when an instrument is available, and monotonicity of the true (not observed) treatment in the instrument holds. Identification is achieved by exploiting the relationship between the probability of being a complier and the total variation distance\footnote{The total variation distance between two probability measures $P$ and $Q$ on a sigma-algebra $\mathcal {F}$ of subsets of the sample space $\Omega$ is defined via $\delta (P,Q)=\sup _{A\in {\mathcal {F}}}\left|P(A)-Q(A)\right|$. It can alternatively be defined for probability measures that have densities to be $\frac{1}{2} \int |p-q|d\nu$ where $\nu$ is a measure dominating both probability measures. In the context of U paper the total variation distance calculated is $TV_{Y,D}=\frac{1}{2}\int\Big(\sum_{d=0,1}|f_{y,d|Z=1}(y,d)-f_{y,d|Z=0}(y,d)|\Big)d\nu$ where $f_{y,d|z}$ is the joint density of the observed outcome variable and the observed treatment variable conditional on the value of the instrument.} between people assigned to treatment and the ones that are not. The under-identification for $LATE$ is a consequence of the under-identification for the size of compliers; with no measurement error, one could compute the size of compliers based on the measured treatment and, therefore, $LATE$ would be the Wald estimand. The total variation distance plays a key role in determining the sharp identified set in U. First, it measures the strength of the instrumental variable; when the total variation distance is positive, the identified set of $LATE$ is a strict subset of the whole parameter space, which implies that $Z$ has some identifying power. Secondly, as shown in U lemma 3, the total variation distance is a lower bound for the proportion of compliers which is the under-identified element in the presence of measurement error. CLT, TZ extend U's results for the case where the instrument can take multiple discrete values. AK focuses on bounding the marginal treatment effects when there is a continuous instrument. KPGJ using auxiliary information about the possibility of misreporting and under different combinations of the outcome, treatment, and instrumental monotonicity bounds the $ATE$ for a binary outcome. VP focuses on partially identifying the $MTE$ with a continuous instrument and imposing sign and functional relationships between the derivatives of the true propensity score and the observed one with respect to the continuous instrument. This current paper complements the previously mentioned papers. Fundamentally this paper focuses on identifying the $MTE$ when discrete instruments are available. Such a task requires a different set of assumptions than the ones used to recover directly $ATE$, $LATE$, or $MTE$ with continuous instruments. We complement U, KPGJ and TZ because we are interested in identifying $MTE$ (which can then be used to achieve identification of $LATE$ and $ATE$) instead of the $LATE$ and $ATE$. It is also complementing AK since their analysis relies on the continuity of the instrument. It is worth noticing that it is more common to observe discrete (mostly binary) instruments such as random selection to receive treatment like in medical studies or random selection to receive a treatment conditional on covariates in social sciences ( e.g., Supplemental Nutrition Assistance Program, SNAP). We complement VP since we provide an alternative set of assumptions to identify the $MTE$, and also, we are focusing on a discrete instrument. Identifying marginal treatment effects with discrete instruments BMW show how a discrete instrument can be used to identify the marginal treatment effects under a functional structure that allows for treatment heterogeneity among individuals with the same observed characteristics and self-selection based on the unobserved gain from treatment. This paper builds upon BMW results by considering the case with (endogenous) misreporting and more flexible restrictions (such as shape restrictions instead of parametric assumptions) at the cost of losing point identification. The second one is MST which using the observed instrumental variables estimates, develops a linear programming approach to recover policy-relevant treatment effects such as the $MTE$. This paper differs from it by finding analytical bounds under the different smoothness and shape restrictions. Such bounds permit one to have a first-hand insight into how the assumptions are aiding identification. Estimation of the analytical bounds is simple since it can be performed using their respective sample analogs. Additionally, MST does not allow for the possibility of the treatment to be misreported while here is allowed. In the presence of misreporting, the results from MST do not apply directly while the ones derived here do. In the case of no misreporting, our bounds remain valid; in that sense, our results complement the ones from MST and BMW.
The rest of the paper is organized as follows, section (ref) introduces the main framework and assumptions. Section (ref) shows the main identification results and illustrates them. Section (ref) has an application of the identification results to KPGJ. Section (ref) concludes. Additional results are collected in the online appendix. Non analytical results on partial identification without additional shape restrictions extending MST are included the online appendix section (ref). Sections (ref) and (ref) of the appendix focuses on inference for the $ATE$. Section (ref) discusses how to choose the tuning parameter $b$. Section (ref) illustrates the bounds for the $ATE$. Section (ref) illustrates the analytical results on a $DGP$. Section (ref) extends the results with additional monotonicity assumptions. Sections (ref) and (ref) derives the results for the case when the instrument takes more than $2$ values.\footnote{The case of an instrument taking more than two values can also be seen as the generalization to the case of multiple discrete instruments. This is the case because multiple discrete instruments can be combined in one single multi-valued discrete instrument.} Finally, section (ref) collects all the figures from the document.
Consider the following framework (AK, HUV and HV):
Where $Y$ is an outcome variable that can be discrete, continuous, or mixed, the potential outcomes are denoted by $Y_d$, which is the outcome realization for when treatment $D=d$, $D=\{0,1\}$ is a binary unobserved endogenous treatment. Let $Z\in \mathcal Z=\{z_0,z_1,...,z_k\}$ be a discrete instrument,\footnote{The results will focus on the binary $z$ case but the generalization is natural for more than two values of $z$.} $V$ is a latent scalar random variable normalized to be uniformly distributed between $(0,1)$. $D^*$ is a misreported binary proxy of $D$, the true unobserved treatment status. $\varepsilon \in \{0,1\}$ is a random variable indicating the presence of misreporting or not. The vector $(Y,D^*,Z)$ is the observed data while $(Y_1, Y_0, D, \varepsilon, V)$ are latent (unobserved). In the rest of the document, small case letters denote realizations of the respective random variables. Object of interest: In this paper, we care about identifying the $MTE(v^*)$ which is the marginal treatment effect at a particular level $V=v^*$, more precisely, it is defined as $E[Y_1-Y_0|V=v^*]$. To identify the $MTE$ in this context, we introduce baseline assumptions that additional assumptions will aid. The baseline assumptions are:
The previous assumption and the model structure makes innocuous to say that $V$ is uniform between $[0,1]$ and that $p(Z)=P(D=1|Z)$.
Assumption (ref)-(ref) include the instrument independence and validity assumption as in HV, HUV among others. Assumption (ref) requires that $Z$ be a valid instrument, in the sense that it is statistically independent of the unobservables in the selection equation and the outcome equation. This assumption was used in AK. This assumption does not require the measurement error to be non-differential.\footnote{Non-differential measurement error is that conditional on the unobserved heterogeneity that drives the selection into treatment, misreporting is independent of the potential outcomes.} Non-differential measurement error combined with assumption (ref) implies that misreporting is independent of the outcome conditional on the true treatment, which is in general restrictive. Note that the measurement error can still depend on $z$, but this is through the true treatment since $D^*=D+(1-2D)\varepsilon$. Assumption (ref) is restricting the indicator of the existence of measurement error to be independent of $z$ but not the measurement error itself. Assumption (ref) requires the existence of an instrument that shifts the probability of selection into treatment. In addition, (ref) says that, even though the propensity scores cannot be recovered from the observed data (because $D$ is unobserved in practice), the ascending order of them in $Z$ is still known. This can be seen as a structural restriction imposed on the true treatment $D$. See TZ.\footnote {An example of sufficient condition is a constant-coefficient latent-index model. That is, suppose the treatment is generated by $D = 1(bZ > e)$, where $b$ is a parameter and $e$ is an error term independent of $Z$. Then, the order of $p(z)$, is determined by the sign of $b$. It is plausible in many applications that the sign of $b$ can be retrieved from economic theory. For example, in the study of the returns to schooling, distance to college is often used as an instrument for completed college education. In this specific example, the parameter $b$ is negative.}. Under assumption (ref), the sign of $p(z)-p(z\prime)$ for any two $z,z\prime$ is known. Under Assumption (ref),we are imposing that
With the false-positive probability $P(\varepsilon=1|D=0, Z=z)$ and the false-negative probability $P(\varepsilon=1|D=1, Z=z)$. We introduce the working example that will help interpret the assumptions and results through the rest of the document.
Besides the previously mentioned baseline assumptions, the following assumption is introduced.
KKKL introduces smoothness conditions for $E[Y_d|D=d]$ to bound the $ATE$ without an instrument and treatment exogeneity; this approach has the same spirit. In this case, we can build on their insight to provide bounds for the $MTE$ using similar smoothness conditions. The previous assumption states the degree of smoothness of the marginal treatment responses ($E[Y_d|V=v]$). Generally speaking, we may interpret our identification analysis in this section as a conditional one indexed by $b$. Furthermore, we may conduct a sensitivity analysis by looking at different values of $b$. The parameter $b$ is the Lipschitz constant which serves as a measure of smoothness. In this case we are assuming a maximum level of smoothness $b$. Assumption (ref) is restricting the functional form for the marginal responses, but considering all possible functionals in the lipschitz family with smoothness parameter $b$ or smaller instead of a particular parametric family (like for example linear functions). It is stating the degree of smoothness of the potential responses without assuming a particular functional form of it. In this sense $E[Y_d|V=v]$ could be for example linear $E[Y_d|V=v]=\mu_d+a_dv$ (in which case $b=a_d$) or quadratic $E[Y_d|V=v]=\mu_d+a_dv+c_dv^2$ (in which case $b=a_d+2c_d$) among different possibilities. This assumption introduces constraints in the underlying selection mechanism since is imposing restrictions on how the potential outcomes behave in relationship to the underlying cost of selecting into treatment. The smoothness assumption also relies on the choice of the Lipschitz constant which makes the result sensitive to the choice. This later point is discussed in the online appendix. Assumption (ref) might more appropriately be called something like bounded slope, bounded rate of change, or Lipschitz continuity of the $MTR$ functions but they are directly impacting the degree of parsimony of the functions, so we call it smoothness.\footnote{It is worth noticing that in the standard analysis on the $MTE$, the normalization of $V$ does not change any content of the model, but it does change the interpretation of the Lipschitz condition. This is because it is not the same to impose a Lipschitz condition on the conditional mean of $Y_d$ on $E$ where $E$ has a normal distribution, than to put it on $V=F_E(E)$ which is uniform.} The following remark adapted from KKKL is relevant to understand what this type of assumption is imposing on the marginal treatment responses.
More generally one could say $b^{\prime}|v_1-v_2|\leq E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b|v_1-v_2|$ as stated by KKKL, furthermore, letting $b^{\prime}=0$ and saying $E[Y_d|V=v_1]-E[Y_d|V=v_2]\leq b|v_1-v_2|$ for $v_1>v_2$ would be combining monotonicity of the treatment responses with assumption (ref). More precisely:
Note that following HUV and their standard assumptions ((ref)-(ref) above), without further restrictions the $MTE$ at the level of heterogeneity $v^*$ ($E[Y_1-Y_0|V=v^*]$) is not identified in this setting with discrete instruments and a misreported treatment. From standard results, we get:
The second equation takes the difference of the first equation for any two values of the instrument connecting the observed shift in $Y$ caused by changes in $z$ and the underlying treatment effect for all the individuals affected by such a change of the instrument. For any given $z$, say $z\prime$ we can get:
Where the first equality is because we are conditioning on $D^*=D, D^* \neq D$ and applying the properties of probabilities. This last equation reflects that the observed propensity score for the proxy of the true treatment variable conditional on $z\prime$ equals the share of treated individuals who at that particular $z\prime$ report treatment status correctly multiplied by the probability of reporting correctly, plus the share of not treated individuals who at that particular $z\prime$ report treatment status incorrectly multiplied by the probability of reporting incorrectly. The previous expressions depend on unobserved components. While $ E[Y|Z=z\prime], P[D^*=1|Z=z\prime]$ are observed, $P[D^*=D|Z=z\prime],P[D=1|D^*=D,Z=z\prime]$ and $p(Z)$ are not, which without further assumptions do not allow for identification of the true propensity score and also of the $MTE$. If $p(z)$ was observed and $z$ was continuous, then the $MTE(v^*)$ would be identified as $\frac{\partial E[Y|P(Z)=v^*] }{\partial v^*}=\frac{\int_0^{v^*}E[Y_1|V=v]dv + \int_{v^*}^1E[Y_0|V=v]}{\partial v^*}=E[Y_1|V=v^*]-E[Y_0|V=v^*]$. So this displays the two main identification challenges, the non-continuity of $z$ and the fact that $p(z)$ is not observed.
Before proceeding to the identification results, it is worth showing the main elements of the current work and how they differentiate from previous identification results of $MTE$ with discrete instruments. It is also relevant to show the role of misreporting. From the observed data if there is no misreporting from assumptions (ref)-(ref) one can identify:
The first equality comes from the definition of the model, the second one from the laws of probability, the third one by the independence of $Z$ from $Y_1, V$ and the last one from the properties of conditional expectations and the normalization that $V$ is marginally uniform. Similarly, we have:
This then implies the equality expressed at the beginning of this subsection:
Without misreporting the propensity score $p(z)$ is identified and, given assumptions (ref)-(ref) index sufficiency holds and thus $E[YD|Z=z]=E[YD|p(Z)=p]$. So we can rewrite the previous equalities as functions of $p$ instead of $z$. Where $p \equiv p(z)$. In this context without differentiability of $p$ the key insight from BMW is to introduce parametric restrictions that for example say that $E[Y_d|V=v]=\mu_d +a_d v$ and thus $E[Y_1-Y_0|V=v]=\mu+a v$, where $\mu\equiv \mu_1-\mu_0$ and $a \equiv a_1-a_0$. Additionally define $c=\mu_0+\frac{a_0}{2}$. In this case
Then from $E[YD|p(Z)=p]$ for different values of $p$ (at least two which is enough with a binary instrument) we can solve for $\mu_1,a_1$. Similarly for $a_0,\mu_0$ from $E[Y(1-D)|p(Z)=p]$. Note that then given the marginal treatment responses, $E[Y_d|V=v]$ is identified, so it is the $MTE$ as their difference. One might not be willing to assume particular parametric specifications for the conditional expectations of the potential outcomes since they are restrictive. One of the contributions of the current work is relaxing such restrictions and still recovering analytically tractable expression for the bounds of the $MTE(v^*)$. The current work relates MST in the following way. MST relies on recovering the set of marginal treatment responses consistent with observed $IV$-like estimands. In this setting, such strategy would rely on finding all the candidates $E[Y_d|V=v]$ functions consistent with:
Their strategy relies on the fact that $w_1 w_0$ are known (or identified). In the case of misreporting, where we do not know exactly the rate of false positives and false negatives for every value of $z$, we have that $p(z)$, and thus, the weights are not identified. This makes the current work to differ from the existing literature since developed computational methods rely on the weights being known or identified. In this context one of the main contributions of the current paper is working in the context where the weights are not identified but actually can be partially identified. In section (ref) we start from the same insight as MST, but instead of solving a linear problem, we aid identification with shape restrictions to get analytical bounds on the $MTE$. In the online appendix an extension of MST is discussed without aiding identification with shape restrictions by solving the same linear problem as in MST but for different values of $p(z)$ score in the identified set. As the identified set of $p(z)$ is not finite, the solution can only be approximated.
In subsection (ref) identification without misreporting will be discussed. Subsection (ref) incorporates misreporting.
Note than since
For any values $z$ and $z\prime$ with $p(z) \geq p(z\prime)$:
Where we are adding and subtracting the marginal treatment responses ($E[Y_d|V=v^*]$) related to the marginal treatment effect at the $v^*$ of interest ($E[Y_1|V=v^*]-E[Y_0|V=v^*]$), then using $E[Y_d|V=v]-E[Y_d|V=v^*]\leq b|v-v^*|$ twice and also the definition of $MTE$. Similarly we can get:
Then:
The bounds depend on the propensity score and the difference of the propensity score for different values of $z$.
In order to identify the $MTE$ first we need to identify $p(z)$. In AK such identification is discussed. Subsequent subsections, builds upon the results from AK. For clarity, let $\Delta_p \equiv p(z)-p(z'), \Delta_{D^*Z}(z',z) \equiv P(D^*=1|Z=z)-P(D^*=1|Z=z'), \alpha \equiv P(\varepsilon=1) $, then let the bounds derived in AK be:
The bounds on the propensity score rely on any given level of unconditional misreporting ($\alpha$). Misreporting enters in two ways. First of all, if the probabilities of misreporting are unknown, $p(z)$ is no longer identified (see AK for more details), which then creates a problem since such quantity appears systematically in the bounding strategies. To solve this problem, we use the previously defined bounds on the misreporting probabilities. The lack of point identification of $p(z)$ affects the strategy using smoothness restrictions. See for example that:
The bound of $MTE$ depends on the sign of it since it is not always true that $MTE(v^*)(p(z)-p(z\prime)) \leq MTE(v^*)(p_u(z)-p_l(z\prime))$. These considerations are taken into account in theorem (ref). Smoothness assumptions for marginal treatment responses
The bounds from theorem (ref) builds on the following inequalities due to assumption (ref)
Where the inequalities use the smoothness assumption as in the previous section without misreporting and fact that $\int_{p(z\prime)}^{p(z)}2b|v-v^*|dv \leq \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv$. Note that we have a sufficient condition for identifying the sign of the $MTE(v^*)$. If $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$ then the $MTE(v^*)$ is positive. If $E[Y|Z=z]-E[Y|Z=z\prime]+ \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\leq 0$ then the $MTE(v^*)$ is negative. If $E[Y|Z=z]-E[Y|Z=z\prime]- \int_{p_l(z\prime)}^{p_u(z)}2b|v-v^*|dv\geq 0$ then from equations (ref) and (ref) combined with the bounds on $p(z)-p(z\prime)$ we get:
Note that the form of the bounds depend on $|v-v^*|$. In the integral of the absolute value is where the relative position of $v^*$ with respect of $p_l(z\prime),p_u(z)$ will matter. Note that if the $v^*$ of interest is such that $v^*\leq p_l(z\prime)$, the integral involving $|v-v^*|$ is $ [\frac{p_u(z)^2-p_l(z\prime)^2}{2}+v^*(p_l(z\prime)-p_u(z))]$. If the $v^*$ of interest is such that $v^*\geq p_u(z)$ then $[\frac{-p_u(z)^2+p_l(z\prime)^2}{2}+v^*(-p_l(z\prime)+p_u(z))]$. If the $v^*$ of interest is such that $p_l(z\prime)\leq v^*\leq p_u(z)$ then $[v^{*2}-v^*(p_u(z)+p_l(z\prime))+(\frac{p_u(z)^2+p_l(z\prime)^2}{2})]$
The following theorem summarizes the previous discussion.
The previous theorem is extended for discrete instruments taking more than 2 values in the online appendix. A researcher might be interested in combining the monotonicity assumption on the treatment responses and the smoothness assumption. So instead of using assumption (ref), the researcher might be willing to use (ref). This result is collected in the online appendix. The choice of $b$ is not arbitrary. The fact that it operates as a tuning parameter might lead to a discretionary use of it to get the desired result; different ways of choosing this parameter are discussed in the online appendix. The previous identification results are illustrated in the appendix.
In this section, the developed methods are applied to get bounds on the $MTE$. We then integrate them over $v^*$ to get bounds on the $ATE$ of receiving SNAP on the outcome of being food insecure. As stated by KPGJ SNAP, formerly known as the Food Stamp Program, is by far the largest food assistance program in the United States and, as such, constitutes a crucial component of the social safety net in the United States. In any given month during 2009, SNAP assisted more than 15 million children, and it is estimated that nearly one in two American children will receive assistance during their childhood. Concluding about the program's impact is complex due to two of the fundamental problems studied in this paper. First, a selection problem arises because the decision to participate in SNAP is unlikely to be exogenous. On the contrary, unobserved factors such as expected future health status, parents’ human capital characteristics, financial stability, and attitudes towards work and family are all thought to be jointly related to participation in the program and health outcomes such as food security. Families may decide to participate precisely because they expect to be food insecure or in poor health. Second, a nonrandom measurement error problem arises because a large fraction of food stamp recipients fails to correctly report their program participation in household surveys. Using administrative data matched with data from the Survey of Income and Program Participation (SIPP), for example, BD find that errors in self-reported receipt of food stamps exceed 12 percent and are related to respondents’ characteristics, including their true participation status, health outcomes, and demographic attributes. MMS provide evidence of extensive underreporting of food stamps in the SIPP, the Current Population Survey (CPS), and the Panel Study of Income Dynamics (PSID). In this context KPGJ studies the average effects using the December Supplement of the 2003 Current Population Survey (CPS). In the data, we can observe a self-reported measure of food stamp receipt over the past year, food insecurity over the past year, and the ratio of income to the poverty line.\footnote{For further details about the data see GK. In there, they state that just over 40 percent of the households report receiving food stamps, and the food insecurity rate among self-reported recipients is 17.9 percentage points higher than among eligible non-recipients (52.3 percent vs. 34.4).} As KPGJ states, the data is rich enough to allow the construction of instrumental variables for SNAP participation used in previous literature. In particular, state identifiers in the CPS apply a more traditional instrumental variable (IV) assumption based on cross-state variation in program eligibility rules.\footnote{In general terms, program eligibility rules are income requirements (most households must meet both gross and net income limits to qualify for SNAP benefits), resource requirements (households must also meet a resource limit in their bank accounts), work requirements (If you are an able-bodied adult without dependents, between the ages of 18 and 49, and able to work but currently unemployed, you may only be eligible for SNAP benefits for three months within a three-year period) and other eligibility requirements (to be eligible for SNAP benefits, households must also, meet other conditions in addition to the income and resource requirements, such as everyone in your household having, or have applied for, a social security number). To establish the income and resource requirements, each state computes asset tests, but there is variation in how these states evaluate the assets of individuals. More specifically, they may or may not include certain assets that will affect the individuals' eligibility.} Merging the Urban Institute’s database of state program rules with the CPS data KPGJ create two instrumental variables: an indicator for whether the state uses a simplified semi-annual reporting requirement for earnings and an indicator for whether cars are exempt from the asset test.\footnote{For more details on the construction see KPGJ.} These instrumental variables if valid and independent allows using the current methods to bound the $MTE$. For example, if one is willing to assume that the state variation in the asset test is exogenous, then instrument independence is satisfied. It is worth noticing that only $LATE$ can be identified from an instrumental variable regression under individual heterogeneity. In this sense, the methods developed here can be used to recover bounds on the average treatment effect, and complement KPGJ results. In this context, the $MTE$ at a particular level of $v^*$ represents the treatment effect receiving SNAP has for a particular level of stigma. Stigma mat is connected with the potential outcomes since, as stated by Food2 stigma manifestations affect health outcomes (and food security as such). More precisely, in this context to illustrate our methods, $Y$ is a food insecurity indicator over the past year, $D^*$ is the self-reported SNAP participation (subject to potential measurement error as stated before), $Z$ is a binary variable for cars exempt from the asset test. Concerning choosing $b$ for the sake of exposition, we present the results here for several potential values of $b$. In the online appendix we discuss different methods on how to choose $b$ which could be used in this context. Intuitively choosing $b$ is restricting the degree of smoothness the $MTE$ would have. One can draw a parallelism between choosing a linear form (a low $b)$ for the $MTE$ versus choosing a high order polynomial (high $b$). The linear form is rather restrictive on the behavior of higher-order derivatives (and thus smoothness) compared with the polynomial. A researcher choosing $b$ for SNAP should consider what he thinks the underlying decision-maker optimization problem looks like. If the expected utility, for example, has a quadratic form, or the optimal expected demand of food security is linear under both receiving and not receiving SNAP, then the expected benefit on the optimal choice of food security both under getting SNAP and not getting it could be considered linear. Consistently with KPGJ a treatment cannot hurt assumption ($P(Y_1<Y_0|V=v)=1$) will be introduced. Making then the conservative upper bounds of the $MTE$ to be $0$. The following table summarizes the data and shows the average and median characteristics in the sample.
We can see that around forty-two percent of the people are food insecure, while forty-one report receiving SNAP. Thirty percent have their cars excluded from the asset test, which implies that thirty percent of the individuals in the sample live in states where cars are excluded from the asset test. On average (and in the median), individuals in this sample are below the poverty line. Each of the graphs in the following figure is computed in the following way. For any given level of $b$ and $\alpha$, point estimates of the bounds are constructed for the $MTE$ at different values $v^*$ using their sample analogs. To decide which type of bound to use, the relative position of the sample analog estimates of the bounds for the propensity score is calculated. The maximum level of $\alpha$ is twenty percent which is chosen as an arbitrary big upper bound above the existing results from KPGJ. See Figure (ref).
In orange, we have the upper and lower bounds assuming $\alpha=0.2$. In blue, we have the upper and lower bounds assuming $\alpha=0.1$, and in red, we report the upper and lower bounds assuming $\alpha=0$ (no misreporting). The limits of the $Y$-axis are the worst-case upper bound ($0$ under treatment cannot hurt) and the worst-case lower bound ($-1$). The $X$-axis goes from $0$ to $1$, the different potential values of $v^*$. We can see that the bounds have identification power over different regions of $v^*$ support. We can see that when the level of misreporting decreases, the bounds become tighter since there is less lack of identification due to misreporting. Similarly, when $b$ decreases, we also get tighter bounds consistent with reducing the potential functional forms of the marginal treatment responses. These bounds are computed by fixing a level of $\alpha$. If, for example, in the case of $b=0.1$, the researcher is not interested in $\alpha=0.2$ rather in $\alpha \leq 0.2$ then all the region between the upper orange curve and the lower orange curve is the identified set consistent with $b=0.1$ and all the $\alpha$'s less or equal to twenty percent. The length of the identified set becomes tighter in the region between the estimated observed propensity scores ($\widehat{P}[D^*=1|Z=1]=0.49,\widehat{P}[D^*=1|Z=0]=0.38$) since there is more information being used to bound the $MTE$. The previous display also implies a simple way of computing bounds of a parameter of interest such as the $ATE$. It is known that $ATE=\int_{0}^1MTE(v^*)dv^*$. Then we know that $\int_{0}^1MTE_{lb}(v^*)dv^* \leq ATE \leq \int_{0}^1MTE_{ub}(v^*)dv^*$. So then we can approximate bounds for the $ATE$ as:
Where $N_{v^*}$ is the number of grid points where $MTE(v^*)$ was evaluated and where $\widehat{MTE}$ are the bounds from the previous graphs. On the online appendix estimates of the bounds on the $ATE$ are computed. In the online appendix, there is also an alternative way of estimating (and doing inference) on the $ATE$ for the case of smoothness restrictions. On the online appendix a method for asymptotic normality for an outer-set of the $ATE$ in the case of no misreporting and with smoothness conditions is developed. On the online appendix a method for asymptotic normality for an outer-set of the $ATE$ in the case of misreporting, treatment cannot hurt assumption, and smoothness conditions can be found. The $ATE$ bounds are easily estimated and used for inference, since as shown in the appendix, each component can be replaced by their sample analogs, which themselves are asymptotically normal, and thus, by the continuous mapping theorem the upper bounds and lower bounds for the $ATE$ are also. Then, an asymptotically valid bootstrap procedure can be used to build confidence intervals for the entire identified set, such as those constructed by MN. The population identification region is an interval $[L, U]$, we can estimate each side of the interval with consistent and asymptotically normal estimators $\widehat{L}, \widehat{U}$ and via this procedure we get confidence interval $[\widehat{L}-z_{\frac{\alpha+1}{2}}\widehat{\sigma}_l/\sqrt{N}, \widehat{U}+z_{\frac{\alpha+1}{2}}\widehat{\sigma}_u/\sqrt{N}] \equiv [\widehat{L}_{\alpha}, \widehat{U}_{\alpha}]$ such that
In order to control for covariates $X$ such that independence of $Z$ holds conditional on it in a tractable manner, we can assume $E[Y_d|V=v,X=x]=m_d(v)+\beta X$. Then,
Where $P(z,x)=P(D=1|X=x,Z=z)$. We can then any two values of $Z$:
We can then follow a similar display is in Section (ref). In this case, we can estimate $E[Y|Z=z,X=x]$ with a partial linear regression while $p(z,x)$ is estimated with a non-parametric regression. So far, we have reported results for the instrument taking only two values. The data set used in this problem counts with two potential instruments. The already used one, and an eligibility criterion specifying if earners report twice a year or not. Based on this, a three-valued instrument can be constructed related to the intensity of the likelihood of receiving SNAP. That is, it takes $0$ if both instruments take the value $0$, takes the value $1$ if either of them takes the value $1$, and it takes the value $2$ if both of them take the value $1$. The details of how the bounds look for more discrete non-binary instruments and the particular case of an instrument taking three values are collected in the online appendix. In such a case, the length of the identified set for the $ATE$ becomes smaller; this is intuitive since we now have a more exogenous variation to exploit. We we illustrate it with the case of $b=0.5, \alpha=0.1$. Additionally, we can see that the form of the identified set of the $MTE$ changes, since now the regions rely on the different exogenous variation in zones where previously, only the shape restrictions could be used. These results are collected in the online appendix.
In this paper, we provided partial identification results for the Marginal Treatment Effect in the presence of measurement error and a discrete instrument building over MST, BMW and AK. To do so, given the discrete nature of the instruments, we introduced smoothness restrictions. Results are illustrated via a numerical example and quantifying the marginal treatment effect of SNAP on food insecurity, a case in which measurement error and endogeneity of treatment are known to be an issue. In a more general way, our results can serve as a sensitivity analysis tool for when researchers are interested in recovering the $MTE$ in the presence of a discrete instrument and suspect measurement error, and have doubts about their parametric assumptions. This sensitivity analysis is executed by varying $\alpha, b$. If no measurement error exists, the results from this paper provided analytical partial identification results of the $MTE$ in the presence of discrete instruments that can serve as a complement for the already existing results. \singlespace
\setcounter{table}{0} \setcounter{figure}{0} \setcounter{equation}{0}