Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,184 characters · 11 sections · 91 citation commands
Inference on Local Average Treatment Effects for Misclassified Treatment
The local average treatment effect (LATE) is a popular causal parameter in the microeconometric literature (e.g., angrist2008mostly). It represents the average causal effect of a binary endogenous treatment $T^*$ on an outcome $Y$ for the unit whose treatment status changes depending on the value of a binary instrument $Z$. imbens1994identification show that the instrumental variable (IV) estimator may identify the LATE. While the identification requires the precise measurement of the {\it true} binary treatment $T^*$ in addition to identification conditions, in practice, the observed binary treatment $T$ may be mismeasured for $T^*$. LATE applications may involve such a degree of misclassification that the actually treated unit can even be misrecorded as untreated and vice versa. While we are not aware of the presence of measurement errors in practice, ignoring such errors may lead to drawing misleading empirical implications.
As an application, let us consider the causal analysis of returns to schooling, one of the most important applications of LATE inferences. In this case, $Y$ is an individual outcome such as wages, $T^*$ is an indicator of educational attainment such as a college degree, and $Z$ is a variable indicating whether the individual lives close to a college (e.g., card1993using). Educational attainment may be mismeasured because of a random recording error or the provision of intentionally/unintentionally false statements. Indeed, several studies have pointed out the prevalence of misreported schooling (kane1995labor; card1999causal; black2003measurement; battistin2014misreported).
The present study contributes to the econometric literature by proposing a novel inference procedure for the LATE with the mismeasured treatment. The measurement error brings out a bias such that the conventional IV estimator under- or overestimates the LATE. While the IV estimation solves the problem of the {\it classical} measurement error that is uncorrelated with the true variable (e.g., wooldridge2010econometric), the bias for the LATE is caused in our situation since the measurement error for the binary variable must be {\it non-classical}. That is, the error is correlated with the true unobserved treatment, because the support of the measurement error for the discrete variable depends on the true variable (e.g., aigner1973regression). The bias for the LATE depends on the distribution of the measurement error, meaning a failure to identify the LATE.
To correct the identification bias, we develop point-identification for the LATE and the distribution of the measurement error based on the availability of an exogenous observable variable, say $V$. The exogenous variable $V$ can be a covariate, instrument, or repeated measure of the treatment as long as $V$ satisfies our identification conditions. Intuitively, Figure (ref) illustrates the relationship between such exogenous variables and other variables. Many previous studies have corrected problems due to measurement errors by using such exogenous variables (e.g., hausman1991identification; lewbel1997constructing,lewbel1998semiparametric; schennach2007instrumental; hu2008identification; hu2008instrumental), and our identification builds on this strand of the literature. In particular, we extend the results by mahajan2006identification and lewbel2007estimation on the mismeasured exogenous treatment to the LATE inference with the mismeasured endogenous treatment.
The main idea behind our identification is, under a set of empirically plausible conditions, to derive moment conditions no fewer than the parameters including the LATE, true first-stage regression, and distribution of the measurement error. We introduce three key conditions when $V$ is non-binary, and we require an additional condition when $V$ is binary. Firstly, the measurement error of the treatment is {\it non-differential} in the sense that mismeasured $T$ does not affect the mean outcome once true $T^*$ is conditioned on. Secondly, $V$ has to satisfy a relevance condition that requires that $V$ relates to $T^*$. Thirdly, $V$ has to satisfy an exclusion restriction under which $V$ cannot affect the difference in the conditional means of $Y$ and the distribution of the measurement error of $T$. Finally, when $V$ is binary, we need the additional condition under which the distribution of the measurement error does not depend on $Z$. These conditions may be satisfied when $V$ is a covariate, instrument, or repeated measure of the treatment. Importantly, our exclusion restriction does not rule out the possibility that $V$ directly affects $Y$ and does hold if $V$ affects the outcomes with and without the true treatment equally.
The moment conditions derived from the identification lead to moment-based estimators for the parameters. We propose adopting hansen1982large's (hansen1982large) generalized method of moments (GMM) estimator because of its popularity in the literature. Desirably, the GMM inference is easy to implement in practice, and its asymptotic properties are well understood. The asymptotically valid inferences can be developed in the usual manner.\footnote{An { R} package to implement the proposed GMM inference is available on the author's website.}
We illustrate the usefulness of the proposed GMM inference based on Monte Carlo simulations and an empirical illustration. The simulations show that our GMM inference works well in finite samples. The empirical illustration estimates the returns to schooling and the distribution of misreported schooling using card1993using's (card1993using) data by utilizing two indicators of college proximity as exogenous variations $Z$ and $V$.\footnote{While our empirical illustration utilizes two indicators of college proximity, the literature on the returns to schooling suggests many other potential exogenous variables such as the quarters of birth (angrist1991does), parental/sibling education (altonji1996effects), and the sex of siblings (butcher1994effect). See, for example, card1999causal, card2001estimating for a survey on returns to schooling analyses based on such exogenous variables. Such exogenous variables are candidates for exogenous variables to overcome endogeneity and measurement errors simultaneously in our proposed inference. } Our relevance condition requires that whether living close to colleges relates to schooling, and we demonstrate its validity. Our exclusion restriction allows college proximity to relate to the factors affecting wages and seems more plausible than the exclusion restriction for the standard IV inference. We find that the conventional inference overestimates the LATE because of misreported education and our GMM inference corrects the problem successfully.
\paragraph{Paper organization.} Section (ref) reviews related studies. Section (ref) introduces the setup, reviews the LATE inference, and explains the problem due to the measurement error. Section (ref) develops the identification. Sections (ref) and (ref) present the simulations and empirical illustration, respectively. Section (ref) concludes. Appendix (ref) gives the proofs of the theorems. The supplementary appendix contains auxiliary results.
The present study builds on the literature on misclassification problems (i.e., problems due to measurement errors for discrete variables) on which many studies contribute. For example, aigner1973regression, bollinger1996bounding, hausman1998misclassification, frazis2003estimating, molinari2008partial, hu2008identification, imai2010causal, and ura2016instrumental consider linear and non-linear models with possibly misclassified discrete variables. Since misclassification causes a non-classical measurement error that is correlated with the true variable, it fails to identify the parameter of interest. Since the identification bias due to misclassification depends on the misclassification probabilities (i.e., the probabilities that the truly treated is misclassified as the untreated and vice versa), the literature has developed how to infer the misclassification probabilities.
In particular, our identification of the misclassification probabilities relates to the results presented by mahajan2006identification and lewbel2007estimation among the studies in the misclassification literature. Both works identify and estimate the conditional mean of the outcome $Y$ given the binary exogenous treatment $T^*$ and its difference $E(Y|T^*=1) - E(Y|T^*=0)$ based on the availability of the misclassified binary treatment $T$ and an exogenous variable $V$.\footnote{While mahajan2006identification and lewbel2007estimation allow for the presence of other control variables, we omit such variables here for perspicuity. } However, their identification conditions have important distinctions for the relationship between $Y$ and $V$ and for the number of the elements that $V$ takes. On the one hand, mahajan2006identification requires that $E(Y|T^*, V) = E(Y|T^*)$ so that $V$ cannot be a covariate in the sense that $V$ cannot directly affect $Y$, although he allows $V$ to be binary. On the other hand, lewbel2007estimation allows $V$ to be a covariate by introducing an exclusion restriction under which $E(Y|T^*=1, V) - E(Y|T^* = 0, V)$ does not depend on $V$, but he requires that $V$ has to take at least three values, meaning that $V$ cannot be binary.
Our identification extends the results presented by mahajan2006identification and lewbel2007estimation by allowing $V$ to be binary and a covariate simultaneously. These features of $V$ are empirically important since we are rarely confident about whether an exogenous variable affects the outcome and often observe only a binary exogenous variable. Our identification allows for both these features by utilizing the condition that the instrument $Z$ is not related to the misclassification probabilities, instead of excluding the direct effect of $V$ on $Y$ and binariness of $V$. This condition for the misclassification probabilities could be empirically plausible since $Z$ should be assigned as good as randomly in the LATE inference (see angrist2008mostly). Indeed, some of the studies of LATE inferences with a mismeasured treatment discussed below utilize the same condition to identify a parameter similar to the LATE or to partially identify the LATE. We also stress that we cannot use the estimation procedures of mahajan2006identification and lewbel2007estimation to infer the LATE since they intend to infer the conditional mean of $Y$ given $T^*$ and its difference.
mahajan2006identification also discusses identification when the true binary treatment is endogenous and misclassified as an extension of his main result, although the recent study by ditraglia2015mis points out that mahajan2006identification's (mahajan2006identification) result fails. The intuitive reason for this failure is that it is impossible for his approach to correct both endogeneity and misclassification by using solely one instrument. See ditraglia2015mis for details on this problem. We note that such a problem does not occur in our proposed inference since we utilize two exogenous variations, $Z$ and $V$.
To the best of our knowledge, LATE inferences with the binary, endogenous, and misclassified treatment have been examined by battistin2014misreported, calvi2017women, ditraglia2015mis, and ura2017hetero.
battistin2014misreported use two repeated measures for the treatment in addition to the observed treatment $T$ to develop point-identification and a semiparametric estimation for the LATE in a mixture model. Their necessity is the presence of multiple repeated measures for the true treatment from resurvey data. Unlike their approach, our proposed inference may be applicable even when one may utilize a covariate.
calvi2017women utilize a repeated measure of the treatment in addition to the observed treatment $T$ to identify a causal parameter similar to the standard LATE, which they call the mismeasurement robust LATE (MR-LATE). While the MR-LATE can be identified under less restrictive assumptions, it is not identical to the standard LATE in general. We stress that the MR-LATE is different from the standard LATE under the identification conditions developed in the present study.
ditraglia2015mis develop partial- and point-identification for the LATE in a non-parametric additively separable model with moment restrictions on the measurement error and an additively separable error. They do not require the availability of an exogenous variable and/or resurvey data, but demand that the individual treatment effect does not depend on the unobservables. On the contrary, we allow arbitrary heterogeneous treatment effects, although we still need an exogenous variable.
ura2017hetero proposes sharp partial-identification for the LATE with the mismeasured treatment in a general setting. His approach allows for an endogenous measurement error and does not require resurvey data and/or the use of an exogenous variable, but his result does not attain point-identification for the LATE.
Many econometric studies examine non-classical measurement errors and/or non-linear errors-in-variables models for other settings, and our proposed inference also builds on this body of the literature. For example, amemiya1985instrumental, hsiao1989consistent, horowitz1995identification, hu2008instrumental, hu2015closed, and song2015estimating study such measurement error models for microeconometric applications. See bound2001measurement, chen2011nonlinear, and schennach2016recent for excellent reviews of the literature in this regard.
The present study also builds on recent research using covariates to solve endogeneity or measurement errors. In particular, our exclusion restriction and relevance condition relate to conditional covariance restrictions given covariates in caetano2017identifying. We compare their restrictions with ours in Remark (ref), and we here stress that both papers have different aims. They focus on estimating marginal effects in general IV models without measurement errors, while our aim is to infer the LATE with misclassification. Building on the important result by caetano2017identifying, for example, ben2017identification also utilize analogous restrictions for covariates to solve measurement errors. However, their setting is also different from ours.
This section explains the setting considered in this study. Section (ref) introduces the econometric model and briefly reviews the LATE inference without a measurement error. Section (ref) discusses the problem due to a measurement error for the treatment.
Let $(\Omega, \mathcal{F}, P)$ be the common probability space where the random variables introduced below are defined. We have a random sample of $(Y, T, Z, V)$, where $Y \in \mathbb{R}$ is an outcome, $T \in\{0,1\}$ is a possibly mismeasured treatment, $Z \in\{0,1\}$ is an IV, and $V \in supp(V) \subset \mathbb{R}$ is an exogenous variable such as a covariate, instrument, or repeated measure of the treatment. The true unobservable treatment $T^* \in \{0,1\}$ may be endogenous because of omitted variables in the sense that some unobservables relate to both $Y$ and $T^*$. The observed $T$ may contain a measurement error, meaning that $T \neq T^*$ in general.
The availability of an exogenous variable $V$ is essential for our procedure. While the standard LATE inference does not require such exogenous variables, our analysis needs $V$ to correct the misclassification problem. There are many empirical situations in which we can utilize such exogenous variables. For example, when survey data with rich information are available, we may utilize covariates and/or instruments affecting $T^*$ and/or $Y$ as $V$. Further, when we access survey and resurvey data, we may observe a repeated binary measure $V$ for true $T^*$ in addition to $T$. Note that in this case $V$ might also contain a measurement error for $T^*$ and $V$ is not identical to $T$ in general. Section (ref) discusses that $V$ has to satisfy some conditions such as exclusion restrictions and a relevance condition to identify the LATE and the distribution of the measurement error.
For example, in the returns to schooling analysis, $Y$ is wages, $T^*$ is a college degree, $Z$ is the IV indicating whether the individual lives close to a four-year college, and $V$ is a covariate or instrument (e.g., two-year college proximity or the quarter of birth), or a repeated measure of the degree.
The aim of the inference is to examine the causal relationship between $T^*$ and $Y$. Let $Y_0$ and $Y_1$ be the potential outcomes when the unit is truly untreated ($T^*=0$) and when it is treated ($T^*=1$), respectively. Similarly, let $T^*_0$ and $T^*_1$ be the potential true treatment statuses when $Z=0$ and $Z=1$, respectively. We can write $Y= Y_0 + T^* (Y_1-Y_0)$ and $T^* =T_0^* + Z (T_1^* - T_0^*)$. The individual causal effect of $T^*$ on $Y$ is $Y_1 - Y_0$, which may be heterogeneous across units depending on the observables and/or unobservables.
We define the following subsets of the common probability space $(\Omega, \mathcal{F}, P)$.
Intuitively, $A$ or $N$ is the set of units that always take or deny, respectively, the treatment (in the sense of true $T^*$), $C$ is that of units whose treatment statuses are positively affected by $Z$, and $D$ is that of units whose treatment statuses are negatively affected by $Z$.
The parameter of interest is the LATE, which is defined as
Without the measurement error, imbens1994identification show that the LATE is identified by the IV estimand based on $(Y,T^*,Z)$:
where $\beta^* = \Delta \mu / \Delta p^*$ is the IV estimand with $\Delta \mu \coloneqq \mu_1 - \mu_0$, $\Delta p^* \coloneqq p_1^* - p_0^*$, $\mu_z \coloneqq E(Y|Z=z)$, and $p_z^* \coloneqq E(T^*|Z=z) = \Pr(T^*=1|Z=z)$. Note that, if we observe $(Y,T^*,Z)$, we can consistently estimate $\beta^*$ by the two-stage least squares estimation.
Formally, we need the following conditions to establish the identification in (ref). These are essentially the same as the conditions in imbens1994identification.
Assumption (ref) is an exclusion restriction that requires the instrument to be unrelated to the factors affecting the outcome and/or treatment. Intuitively, the assumption guarantees that the instrument is assigned as good as randomly. Assumption (ref) requires the presence of compliers. This is a relevance condition requiring a positive effect of $Z$ on $T^*$. Assumption (ref) rules out the presence of defiers, which is known as a monotonicity condition in the LATE literature.
In the example of returns to schooling with four-year college proximity $Z$, the LATE is the average returns of college degrees for individuals who graduate if and only if they live close to four-year colleges. Assumption (ref) implies that college proximity is unrelated to the factors affecting wages with and without a college degree and potential educational attainment with and without college proximity. Assumption (ref) requires the presence of individuals who graduate if and only if those individuals live close to colleges. Under Assumption (ref), no individuals graduate when they do not live close to a college, but they do not graduate when they do live close to a college.
Based on the identification result in (ref), we call $\beta^*$ “the LATE” below for convenience of explanation, meaning that we implicitly assume that Assumptions (ref), (ref), and (ref) hold throughout the paper.
This section explores the identification problem of the LATE $\beta^*$ in the situation where observed $T$ may be a mismeasured variable of true $T^*$.
The identification result in (ref) implicitly requires the precise measurement of true $T^*$ in addition to the identification conditions of Assumptions (ref), (ref), and (ref). However, observed $T$ may contain a measurement error in practice, of which there are a few types. For example, in the returns to schooling analysis, educational attainment may be misrecorded randomly during the process of correcting the survey data. Such a measurement error is independent of the factors affecting the outcome. Our proposed procedure allows this type of measurement error. The other possibility of measurement errors is false reporting. Individuals may misunderstand the question or have poor recall about their educational attainment when responding to a survey. Individuals might also have an incentive to make false statements about their academic achievement to enhance their careers. Such measurement errors may be correlated with the observables and/or unobservables. Our analysis can allow for false reporting depending on the observables, and we can explicitly allow the measurement error to depend on the observables (see Remark (ref) below).
To examine the misclassification problem, we define the misclassification probability:
In words, $m_0$ is the probability that an individual who is actually untreated is misclassified as treated, and $m_1$ is analogous. The misclassification probability can also be regarded as the distribution of the measurement error.
The measurement error for the treatment is non-classical in the sense that it is dependent on the true treatment. This is because the support of the measurement error for a discrete variable depends on the true variable. To see this, we denote the measurement error for $T$ as $U_T \coloneqq T - T^*$. By construction, $U_T \in \{0,1\}$ given $T^*=0$ but $U_T \in \{-1,0\}$ given $T^*=1$, meaning that $U_T$ is dependent on $T^*$. Specifically, it is easy to see that $Cov(U_T,T^*) = - (m_0 + m_1) \Pr(T^*=0) \Pr(T^*=1)$, implying that the correlation between the error and true treatment is always negative, whereas its magnitude depends on the misclassification probabilities. We can also show that the correlation between $T^*$ and $T$ depends on the misclassification probabilities:
If the sum of the misclassification probabilities $m_0+m_1$ is less (greater) than one, $T$ is positively (negatively) correlated with $T^*$.
To explore the identification problem for the LATE $\beta^* = \Delta \mu/ \Delta p^*$, we examine the relationship between observable treatment probability $p_z \coloneqq E(T|Z=z) = \Pr(T=1|Z=z)$ and true treatment probability $p_z^* = E(T^*|Z=z) = \Pr(T^*=1|Z=z)$. We have the following relationship according to the law of iterated expectations:
where $m_{tz} \coloneqq \Pr(T \neq T^*|T^*=t,Z=z) = \Pr(T=1-t|T^*=t,Z=z)$ is the conditional misclassification probability and $s_z \coloneqq 1- m_{0z} - m_{1z}$ for $t,z=0,1$. We note that $|s_z| \le 1$ by definition. The above equation implies that
The true treatment probability thus depends on the conditional misclassification probabilities.
The misclassification of the treatment variable brings out the serious problem that $\beta^*$ and $\Delta p^*$ cannot be point-identified based on the observables of $(Y,T,Z)$. Equation (ref) leads to the following system:
Even when $Z$ is not related to the misclassification probability, namely $m_t = m_{tz}$ for $t,z=0,1$ (this implies $s = s_0=s_1$ for $s \coloneqq 1-m_0-m_1$), we cannot identify $p_0^*$ and $p_1^*$ because there are four unknown parameters $(m_0,s,p_0^*,p_1^*)$ in the system of two equations. As a result, neither $\beta^*$ nor $\Delta p^*$ can be identified based on $(Y,T,Z)$.
The IV estimand based on the observables may over- or underestimate $\beta^*$. The IV estimand based on the observables is
From the relationship in (ref), the difference between the denominators of $\beta$ and $\beta^*$ (i.e., the difference between the observable and true first-stage regressions) is expanded as
which can be negative or positive depending on the misclassification probabilities. For example, if the misclassification probabilities for the truly untreated given $Z=0$ and $Z=1$ are the same so that $m_{00}=m_{01}$, the sign of the difference between the first-stage regressions depends on that of $(m_{10}+m_{11})/(m_{00}+m_{01})-p_0^*/p_1^*$. As a result, observable $\beta$ may over- or underestimate true $\beta^*$.
This section develops the identification of the LATE with misclassification. Section (ref) shows the identification result based on an exogenous variable $V$. Section (ref) discusses the moment conditions derived from the identification result and GMM estimation.
We need the following assumptions to identify the LATE and misclassification probabilities based on the use of exogenous $V$, similarly to mahajan2006identification and lewbel2007estimation.
Assumption (ref) implies that mismeasured $T$ has no information on the mean of $Y$ once true $T^*$, $Z$, and $V$ are conditioned on. With the measurement error $U_T = T-T^*$, this assumption means that $E(Y|U_T,T^*,Z,V)=E(Y|T^*,Z,V)$, implying that $U_T$ is a {\it non-differential} measurement error. The assumptions of non-differential errors are popular in the misclassification literature. The non-differential error may require that the error does not depend on the unobservables affecting outcome $Y$ and/or true treatment $T^*$. The non-differential error also rules out a placebo effect such that the misclassified treatment affects the outcome of the individual who does not actually receive the treatment.
For example, consider the returns to schooling analysis. Under Assumption (ref), the average wages for individuals who report graduating and, indeed, not graduating are identical to that for individuals who did not graduate with precise reports. This could be satisfied if misclassification is caused by accident or by false statements depending the observable variables. On the contrary, the assumption might be violated if misclassification depends on the causal effect $Y_1-Y_0$ and/or unobservables affecting wages.
Assumption (ref) means that the sum of the misclassification probabilities is less than one. The assumption is popular and known as the monotonicity condition in the misclassification literature. This is satisfied when the actual reporting for the treatment has better information about the true treatment than the random reporting in which the misclassification probabilities $m_{0z}$ and $m_{1z}$ are equal to half. Indeed, $T^*$ is positively correlated with $T$ under this assumption according to (ref), implying that the observed treatment has positive information on the true treatment.
To introduce the next assumption, we define the following shorthand notations:
for $t,z=0,1$ and $v \in supp(V) \subset \mathbb{R}$. Here, $m_{tzv}$ is the conditional misclassification probability, $p_{zv}^*$ is the conditional true treatment probability, and $\tau_{zv}^*$ and $\tau_z^*$ are the differences in the conditional outcome means.
Assumption (ref) contains a set of exclusion restrictions requiring that exogenous $V$ does not affect $\tau_{zv}^*$ and $m_{tzv}$. A sufficient condition of the assumption about $\tau_{zv}^*$ is that the functional form of $E(Y|T^*,Z,V)$ is given by $E(Y|T^*,Z,V) = h_1(T^*,Z) + h_2(Z,V)$ for functions $h_1$ and $h_2$. The assumption could hold, especially when the effect of $V$ on $Y_1$ is the same as that on $Y_0$. Remarkably, Assumption (ref) allows a direct effect of $V$ on $Y$, so that $V$ can be a covariate. For example, even in the linear model of $E(Y|T^*,Z,V) = \gamma_0 + \gamma_1 T^* + \gamma_2 Z + \gamma_3 V + \gamma_4 ZV$ with coefficients $\gamma$s, the assumption is satisfied. The other special case in which the assumption may hold is when $V$ is randomly assigned by an experiment. Also, when $V$ is a repeated measure of the treatment, the assumption of $\tau_{zv}^*$ requires that the error for $V$ is a non-differential error.
Under Assumption (ref), the generating process of the misclassification also does not depend on $V$. However, it allows the misclassification probabilities to depend on $Z$. We note that when $V$ is a repeated measure of the treatment that may contain a measurement error, Assumption (ref) requires that the error for $V$ is independent of the error for $T$. Indeed, if the error for $V$ is correlated with the error for $T$, the misclassification probability $m_{tzv}$ may depend on the value of $V$.
Assumption (ref) further includes a relevance condition under which the true treatment probability $p_{zv}^*$ depends on the value of $V$. This holds when $V$ has a direct effect on $T^*$ or when $V$ is a repeated measure so that $V$ relates to $T^*$. Importantly, the relevance condition is testable under the exclusion restriction that $m_{tz}=m_{tzv}$ in Assumption (ref). Indeed, with the definition of the conditional treatment probability
the same procedure for showing (ref) leads to $p_{zv} = m_{0zv} + (1-m_{0zv}-m_{1zv}) p_{zv}^*$. This implies that $p_{zv}^* \neq p_{zv'}^*$ for $v, v' \in \Omega_z$ if and only if $p_{zv} \neq p_{zv'}$ under exclusion restriction $m_{tz}=m_{tzv}$.
To understand the practical implications of Assumption (ref), let us consider the returns to schooling analysis. Suppose that $Z$ is the indicator of four-year college proximity and that $V$ is the indicator of two-year college proximity. The exclusion restriction for $\tau_{zv}^*$ is satisfied if the mean effect of the college degree on wages for individuals' proximity to two-year colleges is identical to that for individuals who do not live close to two-year colleges. The condition for $m_{tzv}$ holds when two-year college proximity does not determine the generation of the measurement error. The relevance condition for $p_{zv}^*$ is satisfied when two-year college proximity is related to true educational attainment.
For the next assumption, we define the following quantity:
for $z=0,1$ and $v \in supp(V) \subset \mathbb{R}$. Here, $\tau_{zv}$ is the difference in the conditional outcome means, which is identified from the observable data.
The following assumption includes the conditions to solve the systems of linear equations for the identification of the misclassification probabilities. Condition (i) is for the case where $V$ takes at least three values. On the contrary, condition (ii) allows the situation where $V$ is a binary covariate, instrument, or repeated measure of the treatment. We remember the set $\Omega_z \subset supp(V)$ in Assumption (ref).
Assumption (ref) contains somewhat technical inequalities that ensure the unique solutions for the systems of linear equations. Note that the inequalities are testable in principle since their components are identified by the data. Also, as shown in the supplementary appendix, the necessary and sufficient conditions of the inequalities are $\tau_z^* \neq 0$ and $m_{0z} + m_{1z} \neq 1$. For example, in the returns to schooling example, the inequalities hold if and only if average wages for college graduates are different from those for individuals without college degrees.
Condition (ii) also requires the additional exclusion restriction that the misclassification probabilities do not depend on $Z$. We need the additional exclusion restriction since binary $V$ is less informative for the distribution of $T^*$ than non-binary $V$ under the exclusion restrictions in Assumption (ref). We note that calvi2017women, ditraglia2015mis, and ura2017hetero also assume the same exclusion restriction.
The following theorem states the main identification result of this study.
The main idea behind the identification is, similar to mahajan2006identification and lewbel2007estimation, to construct moment equalities whose number is no fewer than the unknown parameters. The idea can be understood by a simple sketch of the proof of Theorem (ref). This proof depends on whether we assume Assumption (ref) (i) or (ii). Under Assumptions (ref), (ref), (ref), and (ref) (i), we can show the following system of equations for each $z=0,1$:
where the $B$s are parameters to be identified in our analysis and $w$s are identified by data. The definitions of the $B$s and $w$s are given in (ref) in the proof of Theorem (ref). Since the $w$s are identified from the observable data and Assumption (ref) (i) guarantees the existence of the unique solutions of the $B$s, the $B$s are identified. Similarly, under the assumption of $m_{t}=m_{t0} = m_{t1}$ for $t=0,1$ in Assumption (ref) (ii), we instead have
We note that the $B$s in (ref) do not depend on $z$ unlike the $B$s in (ref). In this case, Assumption (ref) (ii) guarantees that the $B$s are identified as the solution of simultaneous equations (ref). As a next step, the identification of the $B$s leads to identifying the misclassification probabilities. Specifically, under Assumption (ref) (i), we have
for each $z=0,1$. On the contrary, under Assumption (ref) (ii), we have
These equations mean that the misclassification probabilities are identified. Finally, the true first-stage regression is identified based on the information on the misclassification probability and observable treatment probability. Under Assumption (ref) (i), we have
and, under Assumption (ref) (ii), we have
Hence, the LATE $\beta^* = \Delta \mu/\Delta p^*$ is identified.
We derive the moment conditions based on the identification result in Theorem (ref). These moment conditions lead to estimation procedures such as the non-linear GMM estimation and empirical likelihood estimation. Here, we focus on the moment conditions under Assumption (ref) (i) since those under Assumption (ref) (ii) are simpler. We assume that $V$ is a discrete random variable since covariates, instruments, and repeated measures are often discrete in practice. We can easily derive similar moment conditions when $V$ is continuously distributed and/or when other control variables exist.
Define the vector of the observable variables $X \coloneqq (Y,T,Z,V)^\top$. For simplicity, suppose that discrete $V$ takes $K$ values in $supp(V) = \Omega_0 = \Omega_1 = \{v_1, v_2, \dots, v_K\}$, where $\Omega_0$ and $\Omega_1$ are introduced in Assumptions (ref) and (ref). Define a vector of the $K+3$ parameters:
The vector $\theta_0^{(z)}$ for $z=0,1$ contains the parameters conditional on $Z=z$ introduced in Section (ref). The vector of the $2K+9$ parameters to be estimated is
where $r \coloneqq E(Z)$. The parameters except for the LATE $\beta^*$ and the true first-stage regression $\Delta p^*$ are nuisance parameters to overcome the misclassification problem. Nonetheless, as well as $\beta^*$ and $\Delta p^*$, the misclassification probability $m_{tz}$ could be of interest in practice. Let $\theta$ be a general parameter value in the parameter space $\Theta \subset \mathbb{R}^{2K+9}$ to which $\theta_0$ also belongs.
Let $g(X, \theta_0)$ be the vector valued function with $4K+3$ elements, which are the following components: for $k=1,2,\dots,K$ and $z=0,1$,
where $I_{zv_k} \coloneqq \mathbf{1}(Z=z,V=v_k)$ is the indicator.
The following theorem shows that the moment condition is $E[g(X,\theta_0)]=0$ with the unique solution $\theta_0 \in \Theta$. This is directly shown by the identification result in Theorem (ref).
The number of overidentification restrictions depends on the number of elements in support of $V$, that is, $K$. As stated above, the numbers of parameters and moment equations are $2K+9$ and $4K+3$, respectively, under Assumption (ref) (i). Hence, there are $2K-6$ overidentification restrictions and $\theta_0$ is just-identified when $K=3$.
Similarly, under Assumption (ref) (ii), we can show that there are $2K-4$ overidentification restrictions, and just-identification is achieved when $K=2$.
This section reports the results of the Monte Carlo simulations. The simulations are conducted by { R} with 2,000 simulation replications.
\paragraph{DGP.} By using sample sizes of 200, 500, and 1,000, we generate the random variables by adopting the following six data-generating processes. For all designs, we generate $Z \in \{0,1\}$ with probabilities $\Pr(Z=0) = \Pr(Z=1) = 0.5$ and the following unobservables affecting the outcome and the true treatment:
In designs 1 and 2, we presume $V$ to be a covariate and consider two different designs for $Y$. Assuming Assumption (ref) (ii), we generate binary $V$ with probabilities $\Pr(V=0)=\Pr(V=1)=0.5$. The true treatment is generated by the probit model $T^*=\mathbf{1}(-1 + Z + V - U_1>0)$. The observed treatment $T$ is misclassified independently of the other variables with misclassification probability $m_t = \Pr(T \neq T^*|T^*=t)=0.25$ for each $t=0,1$. Two designs are considered for the outcome variable:
Design 1 considers the homogeneous causal effect, whereas the causal effect in design 2 is heterogeneous depending on unobserved $U_2$. In both designs, $V$ directly affects $Y$.
In designs 3 and 4, we presume $V$ to be an instrument. Generating binary $V$ with probabilities $\Pr(V=0)=\Pr(V=1)=0.5$, the true treatment is generated by $T^*=\mathbf{1}(-1 + Z + V - U_1>0)$. The misclassification probability is $m_t = \Pr(T \neq T^*|T^*=t)=0.25$ for each $t=0,1$. The outcome is generated by
In designs 5 and 6, we presume $V$ to be a binary repeated measure. With the true treatment $T^*=\mathbf{1}(-0.5 + Z - U_1 > 0)$, the observed treatment $T$ and repeated measure $V$ have the misclassification probabilities $m_t = \Pr(T \neq T^* | T^* = t) = 0.25$ and $\Pr(V \neq T^* | T^* = t) = 0.3$ for $t= 0, 1$ independently of other variables. Note that $V \neq T$ in general. The outcome is generated by
\paragraph{Estimators.} We consider two estimators. The first is the GMM estimator based on the identification result developed in this study. The second is the naive IV estimator $\beta$ in (ref), which is a benchmark estimator.
In this simulation, the GMM estimation involves the 11 parameters to be estimated, that is, $\theta_0 = (\beta^*, \Delta p^*, r,m_0,p_{00}^*,p_{01}^*,\tau_0^*,m_1,p_{10}^*,p_{11}^*,\tau_1^*)^\top$. We focus on the parameters of interest of the LATE $\beta^*$, first-stage regression $\Delta p^*$, and misclassification probabilities $m_0$ and $m_1$. The moment condition $E[g(X, \theta_0)]=0$ is composed of 11 elements. Since $\theta_0$ is just-identified here, we select the identity matrix as the weighting matrix.
\paragraph{Result.} Tables (ref), (ref), and (ref) summarize the Monte Carlo simulation results for $\beta^*$, $\Delta p^*$, $m_0$, and $m_1$. Since the main parameter of interest is $\beta^*$, we focus on the results reported in Table (ref). The table reports the true value and biases, standard deviations (SDs), and root mean squared errors (RMSEs) for the proposed GMM estimator and naive IV estimator ignoring the measurement error. It also indicates the coverage probabilities (CPs) of the 95% confidence intervals based on the asymptotic normal approximations for the estimators with heteroscedasticity robust standard errors.
The performance of our GMM estimation is successful. The bias and SD of the GMM estimator are satisfactory in all designs. For example, the biases of the GMM estimator for the LATE when $n=1000$ can be about 10% of the true values. The RMSE of the GMM estimator in each design is also moderate and becomes smaller as the sample size increases, which is expected owing to the asymptotics of the GMM estimator. Further, the CPs of the 95% confidence interval are close to 0.95 even with small sample sizes.
Table (ref) also shows that the naive IV estimator $\beta$ based on the observables exhibits large biases in all designs. In each design, the bias is over 100% of the true value of the LATE, and the naive IV estimator overestimates the LATE. The result is driven by the identification failure of the naive IV estimand for the LATE, as we discuss in Section (ref), and the magnitude of the bias is consistent with our theoretical investigation in (ref).
We illustrate our proposed procedure by analyzing returns to schooling. We use the same dataset used by card1993using. The data are originally drawn from the National Longitudinal Survey of Young Men, and the number of the observations is 3,010. See card1993using for the details on the features of the dataset.
The variables we use are log hourly wages ($lwage$), the schooling indicator ($college$) whether years of schooling are no less than 14, that is, the indicator corresponding degrees greater than two-year college degrees, the indicator of two-year college proximity ($proximity2$), and the indicator of four-year college proximity ($proximity4$). We call individuals with $college = 1$ college graduates, while we recognize that those individuals might include college dropouts. We may also use the other control variables contained in the original dataset such as age, although we do not use those variables here for two reasons. First, to incorporate those variables into the GMM estimation, we need functional form specifications such as linear models for the functions in the parameters, which might cause misspecification problems. Second, the results of the naive estimation ignoring the measurement error are not sensitive to whether we control for those variables.
Table (ref) presents the descriptive statistics for our variables. Table (ref) reports the results of the naive OLS and IV estimation ignoring the presence of the measurement error.
For our proposed inference, we utilize two indicators of college proximity to deal with the endogeneity and misclassification problems simultaneously. Our main result in this illustration uses $proximity4$ as $Z$ and $proximity2$ as $V$. This specification is based on the fact that the observed first-stage regression when using $proximity4$ as $Z$ is significantly stronger than that when using $proximity2$ as $Z$, as shown in columns (4) and (5) in Table (ref). Nonetheless, we also present the GMM estimates when using $proximity2$ as $Z$ and $proximity4$ as $V$ in Table (ref) as an auxiliary result; however, these estimates are less precise than our main estimates below.
The validity of the exclusion restrictions and relevance condition in Assumption (ref) should be examined since they are especially important among our identification conditions. These exclusion restrictions require that the effect of the true college degree on wages and the distribution for misreporting the college degree do not depend on whether individuals live close to two-year colleges. Importantly, college proximity can affect wages for our inference, which would be desirable since several studies point out that college proximity might be related to factors affecting wages (e.g., carneiro2002evidence). The relevance condition requires that the true college degree is related to two-year college proximity. Its validity is statistically testable as discussed in Section (ref); columns (6) and (7) in Table (ref) present the results of such a test. Column (6) shows evidence of $p_{11}^* \neq p_{10}^*$ at a significance level of 5%. Column (7) shows evidence of $p_{01}^* \neq p_{00}^*$ at a significance level of 10%, owing to the small sample size of the data with $proximity4 = 0$, which leads to a relatively large standard error. The estimates in columns (6) and (7) are non-negligible amounts, meaning that they imply the validity of our relevance condition.
Table (ref) reports the result of our GMM estimation when using $proximity4$ as $Z$ and $proximity2$ as $V$. The GMM estimate for the LATE shows that having a college degree increases average wages for compliers by about 42%, whereas the naive estimation shows an average effect of about 131%. This severe difference is caused by the misclassification problem, namely that the observed first-stage regression ignoring the measurement error underestimates the true first-stage regression. These results demonstrate the importance of taking account of the misclassification problem as well as the usefulness of our GMM inference.
Table (ref) also reports the estimates for the misclassification probabilities, showing an interesting result that individuals with a true college degree tend to misreport more than those without. These estimated misclassification probabilities might seem to be significantly high, but they could be reliable estimates since the estimated effect of being a college graduate from the GMM estimation is more consistent with existing empirical results than is the naive IV estimate. For example, by using other datasets, kane1999estimating, lewbel2007estimation, and battistin2014misreported show that the average effects of undergraduate education on wages are about 25%, 37%, and 27%, respectively if taking account of the presence of misreported schooling.
This study presents novel point-identification for the LATE when the binary endogenous treatment may contain a measurement error. Since the measurement error must be non-classical by construction, the standard IV estimator cannot consistently estimate the LATE. To correct the bias due to the measurement error, we build on the results presented by mahajan2006identification and lewbel2007estimation to point-identify the LATE, true first-stage regression, and misclassification probabilities based on the use of an exogenous variable such as a covariate, instrument, or repeated measure of the treatment. The moment conditions derived from the identification lead to the GMM estimator for the parameters with asymptotically valid inferences. The simulations and empirical illustration demonstrate the usefulness of the proposed inference.
Several important future research topics can be proposed on the basis of our findings. First, it is desirable to develop point-identification for the LATE with a {\it differential} (i.e., endogenous) measurement error for the treatment. The measurement error is differential when it is related to the outcome variable, even conditional on the true variable. How to handle such an error would be of interest but challenging. Second, it would be of interest to develop inferences for the LATE with a measurement error when the treatment is a general discrete variable. Without a measurement error for the discrete treatment, the IV estimand identifies the LATE with the variable treatment intensity (angrist1995two). Our inference may be extended to such situations. Third, the proposed identification may be extended to other settings such as regression discontinuity (RD) designs with the mismeasured treatment. Since RD designs require local inferences around thresholds, we cannot directly use our proposed estimation for RD inferences. The author develops this topic as another project (yanagi2015regression).