EconBase
← Back to paper

Identifying the effect of a mis-classified, binary, endogenous regressor

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

41,650 characters · 8 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identifying the Effect of a Mis-classified, Binary, Endogenous Regressor

\thispagestyle{empty}

abstract\singlespacing This paper studies identification of the effect of a mis-classified, binary, endogenous regressor when a discrete-valued instrumental variable is available. We begin by showing that the only existing point identification result for this model is incorrect. We go on to derive the sharp identified set under mean independence assumptions for the instrument and measurement error. The resulting bounds are novel and informative, but fail to point identify the effect of interest. This motivates us to consider alternative and slightly stronger assumptions: we show that adding second and third moment independence assumptions suffices to identify the model. \noindentKeywords: Instrumental variables, Measurement error, Endogeneity \noindentJEL Codes: C10, C25, C26

\setcounter{page}{1}

\singlespacing

Introduction

Measurement error and endogeneity are pervasive features of economic data. Conveniently, a valid instrumental variable corrects for both problems when the measurement error is classical, i.e.\ uncorrelated with the true value of the regressor. Many regressors of interest in applied work, however, are binary and thus cannot be subject to classical measurement error.\footnote{The only way to mis-classify a true one is downwards, as a zero, while the only way to mis-classify a true zero is upwards, as a one. This creates negative dependence between the truth and measurement error.} When faced with non-classical measurement error, the instrumental variables estimator can be severely biased. In this paper, we study an additively separable model of the form

equation[equation omitted — 92 chars of source]

where $\varepsilon$ is a mean-zero error term, $T^*$ is a binary, potentially endogenous regressor of interest, and $\mathbf{x}$ is a vector of exogenous controls.\footnote{Because $T^*$ is binary, there is no loss of generality from writing the model in this form rather than the more familiar $y = h(T^*,\mathbf{x})+\varepsilon$. Simply define $\beta(\mathbf{x}) = h(1,\mathbf{x}) - h(0,\mathbf{x})$ and $c(\mathbf{x}) = h(0,\mathbf{x})$.} We ask whether, and if so under what conditions, a discrete instrumental variable $z$ suffices to non-parametrically identify the causal effect $\beta(\mathbf{x})$ of $T^*$, when we observe not $T^*$ but a mis-classified binary surrogate $T$.

We proceed under the assumption of non-differential measurement error. This condition has been widely used in the existing literature and imposes that $T$ provides no additional information beyond that contained in $(T^*,\mathbf{x})$. Even in this fairly standard setting, identification remains an open question: we begin by showing that the only existing identification result for this model is incorrect. We then go on to derive the sharp identified set under the standard first-moment assumptions from the related literature. We show that regardless of the number of values that $z$ takes on, the model is not point identified. This motivates us to consider alternative, and slightly stronger assumptions. We show that, given a binary instrument, the addition of a second moment independence assumption suffices to identify a model with one-sided mis-classification. Adding a second moment restriction on the measurement error along with a third moment independence assumption for the instrument suffices to identify the model in general. This result likewise requires only a binary $z$.

Our work relates to a large literature that considers departures from classical measurement error, by allowing the measurement error to be related to the true value of the unobserved regressor. ChenHongTamer obtain identification in a general class of moment condition models with mis-measured data by relying on the existence of an auxiliary dataset from which they can estimate the measurement error process. In contrast, HuSchennach and song2015 rely on an instrumental variable and an additional conditional location assumption on the measurement error distribution. More recently, HuShiuWoutersen use a continuous instrument to identify the ratio of partial effects of two continuous regressors, one measured with error, in a linear single index model. Unfortunately, these approaches cannot be applied to the case of a mis-measured binary regressor.

A number of papers have studied models with an exogenous binary regressor subject to non-differential measurement error. One group of papers asks what can be learned without recourse to an instrumental variable. An early contribution by Aigner characterizes the asymptotic bias of OLS in this setting, and proposes a correction using outside information on the mis-classification process. Related work by Bollinger provides partial identification bounds. More recently, ChenHuLewbel use higher moment assumptions to obtain identification in a linear model, and ChenHuLewbel2 extend these results to the non-parametric setting. HasseltBollinger and BollingerHasseltWP provide additional partial identification results. For results on the partial identification of discrete probability distributions under mis-classification, see molinari.

Continuing under the assumption of exogeneity and non-differential measurement error, another group of papers relies on the availability of either an instrumental variable or a second measure of $T^*$. BBS and KRS consider a linear model and show that when two alternative measures $T_1$ and $T_2$ of $T^*$ are available, a non-linear GMM estimator can be used to recover the effect of interest. Subsequently, FL note that an instrumental variable can take the place of one of the measures. Mahajan extends the results of BBS and KRS to a more general setting using a binary instrument in place of one of the treatment measures, establishing non-parametric identification of the conditional mean function. When $T^*$ is in fact exogenous, this coincides with the causal effect. hu2008 derives related results when the mis-classified discrete regressor may take on more than two values. Lewbel provides an identification result for the same model as Mahajan under different assumptions. In particular, his “instrument-like variable” need not satisfy the usual exclusion restriction so long as it does not interact with $T^*$ and takes on three or more values.

Much less is known about the case in which a binary, or discrete, regressor is not only mis-classified but endogenous. The first paper to provide a formal result for this case is Mahajan. He extends his main result to the case of an endogenous treatment, providing an explicit proof of identification under the usual IV assumption in a model with additively separable errors. As we show below, however, this result is false.\footnote{Appendix (ref) provides a detailed explanation of the error in Mahajan's proof.} Several more recent papers also consider the case of a mis-classified, endogenous, binary regressor. kreider2012, partially identify the effects of food stamps on health outcomes of children under weak measurement error assumptions by relying on auxiliary data. Similarly, Batt study the returns to schooling in a setting with multiple mis-reported measures of educational qualifications. Unlike these two papers, our approach does not depend on the availability of auxiliary data. In a different vein, shiu2015 uses an exclusion restriction for the participation equation and an additional valid instrument to identify the effect of a discrete, mis-classified endogenous regressor in a semi-parametric selection model. Similarly, nguimkeu2016estimation use exclusion restrictions for both the participation equation and measurement error equation to identify a parametric model with endogenous participation and one-sided endogenous mis-reporting. Unlike those of the preceding two papers, our results rely neither on parametric assumptions nor additional exclusion restrictions. Other than Mahajan, the paper most closely related to our own is that of Ura, who derives partial identification results for a local average treatment effect without the non-differential assumption. In contrast, we study an additively separable model under non-differential measurement error and derive both partial and point identification results.

The remainder of the paper is organized as follows. Section (ref) describes our model and assumptions, Section (ref) relates our results to existing work, and Sections (ref)--(ref) present our identification results. Section (ref) provides a brief discussion of how to carry out inference using our identification results, and Section (ref) concludes. Proofs appear in Appendix (ref), and we give a detailed explanation of the error in Mahajan in Appendix (ref). Appendix (ref) explains how our partial identification bounds from Section (ref) can be interpreted in a local average treatment effects (LATE) setting.

Identification

Baseline Assumptions

As defined in the preceding section, our model is $y = c(\mathbf{x}) + \beta(\mathbf{x}) T^* + \varepsilon$, where $\varepsilon$ is a mean-zero error term, and the parameter of interest is $\beta(\mathbf{x})$ -- the effect of an unobserved, binary, endogenous regressor $T^*$. Suppose we observe a valid and relevant binary instrument $z$. In the discussion following Corollary (ref) below, we explain how these results generalize to the case of an arbitrary discrete-valued instrument. We assume that the model and instrument satisfy the following conditions:

assump\begin{enumerate}[(i)] • $y = c(\mathbf{x}) + \beta(\mathbf{x})T^* + \varepsilon$ where $T^* \in \left\{ 0,1 \right\}$ and $\mathbb{E}[\varepsilon]=0$; • $z \in \left\{ 0,1 \right\}$, where $0 < \mathbb{P}(z=1|\mathbf{x}) < 1$, and $\mathbb{P}(T^*=1|\mathbf{x},z=1) \neq \mathbb{P}(T^*=1|\mathbf{x},z=0)$; • $\mathbb{E}[\varepsilon|\mathbf{x},z] = 0$. \end{enumerate}

Assumption (ref)(i) is a restatement of the additively separable model from Equation (ref), which includes as a special case the linear model $y = c + \beta T^* + \mathbf{x}'\boldsymbol{\gamma} + \varepsilon$ that is pervasive in empirical economics. Assumptions (ref)(ii) and (iii) are the textbook instrumental variable relevance and validity conditions, respectively. Under Assumption (ref), the Wald estimator \[ \left[\mathbb{E}\left(y|z=1,\mathbf{x}\right)-\mathbb{E}\left(y|z=0,\mathbf{x}\right)\right] / \left[ \mathbb{E}\left(T^*|z=1,\mathbf{x}\right) - \mathbb{E}\left( T^*|z=0,\mathbf{x} \right) \right] \] identifies $\beta(\mathbf{x})$. Unfortunately this estimator is infeasible, as we observe not $T^*$ but a mis-classified binary surrogate $T$.\footnote{Although it involves $T^*$, Assumption (ref)(ii) is testable: see the discussion following Lemma (ref).} To make further progress, we must impose conditions on the process that generates $T$. Accordingly, define the following mis-classification probabilities:

align*[align* omitted — 316 chars of source]
assump\begin{enumerate}[(i)] • $\alpha_0(\mathbf{x},z) = \alpha_0(\mathbf{x})$, $\alpha_1(\mathbf{x},z) = \alpha_1(\mathbf{x})$$\alpha_0(\mathbf{x}) + \alpha_1(\mathbf{x}) <1$$\mathbb{E}[\varepsilon|\mathbf{x},z,T^*,T] = \mathbb{E}[\varepsilon|\mathbf{x},z, T^*]$ \end{enumerate}

Assumption (ref), or a variant thereof, is standard in the theoretical literature on mis-classification Mahajan,BBS,FL,Lewbel,hu2008 and in empirical studies that allow for measurement error in a binary or discrete variable KRS,hu2013,Batt. Assumption (ref) (i) states that the mis-classification probabilities do not depend on $z$. Assumption (ref) (ii) restricts the extent of mis-classification and is equivalent to requiring that $T$ and $T^*$ be positively correlated. Assumption (ref) (iii) is often referred to as “non-differential measurement error.” Intuitively, it maintains that $T$ provides no additional information about $\varepsilon$, and hence $y$, given knowledge of $(T^*,z,\mathbf{x})$. While Assumption (ref)(ii) is quite mild, Assumptions (ref) (i) and (iii) are more restrictive, as discussed by Bound2001. To take a specific example, suppose that $y$ is log wage and $T^*$ is an indicator for college completion. If $T$ is a potentially erroneous measure of college completion taken from a university's administrative records, then the assumption of non-differential measurement error is quite plausible. If, on the other hand, $T$ is a self-report of college completion and there are “returns to lying” about college completion, i.e.\ employers only imperfectly observe worker ability, this assumption is less plausible.\footnote{See huLewbel for a proposal to estimate the “returns to lying” in this context.} Note, however, that our assumptions on the mis-classification process are conditional on $\mathbf{x}$: we place no restrictions on the relationship between observed covariates and the mis-classification errors. In contrast, Bound2001 considers unconditional versions of our Assumption (ref). Instrument validity -- Assumption (ref) (iii) -- is more plausible after conditioning on a rich set of exogenous controls, and the same is true of our mis-classification assumptions. For more discussion of settings in which the assumption of non-differential measurement error is warranted, see carroll2006.

Point Identification Results from the Literature

Existing results from the literature -- see for example FL and Mahajan -- establish that $\beta(\mathbf{x})$ is point identified if Assumptions (ref)--(ref) are augmented to include the following condition:

assump[Joint Exogeneity] $\mathbb{E}[\varepsilon|\mathbf{x},z, T^*] = 0$.

Assumption (ref) strengthens the mean independence condition from Assumption (ref) (iii) to hold jointly for $T^*$ and $z$. By iterated expectations, this implies that $T^*$ is exogenous, i.e.\ $\mathbb{E}[\varepsilon|\mathbf{x},T^*] = 0$. If $T^*$ is endogenous, Assumption (ref) clearly fails. Mahajan argues, however, that the following restriction, along with our Assumptions (ref)--(ref), suffices to identify $\beta(\mathbf{x})$ when $T^*$ may be endogenous:

assump[Mahajan Equation 11] $\mathbb{E}[\varepsilon|\mathbf{x}, z, T^*, T] = \mathbb{E}[\varepsilon|\mathbf{x},T^*]$.

Assumption (ref) does not require $\mathbb{E}[\varepsilon|\mathbf{x},T^*]$ to be zero, but maintains that it does not vary with $z$. We show in Appendix (ref), however, that under Assumptions (ref)--(ref), Assumption (ref) can only hold if $T^*$ is exogenous. If $z$ is a valid instrument and $T^*$ is endogenous, then Assumption (ref) implies that there is no first-stage relationship between $z$ and $T^*$. As such, identification in the case where $T^*$ is endogenous is an open question.

Partial Identification

In this section we derive the sharp identified set under Assumptions (ref)--(ref) and show that $\beta(\mathbf{x})$ is not point identified. For a discussion of how our partial identification results can be interpreted in a local average treatment effects (LATE) setting, see Appendix (ref).

To simplify the notation, define the following shorthand for the unobserved and observed first stage probabilities

equation[equation omitted — 149 chars of source]

We first state two lemmas that that will be used repeatedly below.

lemUnder Assumption (ref) (i), \begin{align*} \left[ 1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x}) \right]p^*_k(\mathbf{x}) &= p_k(\mathbf{x}) - \alpha_0(\mathbf{x})\\ \left[ 1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x}) \right]\left[1 - p^*_k(\mathbf{x}) \right]&= 1 - p_k(\mathbf{x}) - \alpha_1(\mathbf{x}) \end{align*} where the first-stage probabilities $p_k^*(\mathbf{x})$ and $p_k(\mathbf{x})$ are as defined in Equation (ref).
lemUnder Assumptions (ref) and (ref) (i)--(ii), $$\beta(\mathbf{x}) \mbox{Cov}(z,T|\mathbf{x}) = \left[ 1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x}) \right]\mbox{Cov}(y,z|\mathbf{x})$$

Lemma (ref) relates the observed first-stage probabilities $p_k(\mathbf{x})$ to their unobserved counterparts $p^*_k(\mathbf{x})$ in terms of the mis-classification probabilities $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$. By Assumption (ref) (ii), $1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x}) > 0$ so that Lemma (ref) bounds $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$ in terms of the observed first-stage probabilities. Moreover, by taking differences evaluated at $k=1$ and $k=0$, this Lemma shows that $p_0^*(\mathbf{x}) = p_1^*(\mathbf{x})$ if and only if $p_0(\mathbf{x})=p_1(\mathbf{x})$. In other words, Assumption (ref) (ii) is testable under Assumption (ref) (ii). Lemma (ref) relates the instrumental variables (IV) estimand, $\mbox{Cov}(y,z|\mathbf{x})/\mbox{Cov}(z,T|\mathbf{x})$, to the mis-classification probabilities. Since $1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x}) > 0$, IV is biased upwards in the presence of mis-classification. Together these lemmas bound the causal effect of interest: $\beta(\mathbf{x})$ lies between the reduced form and IV estimators. Without Assumption (ref) (iii), non-differential measurement error, these bounds are sharp.

thmUnder Assumptions (ref) and (ref) (i)--(ii), $\alpha_0(\mathbf{x}) \leq p_k(\mathbf{x}) \leq 1 - \alpha_1(\mathbf{x})$ for $k = 0, 1$ and \begin{equation} \mathbb{E}[y|\mathbf{x},z=k] = c(\mathbf{x}) + \beta(\mathbf{x}) \left[\frac{p_k(\mathbf{x}) - \alpha_0(\mathbf{x})}{1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x})}\right]. \end{equation} Provided that $p_0(\mathbf{x}) \neq p_1(\mathbf{x})$, these expressions characterize the sharp identified set for $c(\mathbf{x})$, $\beta(\mathbf{x})$, $\alpha_0(\mathbf{x})$, and $\alpha_1(\mathbf{x})$.
corUnder the conditions of Theorem (ref), the sharp identified set for $\beta(\mathbf{x})$ is the closed interval between the reduced form estimand $\mbox{Cov}(y,z|\mathbf{x})/\mbox{Var}(z|\mathbf{x})$ and the IV estimand $\mbox{Cov}(y,z|\mathbf{x})/\mbox{Cov}(z,T|\mathbf{x})$.

Corollary (ref) follows by taking differences of the expression for $\mathbb{E}[y|\mathbf{x},z=k]$ across $k=1$ and $k=0$, and substituting the maximum and minimum value for $\alpha_0(\mathbf{x}) + \alpha_1(\mathbf{x})$ consistent with the observed first-stage probabilities.\footnote{If a priori restrictions on $\alpha_0$ and $\alpha_1$ are available, e.g.\ $\alpha_0 = 0$, $\alpha_1 = 0$, or $\alpha_0 = \alpha_1$, these bounds can be improved. For more discussion, see Corollary 2.2 of DiTragliaGarciaWP2017.} Note that the only role of the condition $p_0(\mathbf{x}) \neq p_1(\mathbf{x})$ in the preceding two results is to ensure that it is possible to satisfy Assumption (ref) (ii). FL point out that the IV estimand provides an upper bound for $\beta(\mathbf{x})$, and Lemmas (ref)--(ref) are well-known in the literature FL,Mahajan. Nevertheless, we are unaware of any published result that explicitly states both bounds from Corollary (ref) or proves that they are sharp under Assumptions (ref) and (ref) (i)--(ii).

Neither Theorem (ref) nor Corollary (ref) imposes Assumption (ref) (iii) -- non-differential measurement error. While this assumption plays an important role in existing identification results for an exogenous $T^*$ (see Section (ref)), its identifying power under endogeneity has not been addressed in the literature.\footnote{The only exception is the incorrect result of Mahajan described in Section (ref) and Appendix (ref).} We now show that this assumption in general yields further restrictions on probabilities $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$, but fails to point identify $\beta(\mathbf{x})$. To simplify the proof of sharpness, we assume that $y$ is continuously distributed, which is natural in an additively separable model. Without this assumption, the bounds that we derive are still valid, but may not be sharp. Nevertheless, the reasoning from our proof can be generalized to cases in which $y$ does not have a continuous support set.

thmSuppose that the conditional distribution of $y$ given $(\mathbf{x}, T,z)$ is continuous. Further suppose that the conditions of Theorem (ref) and Assumption (ref) (iii) hold. For any $k$ such that $\mathbb{E}\left[ y|\mathbf{x},T=0,z=k \right] \neq \mathbb{E}\left[ y|\mathbf{x},T=1,z=k \right]$, let $\mathcal{A}_k$ denote the set of pairs $\big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}) \big)$ such that $\alpha_0(\mathbf{x}) < p_k(\mathbf{x}) < 1 - \alpha_1(\mathbf{x})$ and \[ \underline{\mu}_{tk}\bigg( \underline{q}_{tk}\big( \alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}), \mathbf{x}\big) , \,\mathbf{x} \bigg)\leq \mu_{k}\big( \alpha_0(\mathbf{x}),\mathbf{x} \big)\leq \overline{\mu}_{tk}\bigg(\overline{q}_{tk}\big( \alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}), \mathbf{x}\big), \,\mathbf{x} \bigg) \] for all $t = 0,1$ where \begin{align*} \mu_{tk}\big( q,\mathbf{x} \big) = \mathbb{E}\left[ y\left|\right.y\leq q, \mathbf{x},T=t, z=k\right], \quad \quad \overline{\mu}_{tk}\big(q,\mathbf{x} \big) = \mathbb{E}\left[ y\left|\right. y > q, \mathbf{x}, T=t, z=k\right] \end{align*} \begin{align*} \mu_k\big(\alpha_0(\mathbf{x}),\mathbf{x}\big) &= \frac{p_k(\mathbf{x}) \mathbb{E}[y|\mathbf{x},z=k,T=1] - \alpha_0(\mathbf{x}) \mathbb{E}[y|\mathbf{x},z=k]}{p_k(\mathbf{x}) - \alpha_0(\mathbf{x})} \end{align*} and we define \begin{align*} q_{tk}\big(\alpha_0(\mathbf{x}),\alpha_1(\mathbf{x}),\mathbf{x}\big) &= F^{-1}_{tk}\bigg(r_{tk}\big(\alpha_0(\mathbf{x}),\alpha_1(\mathbf{x}), \mathbf{x}\big)\, \bigg|\,\mathbf{x}\bigg)\\ \overline{q}_{tk}\big(\alpha_0(\mathbf{x}),\alpha_1(\mathbf{x}),\mathbf{x}\big) &= F^{-1}_{tk}\bigg(1 - r_{tk}\big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}),\mathbf{x}\big) \,\bigg|\,\mathbf{x}\bigg) \end{align*} where $F_{tk}^{-1}(\cdot|\mathbf{x})$ is the conditional quantile function of $y$ given $(\mathbf{x},T=t,z=k)$, \begin{align*} r_{0k}\big(\alpha_0(\mathbf{x}),\alpha_1(\mathbf{x}),\mathbf{x}\big) &= \frac{\alpha_1(\mathbf{x})}{1 - p_k(\mathbf{x})} \left[ \frac{p_k(\mathbf{x}) - \alpha_0(\mathbf{x})}{1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x})} \right]\\ r_{1k}\big(\alpha_0(\mathbf{x}),\alpha_1(\mathbf{x}),\mathbf{x}\big) &= \frac{1 - \alpha_1(\mathbf{x})}{p_k(\mathbf{x})} \left[ \frac{p_k(\mathbf{x}) - \alpha_0(\mathbf{x})}{1 - \alpha_0(\mathbf{x}) - \alpha_1(\mathbf{x})} \right] \end{align*} and $p_k(\mathbf{x})$ is defined in Equation (ref). The sharp identified set for $c(\mathbf{x})$, $\beta(\mathbf{x})$, $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$ is characterized by Equation (ref) and $\big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x})\big) \in \mathcal{A}^*$ where \begin{enumerate}[(i)] • $\mathcal{A}^* \equiv \mathcal{A}_0 \cap \mathcal{A}_1$ if $\mathbb{E}[y|\mathbf{x},T=0,z=k] \neq \mathbb{E}[y|\mathbf{x},T=1,z=k]$ for all $k=0,1$; • $\mathcal{A}^* \equiv \mathcal{A}_k$ if $\mathbb{E}[y|\mathbf{x},T=0,z=k] \neq \mathbb{E}[y|\mathbf{x},T=1,z=k]$ and $\mathbb{E}[y|\mathbf{x},T=0,z=\ell] = \mathbb{E}[y|\mathbf{x},T=1,z=\ell]$; • $\mathcal{A}^* \equiv \left\{ \big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x})\big)\colon \alpha_0(\mathbf{x}) \leq p_k(\mathbf{x}) \leq 1 - \alpha_1(\mathbf{x}) \mbox{ for all } k \right\}$ if $\mathbb{E}[y|\mathbf{x},T=0,z=k] = \mathbb{E}[y|\mathbf{x},T=1,z=k]$ for all $k=0,1$. \end{enumerate}

Imposing Assumption (ref) (iii) strictly improves upon the identified set from Theorem (ref) unless $\mathbb{E}[y|\mathbf{x},T=0,z=k]=\mathbb{E}[y|\mathbf{x},T=1,z=k]$ for all $k$. Even if $\beta(\mathbf{x}) = 0$, the difference of these observable means is generically nonzero.\footnote{Suppress dependence on $\mathbf{x}$ for simplicity. There are only two settings in which $\mathbb{E}[y|T=0,z=k] = \mathbb{E}[y|T=1,z=k]$. The first is if the true value of either $\alpha_0$ or $\alpha_1$ lies at the upper boundary of the identified set from Theorem (ref). The second is if $\beta = \mathbb{E}[\varepsilon|T^*=0,z=k] - \mathbb{E}[\varepsilon|T^*=0,z=k]$.} The intuition for Theorem (ref) is as follows. For simplicity, suppress dependence on $\mathbf{x}$. Now, fix $(T=t, z=k)$ and $(\alpha_0, \alpha_1)$. The observed distribution of $y$ given $(T=t,z=k)$, call it $F_{tk}$, is a mixture of two unobserved distributions: the distribution of $y$ given $(T=k,z=k,T^*=1)$, call it $F^1_{tk}$, and the distribution of $y$ given $(T=t,z=k,T^*=0)$, call it $F^{0}_{tk}$. The mixing probabilities are $r_{tk}$ and $1-r_{tk}$ from the statement of Theorem (ref) and are fully determined by $(\alpha_0, \alpha_1)$ and $p_k$. Assumptions (ref) (i) and (ref) (ii) imply that the unobserved means $\mathbb{E}[y|T^*,T,z]$ are fully determined by $(\alpha_0, \alpha_1)$ given the observed means $\mathbb{E}[y|T,z]$. The question is whether it is possible, given the observed distribution $F_{tk}$, to construct $F^1_{tk}$ and $F^{0}_{tk}$ with the required values for $\mathbb{E}[y|T^*,T,z]$ such that $F_{tk} = r_{tk} F^{1}_{tk} + (1 - r_{tk}) F^{0}_{tk}$ for all combinations $(t,k)$. If not, then $(\alpha_0, \alpha_1)$ does not belong to the identified set. Our proof provides necessary and sufficient conditions for such a mixture to exist at a given point $(\alpha_0, \alpha_1)$. We can then appeal to the reasoning from Theorem (ref) to complete the argument. By ruling out values for $\alpha_0$ and $\alpha_1$, Theorem (ref) restricts $\beta$ via Lemma (ref). While these restrictions can be very informative, they do not yield point identification.

corUnder Assumptions (ref) and (ref) the identified set for $\beta(\mathbf{x})$ contains both the IV estimand $\mbox{Cov}(y,z|\mathbf{x})/\mbox{Cov}(z,T|\mathbf{x})$ and the true coefficient $\beta(\mathbf{x})$.

Corollary (ref) follows by Lemma (ref) because $\alpha_0(\mathbf{x})=\alpha_1(\mathbf{x})=0$ always belongs to the sharp identified set from Theorem (ref). Non-differential measurement error cannot exclude the possibility that there is no mis-classification because in this case it is trivial to construct the required mixtures. Although we focus throughout this paper on the case of a binary instrument, one might wonder whether point identification can be achieved by increasing the support of $z$, perhaps along the lines of Lewbel. The answer turns out to be no. Suppose that we were to modify Assumptions (ref) and (ref) to hold for all values of $z$ in some discrete support set. By Lemma (ref), a binary instrument identifies $\beta(\mathbf{x})$ up to knowledge of the mis-classification probabilities $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$. It follows that any pair of values $(k,\ell)$ in the support set of $z$ identifies the same object. Accordingly, to identify $\beta(\mathbf{x})$ it is necessary and sufficient to identify the mis-classification probabilities. A binary instrument fails to identify these probabilities because we can never exclude the possibility of zero mis-classification. The same is true of a discrete $K$-valued instrument. Increasing the support of $z$ does, however, shrink the identified set by increasing the number of restrictions available: in this case Theorems (ref)--(ref) continue to apply replacing “$k=0,1$” with “for all $k$.”

Point Identification

The results of the preceding section establish that $\beta(\mathbf{x})$ is not point identified under Assumptions (ref) and (ref). In light of this, there are two possible ways to proceed: either one can report partial identification bounds based on our characterization of the sharp identified set from Theorem (ref), or one can attempt to impose stronger assumptions to obtain point identification. In this section we consider the second possibility. We begin by defining the following functions of the model parameters:

align[align omitted — 541 chars of source]

Now consider the following additional assumption:

assump$\mathbb{E}[\varepsilon^2|\mathbf{x},z] = \mathbb{E}[\varepsilon^2|\mathbf{x}]$

Assumption (ref) is a second moment version of the standard mean exclusion restriction for the instrument $z$ -- Assumption (ref) (iii). It requires that the conditional variance of the error term given the covariates $\mathbf{x}$ does not depend on $z$, but does not require homoskedasticity with respect to $\mathbf{x}, T^*$ or $T$. Assumption (ref) allows us to derive the following lemma:

lemUnder Assumptions (ref), (ref) and (ref), \[ \mbox{Cov}(y^2,z|\mathbf{x}) = 2\mbox{Cov}(yT,z|\mathbf{x}) \theta_1(\mathbf{x}) -\mbox{Cov}(T,z|\mathbf{x})\theta_2(\mathbf{x}) \] where $\theta_1(\mathbf{x})$ and $\theta_2(\mathbf{x})$ are defined in Equations (ref)--(ref).

Lemma (ref) identifies $\theta_1(\mathbf{x})$. Since $\mbox{Cov}(z,T|\mathbf{x}) \neq 0$ by Assumption (ref) (ii), we can solve for $\theta_2(\mathbf{x})$ in terms of observables only, using Lemma (ref). Given knowledge of $\theta_1(\mathbf{x})$, we can solve Equation (ref) for the difference of mis-classification rates so long as $\beta(\mathbf{x}) \neq 0$.

corUnder Assumptions (ref)--(ref) and (ref), $\alpha_1(\mathbf{x}) - \alpha_0(\mathbf{x})$ is identified so long as $\beta(\mathbf{x}) \neq 0$.

Corollary (ref) identifies the difference of mis-classification error rates. Hence, under one-sided mis-classification, $\alpha_0(\mathbf{x}) = 0$ or $\alpha_1(\mathbf{x}) = 0$, augmenting our baseline Assumptions (ref)--(ref) with Assumption (ref) suffices to identify $\beta(\mathbf{x})$. Notice that $\beta(\mathbf{x})=0$ if and only if $\theta_1(\mathbf{x}) = 0$. Thus, $\beta(\mathbf{x})$ is still identified in the case where Corollary (ref) fails to apply.

Assumption (ref) does not suffice to identify $\beta(\mathbf{x})$ without a priori restrictions on the mis-classification error rates. To achieve identification in the general case, we impose the following additional conditions:

assump\begin{enumerate}[(i)] • $\mathbb{E}[\varepsilon^2|\mathbf{x},z,T^*,T] = \mathbb{E}[\varepsilon^2|\mathbf{x},z, T^*]$$\mathbb{E}[\varepsilon^3|\mathbf{x},z] = \mathbb{E}[\varepsilon^3|\mathbf{x}]$ \end{enumerate}

Assumption (ref) (i) is a second moment version of the non-differential measurement error assumption, Assumption (ref) (iii). It requires that, given knowledge of $(\mathbf{x}, T^*,z)$, $T$ provides no additional information about the variance of the error term. Note that Assumption (ref) (i) does not require homoskedasticity of $\varepsilon$ with respect to $\mathbf{x}$ or $T^*$. Assumption (ref) (ii) is a third moment version of Assumption (ref). It requires that the conditional third moment of the error term given $\mathbf{x}$ does not depend on $z$. This condition neither requires nor excludes skewness in the error term conditional on covariates: it merely states that the skewness is unaffected by the instrument. While Assumptions (ref) and (ref) may appear somewhat unusual, they are implied by the more intuitive independence conditions $\varepsilon \rotatebox[origin=c]{90}{$\models$} z |\mathbf{x}$ and $\varepsilon \rotatebox[origin=c]{90}{$\models$} T | (\mathbf{x}, T^*, z)$. Although $\mathbb{E}[\varepsilon|\mathbf{x},z]=0$ and $\mathbb{E}[\varepsilon|\mathbf{x},z,T^*,T] = \mathbb{E}[\varepsilon|\mathbf{x},z,T^*]$ are technically weaker than assuming full independence, we would be somewhat dubious of any supposed “natural experiment” that purportedly satisfied mean exclusion but not independence. Indeed, as discussed by ImbensRubin1997, an instrument satisfying mean exclusion but not independence could become invalid if the outcome variable were transformed, for example by taking logs. As it is not uncommon for applied papers to report results in both logs and levels angrist1990, our view is that researchers implicitly assume more than mean exclusion in typical applications of instrumental variables. Analogous reasoning applies to the non-differential measurement error assumption.

Assumption (ref) allows us to derive the following Lemma which, combined with Lemma (ref), leads to point identification:

lemUnder Assumptions (ref)--(ref) and (ref)--(ref), \[ \mbox{Cov}(y^3,z|\mathbf{x}) = 3 \mbox{Cov}(y^2T,z|\mathbf{x}) \theta_1(\mathbf{x}) -3\mbox{Cov}(yT,z|\mathbf{x}) \theta_2(\mathbf{x}) + \mbox{Cov}(T,z|\mathbf{x}) \theta_3(\mathbf{x}) \] where $\theta_1(\mathbf{x}),\theta_2(\mathbf{x})$ and $\theta_3(\mathbf{x})$ are defined in Equations (ref)--(ref).
thmUnder Assumptions (ref)--(ref) and (ref)--(ref), $\beta(\mathbf{x})$ is identified. If $\mathbf{\beta}(\mathbf{x}) \neq 0$, then $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$ are likewise identified.

Lemmas (ref)--(ref) yield a linear system of three equations in $\theta_1(\mathbf{x}), \theta_2(\mathbf{x})$ and $\theta_3(\mathbf{x})$. Under Assumption (ref) (ii), the system has a unique solution so $\theta_1(\mathbf{x}), \theta_2(\mathbf{x})$ and $\theta_3(\mathbf{x})$ are identified. The proof of Theorem (ref) shows that, so long as $\beta(\mathbf{x})\neq 0$, Equations (ref)--(ref) can be solved for $\beta(\mathbf{x})$, $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$. In particular, using steps from the proof of Theorem (ref) \[ \beta(\mathbf{x}) = \mbox{sign}\big[ \theta_1(\mathbf{x}) \big] \sqrt{3 \big[ \theta_2(\mathbf{x})/\theta_1(\mathbf{x}) \big]^2 - 2\big[\theta_3(\mathbf{x})/\theta_1(\mathbf{x})\big]}. \] If we relax Assumption (ref) (ii) and assume $\alpha_0(\mathbf{x}) + \alpha_1(\mathbf{x}) \neq 1$ only, $\beta(\mathbf{x})$ is only identified up to sign: in this case the sign of $\theta_1(\mathbf{x})$ need not equal that of $\beta(\mathbf{x})$.

Estimation and Inference

We now briefly outline how the identification results from Section (ref) can be used to estimate and carry out statistical inference for the parameters of interest: $\big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}), \beta(\mathbf{x})\big)$. Lemmas (ref)--(ref) yield a system of linear moment equations in the reduced form parameters $\boldsymbol{\theta}'(\mathbf{x}) = \big(\theta_1(\mathbf{x}), \theta_2(\mathbf{x}),\theta_3(\mathbf{x})\big)$. Defining a vector of intercepts $\boldsymbol{\kappa}'(\mathbf{x}) = \big(\kappa_1(\mathbf{x}), \kappa_2(\mathbf{x}), \kappa_3(\mathbf{x})\big)$, and a vector of observables $\mathbf{w}' = (T, y, yT, y^2, y^2 T, y^3)$, we can write this system as

align[align omitted — 614 chars of source]

Using Equations (ref)--(ref), we can re-write $\mathbf{\Psi}$ as a function of $\big(\alpha_0(\mathbf{x}), \alpha_1(\mathbf{x}), \beta(\mathbf{x})\big)$, leaving us with a just-identified, non-parametric conditional moment problem. Because the conditioning variables in Equation (ref) are the same as the arguments of the unknown functions $(\alpha_0, \alpha_1, \beta)$, this problem fits within the framework of Lewbel2007, permitting straightforward estimation and inference via a local GMM procedure. If $\beta(\mathbf{x})$ is close to zero, however, this procedure can perform poorly; in this case the moment conditions from Equations (ref), are only weakly informative about $\alpha_0(\mathbf{x})$ and $\alpha_1(\mathbf{x})$. An earlier version of this paper DiTragliaGarciaWP2017 discusses this problem in more detail and provides a solution based on generalized moment selection AndrewsSoares that combines the moment inequalities implied by our partial identification results from Section (ref) with the moment equalities from Equation (ref).

Conclusion

This paper has studied identification and inference for a mis-classified, binary, endogenous regressor in an additively separable model using a discrete instrumental variable. We have shown that the only existing identification result for this model is incorrect, and gone on to derive the sharp identified set under standard first-moment assumptions from the literature. Strengthening these assumptions to hold for second and third moments, we have established point identification for the effect of interest. An interesting extension of the results presented above would be to consider the case of discrete regressors that take on more than two values.