EconBase
← Back to paper

Partial Identification of Binary Choice Models with Misreported Outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,071 characters · 19 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification of Binary Choice Models with Misreported Outcomes

\sloppy

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long\global\long \global\long\global\long\global\long\global\long\global\long \global\long \global\long\def\mc#1{\mathscr{#1}}

\global\long\global\long \global\long \global\long

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }

\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }

\global\long\def\turd#1{\frac{#1}{3}}

\global\long \global\long\def\sand#1{\left\lceil #1\right\vert }

\global\long\def\wich#1{\left\vert #1\right\rfloor }

\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }

\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\inprod#1{\left\langle #1\right\rangle }

\global\long\def\ol#1{\overline{#1}}

\global\long\def\ul#1{#1}

\global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}}

\global\long \global\long \global\long \global\long \global\long {6pt} {6pt}

abstractThis paper provides partial identification of various binary choice models with misreported dependent variables. We propose two distinct approaches by exploiting different instrumental variables respectively. In the first approach, the instrument is assumed to only affect the true dependent variable but not misreporting probabilities. The second approach uses an instrument that influences misreporting probabilities monotonically while having no effect on the true dependent variable. Moreover, we derive identification results under additional restrictions on misreporting, including bounded/monotone misreporting probabilities. We use simulations to demonstrate the robust performance of our approaches, and apply the method to study educational attainment. Keywords: Misreporting, Binary Choice Models, Instrumental Variables, Partial Identification, Exclusion Restriction JEL classification: C01, C14, C26.

Introduction

This paper provides partial identification of various binary choice models when the dependent variable is potentially misreported. Binary choice models have been widely used in empirical applications such as analyzing participation in social programs, employment status, and educational attainment. However, many applications rely on survey data such as the Survey of Income and Program Participation (SIPP) and the Current Population Survey (CPS), where the binary outcome variable may be misreported or misclassified in survey data due to interviewer or respondent errors. The problem of misreporting is well documented and several studies show that the misreporting probabilities can be significant. For example, meyer2020 show that the probability of misreporting participation in a food stamp program can range from 23% in the SIPP to 50% in the CPS.

Numerous studies have examined the bias introduced by misreporting across various econometric models \citep*{aigner1973, bollinger1997, kane1999, davern2009, nguimkeu2019}. Regarding a binary choice model, meyer2017 show that misreporting in the binary dependent variable can lead to significant biases in parametric estimators. While misreporting might be pervasive in some widely used datasets, these datasets may remain valuable sources of information, often with no appropriate substitute. It is therefore vital to investigate what can still be learned from the contaminated data.

Tackling misreporting issues can be challenging. Firstly, misreporting in a binary variable involves non-classical measurement errors, as the measurement error is always negatively correlated with the true outcome. Moreover, misreporting stem from unobserved incentives among respondents; for example, people who benefit from a food stamp program may conceal their participation out of a sense of shame. As such, misreporting probabilities can depend on observed characteristics in an unknown way.

This paper introduces two different approaches to identify various binary choice models with a potentially misreported dependent variable, including parametric, semiparametric, and panel binary choice models. With potential misreporting in the dependent variable, conventional approaches for binary choice models do not apply, as the true dependent variable is not observed, and the conditional expectation of the true dependent variable is not identified. Our identification strategy derives bounds for the conditional expectation of the true outcome given covariates, by exploiting variation in different instruments. Given the bounds, we derive partial identification for binary choice models by characterizing conditional moment inequalities.

In our first approach, the instrument is assumed to only affect the true outcome while not influencing misreporting probabilities. In the example of program participation, such as job training program, this instrument can be randomly assigned eligibility for the program. This instrument affects the true participation of the program, but it is unlikely to affect misreporting, given its random nature. The second approach uses an instrument that only affects misreporting probabilities monotonically, but does not influence the true outcome. Examples of such a variable could include interview-relevant variables such as an interviewer's evaluation of respondents' accuracy or interview styles in survey data, including in-person, phone, or email interviews. Individuals are more likely to provide truthful responses during in-person interviews than during email interviews.

We derive partial identification results using each instrument. The strategy involves deriving bounds on misreporting probabilities through the variation in each instrument, thereby establishing bounds on the conditional expectation of the true outcome. Furthermore, we explore the identifying power of each instrument with additional restrictions on misreporting, including one-sided misreporting, bounded misreporting probabilities, and monotone misreporting probabilities. These restrictions can potentially provide more informative bounds for misreporting probabilities and thus tighten the bounds for the conditional expectation of the true outcome.

Our approach accommodates various binary choice models and allows for flexible misreporting processes. The identification strategy is applicable to binary choice models under full independence and distributional assumptions, or median independence restrictions, as well as panel models with conditional homogeneity conditions. Additionally, we allow for heterogeneous misreporting probabilities, which can depend on observed characteristics arbitrarily. This flexibility has practical value; for example, bollinger1997, bollinger2001 demonstrate that misreporting probabilities are correlated with participants' characteristics such as demographic characteristics and family income. Furthermore, we do not assume a parametric model for the misreporting process (such as a linear index model) and permit arbitrary dependence between the true outcome and the misreporting process.

We characterize partial identification for various binary choice models using conditional moment inequalities. Through simulations, we evaluate the finite sample performance of our method, taking the semiparametric model as an illustrative example. For comparison, we also implement the maximum likelihood estimation method studied in hausman1998, which assumes constant misreporting probabilities and distributional assumptions. The results demonstrate the robustness of our approaches with respect to heterogeneous misreporting probabilities and parametric assumptions. As an empirical illustration, we apply our method to study a binary choice model of educational attainment using the data from National Longitudinal Surveys in 1976.

In the extension, we examine the joint identifying power of the two instrumental variables together. The two instruments jointly provide a new channel for identification, introducing additional restrictions on misreporting probabilities. This result leads to more informative bounds on the conditional expectation of the true outcome compared to intersecting the bounds obtained by using each instrument separately.

Related Literature

This paper directly contributes to the line of literature on binary choice models with misreported dependent variables. See chen2011nonlinear for a comprehensive survey of nonlinear models with measurement errors, including binary choice models. The literature on binary choice models dates back to chamberlain1980, manski1985, and manski1987. When the dependent variable is misreported, hausman1998 proposes a modified maximum likelihood estimator for binary choice models to correct potential misreporting, and abrevaya1999semiparametric explores a more general linear index model. These studies assume homogeneous misreporting probabilities, where misreporting rates are constant regardless of values of covariates. bollinger1997 allows misreporting rates to depend on covariates in a Probit model, imposing parametric assumptions for both the true binary choice model and misreporting processes. lewbel2000 studies semiparametric identification of binary choice models by using a continuous instrumental variable that only affects the true outcome but is independent of misreporting rates. In a more recent paper, meyer2017 proposes different parametric estimators, relying on a parametric model for misreporting processes or the availability of validation data. Our paper allows for heterogeneous misreporting probabilities, which can depend on covariates in an arbitrary way. Furthermore, we explore the identifying power of discrete instruments (along with additional restrictions on misreporting) and the results are valid even if the instrument only take two values.

Our work also relates to a large body of literature on various models with misreported regressors. mahajan2003misclassified investigates misclassified regressors in binary choice models. Several studies explore regression models with misreported regressors under homogeneous misreporting probabilities, including aigner1973, bollinger1996, frazis2003. mahajan2006, lewbel2007, ditraglia2019 allow misreporting probabilities to be covariates dependent and achieve identification by using a binary instrument for the true regressor. chen2005measurement, chen2006identification, and chen2008nonparametric provide nonparametric identification without instrumental variables. They achieve identification by exploiting auxiliary data, two samples, and higher-order moments, respectively. Additionally, hu2008 and molinari2008 explore misclassification in a general discrete regressor, and hu2008instrumental and hu2022identification extend the study to a continuous regressor with nonclassical measurement errors. These papers investigate different frameworks with misreported regressors, and thus their assumptions and identification approaches/results differ from those in our paper. In a more closely related study, nguimkeu2019 uses two instruments jointly for identification--one for the true regressor and the other for misreporting. They obtain point identification under one-sided misreporting and parametric structures on misreporting, while our paper allows for two-sided misreporting and nonparametric structures on misreporting processes.

The above literature focuses on homogeneous effects of the regressor given covariates in a regression setting. Numerous papers study heterogeneous treatment effects with a misreported treatment such as kreider2007, kreider2012, battistin2014, calvi2017, and ura2018. Their approaches either exploit a repeated measurement, auxiliary administrative data to restrict misreporting errors, or an instrumental variable related to the true treatment. This literature also studies different framework and requires different assumptions to identify heterogeneous treatment effects from our paper.\footnote{Our approach can be potentially applied to study heterogeneous treatment effects, while it still requires substantial work to explore how to combine our approach with additional assumptions on the instrument for identification. } Additionally, these works exploit the instrument for treatment, while we also explore the identifying power of the instrument for misreporting. More differently, horowitz1995identification, kreider2008inferring, and kreider2011identification study identification with corrupted data under a mixing framework on data errors.

The rest of this paper is organized as follows. Section (ref) presents various binary choice models and identification approaches using two different instrument along with additional assumptions on misreporting. Section (ref) characterizes conditional moment inequalities for model parameters, and Section (ref) examines the finite sample performance via simulations. Section (ref) studies the application of educational attainment. Section (ref) explores an extension. We conclude with Section (ref).

Model and Identification

The analysis studies the identification of various binary choice models with potential misreporting (or misclassification) in the binary dependent variable. Let $Y_i^{*}\in\{0 ,1\}$ denote the true binary dependent variable, and $Y_i\in \{0, 1\}$ denote the observed variable which may be subject to misreporting. Let $X_i\in \mathcal{X}$ denotes a vector of observed covariate, which is relevant to the true variable $Y_i^{*}$ and can also affect misreporting probabilities.

When there is potential misreporting in the variable $Y_i^{*}$, the standard identification results do not apply, as the true variable $Y_i^{*}$ is not observed and the conditional probability $p^{*}(x):=\Pr(Y_i^{*}=1 \mid x)$ is not identified. Our identification strategy is to establish bounds $[L(x), U(x)]$ for the true conditional choice probability using observed variables $(X_i, Y_i)$ and exploiting the availability of various instruments. The bounds $[L(x), U(x)]$ will depend only on observed variables, so they are identified from data. Based on the bounds, we can characterize partial identification for several binary choice models.

Before introducing the identification approach, we first present various binary choice models to illustrate how the bounds on $p^{*}(x)$ can be exploited to derive partial identification results in various models.

example[Parametric Binary Choice Model] Consider the following model for the true dependent variable $Y_i^{*}$: \[ Y_i^{*} = \mathbf{\mathbbm1}\{\epsilon_i \leq X_i'\beta_0\}, \] where $\epsilon_i$ is independent of $X_i$ and follow a known distribution: $\epsilon_i \mid X \sim F_{\epsilon}(\cdot)$.\footnote{Common examples for parametric binary choice models include Probit and Logit models.} Under this structure, the true conditional choice probability is given as \[ p^{*}(x)= F_{\epsilon}(x'\beta_0). \] We are interested in identifying and estimating the parameter $\beta_0$. Given the derived bounds $p^{*}(x)\in [L(x), U(x)]$ in Section (ref), the identifying condition for the true parameter $\beta_0$ is characterized by the following conditional moment inequality: for any $x$, \[ L(x) \leq F_{\epsilon}(x'\beta_0) \leq U(x). \]
example[Semiparametric Binary Choice Model] The true dependent variable $Y_i^{*}$ is given as: \[ Y_i^{*} = \mathbf{\mathbbm1}\{\epsilon_i \leq X_i'\beta_0\}, \] where $\epsilon_i$ satisfies the mean independence assumption $E[\epsilon_i \mid X_i]=0$, while the distribution of $\epsilon_i$ is left unknown. When there is no misreporting in the variable $Y_i^{*}$, manski1985 derives the following identifying condition for $\beta_0$: \[ \begin{aligned} x'\beta_0 \geq 0 &\Longrightarrow p^{*}(x) - 0.5 \geq 0, \\ x'\beta_0 \leq 0 &\Longrightarrow p^{*}(x) - 0.5 \leq 0. \\ \end{aligned} \] However, the probability $p^{*}(x)$ is no longer identified when the true outcome $Y_i^{*}$ is not observed. Our approach dervies bounds for the conditional choice probability $p^{*}(x)\in [L(x), U(x)]$, yielding the following identifying condition: \[ \begin{aligned} x'\beta_0 \geq 0\Longrightarrow p^{*}(x)-0.5\geq 0 \Longrightarrow U(x)-0.5\geq 0, \\ x'\beta_0 \leq 0 \Longrightarrow p^{*}(x)-0.5\leq 0 \Longrightarrow L(x)-0.5\leq 0. \end{aligned} \]
example[Panel Binary Choice Model] Consider the following binary choice model for $Y_{it}^{*}$ with panel data structure: \[ Y_{it}^{*} = \mathbf{\mathbbm1}\{\epsilon_{it}+\alpha_i \leq X_{it}'\beta_0\}, \] where $\epsilon_{it}$ satisfies the conditional stationarity (homogeneity) assumption in manski1987: $\epsilon_{it}\mid \alpha_i, X_{is}, X_{it} \sim \epsilon_{is}\mid \alpha_i, X_{is}, X_{it}$. Similar to Example (ref), given the bounds on $p^{*}_t(x):=\Pr(Y_{it}^{*}=1 \mid X_{ist}=x)\in [L_t(x), U_t(x)]$ where $X_{ist}:=(X_{is}, X_{it})$, it has the following implication: \[ \begin{aligned} (x_t-x_s)'\beta_0 \geq 0 &\Longrightarrow p^{*}_t(x) \geq p^{*}_s(x) \Longrightarrow U_t(x)\geq L_s(x), \\ (x_t-x_s)'\beta_0 \leq 0 &\Longrightarrow p^{*}_t(x) \leq p^{*}_s(x) \Longrightarrow L_t(x)\leq U_s(x). \\ \end{aligned} \]

Although our paper mainly focuses on binary choice models, the proposed approach can be potentially applied to estimate treatment effects with misreported treatment.

example[Local Average Treatment Effects] Suppose that we are interested in estimating the causal effects of the true treatment $T^{*}\in \{0, 1\}$ on the outcome $Y$. The true treatment $T^{*}$ can be endogenous, so a binary instrument $Z\in\{0, 1\}$ is used to address endogeneity. Under the assumptions on instrument $Z$ introduced in imbens1994identification, the local average treatment effect (LATE) is given as \[ LATE =\frac{E[Y\mid Z=1]-E[Y\mid Z=0]} {E[T^{*} \mid Z=1]-E[T^{*}\mid Z=0]}. \] We consider that the true treatment $T^{*}$ is not observed, but instead we only observe a reported treatment $T\in\{0, 1\}$ which could be subject to misreporting. Our approach can bound the true conditional probability $p^{*}(z):= E[T^{*}\mid Z=z] \in [L(z), U(z)]$, and thus can also bound LATE.\footnote{The heterogeneous treatment effects framework introduces additional assumptions on instrument $Z$ with $(T^*, Y)$ to address endogeneity. Our approach only imposes assumptions between instruments $(Z, W)$ and variables $(T^{*}, T)$ (excluding $Y$) to address misreporting issues. Therefore, under this framework, we need to combine all these assumptions jointly to identify treatment effects with misreported treatment. }

Identification

We now present our identification approaches to establish bounds on the true conditional probability $p^{*}(x)=\Pr(Y_i^{*}=1\mid x)$. Let $p(x)=\Pr(Y_i=1\mid x)$ denote the reported probability of $Y_i=1$ given $X_i=x$, which is identified from the data. The reported probability $p(x)$ depends on two components: the true probability and misreporting probabilities. Therefore, it is essential to distinguish between these two components to identify the true probability from the reported probability.

We introduce two different approaches to identify the true probability $p^{*}(x)$ by exploiting different exclusion restrictions. In the first approach, we use a discrete instrument $Z_i\in\mathcal{Z}:=\{z_1, z_2,..., z_k\}$ that only affects the true probability $p^{*}(x)$ but does not affect misreporting probabilities. The other approach uses a discrete instrument $W_i\in \mathcal{W}:=\{w_1, w_2,..., w_l\}$ that only affects misreporting probabilities but not the true probability $p^{*}(x)$. We first study the identifying power of each individual instrument and then discuss how the two instruments can jointly identify the true probability $p^{*}(x)$ in the extension.

For simplicity of notation, we suppress subscript $i$ for random variables in the following analysis. The following graph describes the relationship among all variables.

figure[figure omitted — 547 chars of source]

Figure (ref) summarizes our model and main identification strategies. The objective is to learn the effects of covariate $X$ on the true binary outcome $Y^{*}$, but we only observe the reported outcome $Y$ which can be possibly misreported. We study the identifying power of two different instruments respectively: the first one is instrument $Z$ that only affects the true outcome $Y^{*}$ and the other is instrument $W$ that only affects the reported outcome $Y$ by affecting misreporting probabilities. Furthermore, we also study identification results of each instrument with additional assumptions on misreporting in Section (ref).

Instrument $Z$

This section studies the identifying power of instrument $Z$ that only affects the true choice probability but does not influence misreporting probabilities. Instrument $Z$ affects the true variable $Y^*$ directly so it is a component of the covariate vector $X$. Therefore, we divide covariate $X$ into two parts: instrument $Z$ and the remaining covariates denoted as $\tilde{X}=X\setminus Z\in \tilde{\mathcal{X} }$.

Next we state some assumptions on instrument $Z$.

ass[Exclusion] For any $x\in \mathcal{X}, y\in\{0,1\}$, \begin{equation*} \Pr(Y=1-y\mid Y^{*}=y, x)=\Pr(Y=1-y\mid Y^{*}=y, \tilde{x}). \end{equation*}

The exclusion restriction requires instrument $Z$ to be independent of misreporting process, so it only affects the reported probability by shifting the true probability. When the true outcome is participation in social programs, such as a job training program, one example of instrument $Z$ could be randomly assigned eligibility for the program. The random assignment will affect true participation but not influence misreporting, given its random nature. Studies, e.g., mahajan2006, ura2018, and ditraglia2019, provide more examples of this instrument in various applications.

The next assumption is about the extent of misreporting probabilities.

ass[Degree of Misreporting] For any $ \tilde{x}\in \tilde{\mathcal{X}}$, \begin{equation*} \Pr(Y=0\mid Y^{*}=1, \tilde{x})+\Pr(Y=1\mid Y^{*}=0, \tilde{x}) \leq 1. \end{equation*}

Assumption (ref) is about the degree of misreporting, ensuring that the reported data is informative for the true probability. It requires that the misreporting errors are not too large so that the sum of two-sided misreporting probabilities is smaller than one. This assumption is consistent with empirical evidence in multiple studies such as meyer2009 and meyer2020, which document that one-sided misreporting probability is usually less than 50% in survey data. Moreover, our identification analysis only requires the degree of misreporting to be known, and the analysis is applicable when the sum of two-sided misreporting rates is larger than one.

ass[Boundary Condition] The reported choice probability $p(x)$ satisfies that: $\sup \limits_{z\in \mathcal{Z}}p( \tilde{x}, z)>0$ and $\inf\limits_{z\in \mathcal{Z}}p( \tilde{x}, z)<1$ for any $\tilde{x}\in \tilde{\mathcal{X} }$.

Assumption (ref) is a boundary condition for the reported probability. This assumption is relatively weak, requiring that the supremum of the reported probability is bounded away from zero and the infimum of the reported probability is bounded away from one. It can be satisfied when instrument $Z$ strictly affects the reported probability and takes on at least two values.

Under the above assumptions, we are ready to establish identification results for the true probability $p^{*}(x)$.

propUnder Assumptions (ref)-(ref), the sharp bounds for $p^{*}(x)$ are characterized as $p^{*}(x)\in [L_{1}(x), U_{1}(x) ]$ for any $x\in \mathcal{X}$, where \begin{equation*} L_{1}(x)=\frac{p(x)- p\mkern-2mu\mkern2mu _z( \tilde{x})}{1- p\mkern-2mu\mkern2mu _z( \tilde{x})}, \qquad U_{1}(x)=\frac{p(x)}{\bar{p}_z( \tilde{x})}. \end{equation*} where $\underline{p\mkern-2mu}\mkern2mu _z( \tilde{x}):= \inf\limits_{z\in \mathcal{Z}} p( \tilde{x}, z) $ and $\bar{p}_z( \tilde{x}):=\sup \limits_{z\in \mathcal{Z}} p( \tilde{x}, z)$.

Proposition (ref) characterizes partial identification for the true probability $p^{*}(x)$ using variation in instrument $Z$. It also establishes sharpness of the results, showings that the bounds are the best possible given the assumptions and data. According to the definition of the bounds, we know that $L_1(x), U_1(x)\in [0, 1]$. The identifying power of instrument $Z$ depends on the range of the reported probability $p(\tilde{x}, z)$ as one varies $z$: tighter bounds for the true probability are achieved with a larger range of the reported probability. A larger variation in the reported probability can yield smaller bounds for misreporting probabilities, leading to tighter bounds on the true probability.

When the conditional reported probability can vary from zero to one when instrument $Z$ changes, we can infer that there is no misreporting and achieve point identification as $p^{*}(x)=p(x)$. The intuition is as follows. When the reported probability is one (or zero), there are two possibilities: one is that the true conditional probability is one, and individuals with $Y^{*}=1$ all report the truth (or misreport); the other is that the true probability is zero while individuals with $Y^{*}=0$ all misreport (or report the truth). The possibility that everyone misreports can be rejected by Assumption (ref), which requires that the sum of two-sided misreporting probabilities to be smaller than one. Therefore, we can conclude that there is no misreporting, and point identification for the true probability is obtained.

Instrument $W$

This section introduces an alternative method for identifying the true probability using instrument $W$. This instrument is assumed to only affect misreporting probabilities but not the true probability. In addition to covariate $X$ in the true binary choice model, instrument $W$ is an additional variable related to misreporting processes but excluded from the true binary choice model. Our approach does not impose any parametric models for misreporting processes, allowing instrument $W$ to affect misreporting probabilities nonparametrically.

The complete set of observed variables are $(X, W, Y)$ in this section. Let $p_W(x, w)=\Pr(Y=1\mid x, w)$ denote the reported probability conditional on $(X, W)=(x, w)$. The subscript $W$ in the function $p_W$ is used to distinguish it from the function $p(x)=\Pr(Y=1\mid x)$.

The following presents assumptions on instrument $W\in \mathcal{W}$.

ass[Exclusion] \begin{equation*} \Pr(Y^{*}=1\mid x, w)=\Pr(Y^{*}=1\mid x)=p^{*}(x), \end{equation*} for any $ x\in \mathcal{X}$ and $w\in\mathcal{W}$.

Similar to Assumption (ref) for instrument $Z$, Assumption (ref) states a different exclusion restriction, requiring instrument $W$ to be independent of the true dependent variable $Y^*$. Examples for instrument $W$ could include interview-related variables, such as interviewersÕ assessments of respondentsÕ accuracy or different interview styles like phone interviews and in-person interviews. These variables are unlikely to affect the true dependent variable but are related to respondentsÕ probabilities of reporting the truth. These variables are unlikely to affect the true dependent variable, but are relevant to respondents' probabilities of reporting the truth.

ass[Monotonicity] For any $x\in\mathcal{X}$, the following conditions hold for any $y\in\{0, 1\}$ and $w_1>w_2\in \mathcal{W}$, \begin{equation*} \begin{aligned} \Pr(Y=1-y\mid Y^{*}=y, x, w_1) \leq \Pr(Y=1-y\mid Y^{*}=y, x, w_2). \\ \end{aligned} \end{equation*}

Assumption (ref) is the monotonicity condition for instrument $W$: the misreporting probabilities are weakly decreasing with respect to the instrument. For instance, when the instrument is interviewers' evaluations of respondents' accuracy, it is natural that misreporting probabilities are smaller with higher evaluations. In the case of interview styles, individuals are more likely to report the truth during an in-person interview than a phone interview. It is also worth noting that Assumption (ref) is a weak monotonicity condition, as it only requires monotonicity for the average probability but allows for potential violations in certain populations.

ass\begin{enumerate}[(1)] • Degree of Misreporting: for any $x\in \mathcal{X}, w\in \mathcal{W}$, \begin{equation*} \Pr(Y=1\mid Y^{*}=0, x, w)+\Pr(Y=0\mid Y^{*}=1, x, w)\leq 1. \end{equation*} • Boundary Condition: the reported probability $p_W(x, w)$ is bounded away from zero and one: $0<p_W(x, w)<1$ for any $x\in \mathcal{X}, w\in \mathcal{W}$. \end{enumerate}

Assumption (ref) is similar to Assumptions (ref)-(ref) for instrument $Z$. The difference is that all probabilities are also conditional on the additional variable $W$ since the misreporting probabilities and the reported probability depend on instrument $W$ in this scenario.

Under the above assumptions, the next proposition establishes identification results for the true probability $p^{*}(x)$.

propUnder Assumptions (ref)-(ref), the true probability $p^{*}(x)$ can be bounded as $p^{*}(x)\in [L_2(x), U_2(x)]$ for any $x\in \mathcal{X}$, where \begin{equation*} L_2(x)= \sup_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)- p\mkern-2mu\mkern2mu _w(x, w)}{1- p\mkern-2mu\mkern2mu _w(x, w) } \right\}, \quad U_2(x)= \inf_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)}{ \bar{p}_w(x, w) } \right\}, \end{equation*} where $\underline{p\mkern-2mu}\mkern2mu _w(x, w):= \inf\limits_{\tilde{w} \leq w } p_W(x, \tilde{w})$ and $\bar{p}_w(x, w):= \sup\limits_{\tilde{w}\leq w } p_W(x, \tilde{w})$. Moreover, the above bounds are sharp when instrument $W$ is binary.

Proposition (ref) derives partial identification by using a distinct instrument that only monotonically influences misreporting probabilities but not the true probability. Furthermore, it demonstrates that this result exhausts all possible information when the instrument $W$ is binary. The bounds, as per their definitions, satisfy that $L_2(x), U_2(x)\in[0, 1]$. Additionally, the results in Proposition (ref) imply that $U_2(x)\geq L_2(x)$ for all $x$, providing testable implications for our assumptions.

Proposition (ref) mainly exploits the exclusion and monotonicity of instrument $W$. Under the monotonicity condition, the misreporting probabilities at $w$ can be bounded above by all upper bounds of misreporting probabilities evaluated at smaller values $\tilde{w}\leq w$. The identifying power of monotonicity is shown within the bracket of the bounds $(L_2(x), U_2(x))$. The exclusion restriction can help tighten the bounds for the true probability $p^{*}(x)$ by intersecting all bounds derived from any value of instrument $W$, as demonstrated outside the bracket of bounds.

Additional Restrictions on Misreporting

Our analysis can be also combined with other additional information on misreporting probabilities to further tighten the bounds for the true probabilities $p^{*}(x)$. Section (ref)-(ref) studies one-sided misreporting, bounded misreporting probabilities, as well as monotone misreporting probabilities.

One-sided Misreporting

This section studies identification of the true probability $p^{*}(x)$ under one-sided misreporting. One-sided misreporting refers to where only one group with the true outcome $Y^{*}=y$ misreport, while the other group with $Y^{*}=1-y$ always report the truth. This assumption has practical applications and has been used in previous studies, such as nguimkeu2019. In their paper, they investigate the participation in the food stamp program and provide evidence for a small overreporting probability. The assumption regarding which group has no misreporting depends on the specific application, and we present results for both cases.

The next proposition provides bounds for the true probability $p^{*}(x)$ under one-sided misreporting by using one of the two instruments $(Z, W)$ respectively.

prop(1) Under Assumptions (ref)-(ref), the sharp bounds for the true probability $p^{*}(x)$ are characterized as \begin{equation*} p^{*}(x)\in \left\{ \begin{aligned} &\left[p(x), U_1(x) \right] \qquad when \ \Pr(Y=1\mid Y^{*}=0, \tilde{x})=0, \\ &\left[L_1(x), p(x) \right] \qquad when \ \Pr(Y=0\mid Y^{*}=1, \tilde{x})=0. \end{aligned} \right. \end{equation*} (2) Under Assumptions (ref)-(ref), the sharp bounds for the true probability $p^{*}(x)$ are characterized as \begin{equation*} p^{*}(x)\in \left\{ \begin{aligned} &\left[\sup_{w\in \mathcal{W}} p_W(x, w), U_2(x) \right] \qquad when \ \Pr(Y=1\mid Y^{*}=0, x, w)=0, \\ &\left[L_2(x), \inf_{w\in \mathcal{W}} p_W(x, w) \right] \qquad when \ \Pr(Y=0\mid Y^{*}=1, x, w)=0. \end{aligned} \right. \end{equation*}

Proposition (ref) provides sharp identification results for $p^{*}(x)$ under different scenarios of one-sided misreporting, using one of the two instruments, respectively. The one-sided misreporting assumption ensures accurate data from one group and establishes the misreporting probability to be zero for this group. In comparison to the results in Proposition (ref) and (ref), which allow for two-sided misreporting, Proposition (ref) consistently provides tighter bounds on $p^{*}(x)$ with either larger lower bound or smaller upper bound.\footnote{When instrument $W$ is available, the one-sided misreporting assumption implies the monotonicity of $p_W(x, w)$ in $w$, and the upper bound $U_2(x)$ becomes one. This implication can serve as a testable implication for the one-sided misreporting assumption. }

Bounded Misreporting Probabilities

This section considers that the misreporting probabilities are bounded by known numbers. This restriction may come from information such as previous studies on misreporting, auxiliary administrative data, or theoretical models that provide insights into the potential range of misreporting probabilities.

assThe misreporting probabilities satisfy the following condition: for $y \in \{ 1, 2\}$ and any $(x, w)$, \[ \Pr(Y=1-y \mid Y^{*}=y, x, w) \leq \bar{\alpha}_y, \] where $\bar{\alpha}_y \in [0, 1]$ is a known constant.

Assumption (ref) may stem from empirical results in other research or auxiliary information, such as administrative data from other samples. For simplicity, Assumption (ref) adopts uniform bounds for misreporting probabilities regardless of values of $(x, w)$. Our approach can be also adjusted to allow $\bar{\alpha}_y$ to depend on $(x, w)$. When instrument $W$ is not available, then the above conditional misreporting probability can be adjusted by only conditioning on the covariate $x$.

prop(1) Under Assumptions (ref)-(ref) $\&$ (ref), the true probability $p^{*}(x)$ can be bounded as \begin{equation*} p^{*}(x)\in \left[\frac{p(x)- \min \left\{ p\mkern-2mu\mkern2mu ( \tilde{x}), \bar{\alpha}_0 \right\}}{1- \min \left\{ p\mkern-2mu\mkern2mu ( \tilde{x}), \bar{\alpha}_0 \right\} }, \frac{p(x)}{\max \left\{ \bar{p}( \tilde{x}), 1-\bar{\alpha}_1 \right\} } \right]; \end{equation*} (2) Under Assumptions (ref)-(ref) $\&$ (ref), the true probability $p^{*}(x)$ can be bounded as \begin{equation*} \begin{aligned} p^{*}(x)\in \left[ \sup_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)- \min \left\{ p\mkern-2mu\mkern2mu _w(x, w), \bar{\alpha}_0 \right\} }{1- \min \left\{ p\mkern-2mu\mkern2mu _w(x, w), \bar{\alpha}_0 \right\} } \right\}, \inf_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)}{\max\left \{ \bar{p}_w(x, w), 1-\bar{\alpha}_1 \right\} } \right\} \right]. \end{aligned} \end{equation*}

Proposition (ref) derives partial identification for $p^{*}(x)$through additional bounds on misreporting probabilities, with the identifying power depending on the value of the bounds $\bar{\alpha}_y$. A smaller value of $\bar{\alpha}_y$ results in tighter bounds for $p^{*}(x)$, and $\bar{\alpha}_y=0$ corresponds to no misreporting. It is also possible that this additional restriction may not provide any information if the value of the bound $\bar{\alpha}_y$ is larger than that derived by using each instrument.

Monotone Misreporting Probabilities

This section explores monotone misreporting probabilities, where the misreporting probability of one group is smaller than that of the other group. The direction of monotonicity depends on specific applications. To illustrate, we study the following monotonicity on misreporting probabilities.

assThe misreporting probabilities satisfy the following condition: for any $(x, w)$, \[ \Pr(Y=1 \mid Y^{*}=0, x, w) \leq \Pr(Y=0 \mid Y^*=1, x, w). \]

Assumption (ref) says that the misreporting probability for people with $Y^*=0$ is smaller than those with $Y^*=1$. This assumption is applicable for scenarios where $Y^*$ represents participation in social assistance programs or whether one is smoking or not. It is documented in the literature that people who did not participate in social program are more likely to report the truth than people who participated. The direction of the monotonicity can be reversed in some applications, such as $Y^*$ represents educational attainment. Our identification methods can accommodate both cases, where the direction of the monotonicity is required to be known.

prop(1) Under Assumptions (ref)-(ref) $\&$ (ref), the true probability $p^{*}(x)$ can be bounded as \begin{equation*} p^{*}(x)\in \left[\frac{p(x)- \min\left\{ p\mkern-2mu\mkern2mu _z( \tilde{x}), 1-\bar{p}_z( \tilde{x}) \right\} }{1- \min\left\{ p\mkern-2mu\mkern2mu _z( \tilde{x}), 1-\bar{p}_z( \tilde{x}) \right\} }, \frac{p(x)}{ \bar{p}( \tilde{x}) } \right]; \end{equation*} (2) Under Assumptions (ref)-(ref) $\&$ (ref), the true probability $p^{*}(x)$ can be bounded as \begin{equation*} \begin{aligned} p^{*}(x)\in \left[ \sup_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)- \min \left\{ p\mkern-2mu\mkern2mu _w(x, w), 1-\bar{p}_w(x, w) \right\} }{1- \min \left\{ p\mkern-2mu\mkern2mu _w(x, w), 1-\bar{p}_w(x, w) \right\} } \right\}, \inf_{w\in \mathcal{W}} \left\{ \frac{p_W(x, w)}{ \bar{p}_w(x, w) } \right\} \right]. \end{aligned} \end{equation*}

Proposition (ref) shows that Assumption (ref) can further increase the lower bound for $p^{*}(x)$ with each instrument. The idea is that under the monotone misreporting rates assumption, the misreporting probability for group $Y^*=0$ can also be bounded by the upper bound of the misreporting probability for the other group $Y^*=1$, yielding more informative results for the true probability $p^*(x)$. Symmetrically, when assuming a smaller misreporting rate for group $Y^*=1$ compared to $Y^*=0$, the upper bound for $p^{*}(x)$ can be tightened.

Conditional Moment Inequalities

Based on previous identification analysis, this section demonstrates how to characterize identification for various binary choice models using conditional moment inequalities. We focus on the two binary choice models with cross-sectional data in Examples (ref) and (ref) to illustrate the idea, and the analysis can be applied to panel binary choice models.\footnote{For panel models, we can follow the same identification strategy to derive bounds for the true probability $p^{*}_t(x):=E[Y_t^*\mid X_{st}= x]$ at each period $t$. The main distinction is that all variables and results will be indexed by time period $t$.}

Parametric Binary Choice Model

As shown in Example (ref), given the parametric structure of the error term $\epsilon$, i.e., $\epsilon\mid X \sim F_{\epsilon}(\cdot)$, the model parameter $\beta_0$ is characterized by the following restriction:

equation[equation omitted — 80 chars of source]

where $k \in\{1, 2\}$ depends on the availability of instruments.

When the instrument $Z$ is available ($k=1$), plugging into the definition of $(L_1(x), U_1(x))$, the identifying restriction in (ref) is equivalent to the following conditional moment inequality: \[E[g_{z, par}(Y, X, \beta_0)\mid x]\geq 0, \] where $g_{z, par}(Y, X, \beta_0))$ is defined as \[g_{z, par}(Y, X, \beta_0)):=\left\{

aligned&Y-F_{\epsilon}(X'\beta_0) \bar{p}_Z(\tilde{X}) \\ & F_{\epsilon}(X'\beta_0) + p\mkern-2mu\mkern2mu _Z(\tilde{X}) (1-F_{\epsilon}(X'\beta_0) ) -Y.

\right. \]

Similarly, when the instrument $W$ is available ($k=2$), condition (ref) is equivalent to the restriction: \[ E[g_{w, par}(Y, X, W, \beta_0)\mid x, w]\geq 0,\] where $g_{w, par}(Y, X, W, \beta_0)$ is defined as \[g_{w, par}(Y, X, W, \beta_0):=\left\{

aligned&Y-F_{\epsilon}(X'\beta_0) \bar{p}_W(X, W) \\ & F_{\epsilon}(X'\beta_0) + p\mkern-2mu\mkern2mu _W(X, W) (1-F_{\epsilon}(X'\beta_0) ) -Y.

\right. \]

Semiparametric Binary Choice Model

Under the semiparametric framework in Example (ref), the identifying restriction of the model parameter $\beta_0$ is characterized as

equation[equation omitted — 167 chars of source]

for $k \in \{1, 2\}$. To characterize conditional moment inequalities, we first conduct monotone transformations of the bounds $(L_k(x), U_k(x))$ by multiplying them by their respective (positive) denominators. This monotone transformation preserves the sign of those bounds so it does not affect the identification results.

When instrument $Z$ is available with $k=1$, the following condition holds by multiplying the denominators of $(L_1(x), U_1(x))$:

equation*[equation* omitted — 237 chars of source]

Then, restriction (ref) is equivalent to the following conditional moment restriction: \[ E[g_{z, semi}(Y, X, \beta_0)\mid x]\geq 0, \] where $g_{z, semi}(Y, X, \beta_0)$ is defined as \[g_{z, semi}(Y, X, \beta_0):=\left\{

aligned&X'\beta_{0}\mathbbm{1}\{X'\beta_{0}\geq 0\}(Y-0.5 \bar{p}_Z(\tilde{X})), \\ & X'\beta_{0}\mathbbm{1}\{X'\beta_{0} \leq 0 \}(Y-0.5 p\mkern-2mu\mkern2mu _Z(\tilde{X})-0.5).

\right. \]

When the instrument $W$ is available with $k=2$, we can conduct a similar monotone transformation of $(L_2(x), U_2(x))$, yielding the following relationship:

equation*[equation* omitted — 264 chars of source]

Based on the above relationship, restriction (ref) is equivalent to the following conditional moment restriction: \[ E[g_{w, semi}(Y, X, W, \beta_0)\mid x, w]\geq 0, \] where $g_{w, semi}(Y, X, W, \beta_0)$ is defined as \[g_{w, semi}(Y, X, W, \beta_0):=\left\{

aligned&X'\beta_{0}\mathbbm{1}\{X'\beta_{0}\geq 0\}(Y - 0.5 \bar{p}_W(X, W)), \\ & X'\beta_{0}\mathbbm{1}\{X'\beta_{0} \leq 0 \}(Y - 0.5 p\mkern-2mu\mkern2mu _W(X, W)-0.5).

\right. \]

For both parametric and semiparametric models, we characterize partial identification of the model parameter using conditional moment inequalities. Then, we can adopt established methods from the literature developed for general conditional moment inequalities to conduct estimation and inference, such as chernozhukov2007, andrews2013, and chernozhukov2013.\footnote{The conditional moment inequalities contain nuisance parameters, e.g., $\bar{p}_Z( \tilde{x})$ and $\underline{p\mkern-2mu}\mkern2mu _Z( \tilde{x})$, that can be consistently estimated. The inference methods in the literature, such as andrews2013 Section 8, allow for preliminary consistent estimation of nuisance parameters.}

Simulation Study

This section examines the finite sample performance of our identification approaches via Monte Carlo simulations. We focus on the semiparametric binary choice model presented in Example (ref), which allows for flexible misreporting process and does not require distributional assumption. To better evaluate our approach, we also implement the maximum likelihood estimation approach proposed in hausman1998 for comparison, referred to as the HAS approach. Their method accounts for potential misreporting, but assumes distributional assumption on error term and homogeneous (constant) misreporting probabilities. The simulation results demonstrate the robustness of our approach concerning heterogeneous misreporting probabilities and parametric assumptions.

For our approach, we follow chernozhukov2007 and andrews2013 to estimate the identified set based on conditional moment inequalities. We first transform conditional moment inequalities into unconditional moment inequalities using indicator functions of hypercubes in the space of covariates as instrumental functions.\footnote{See andrews2013 for more choices and discussions of instrumental functions.} The number of hypercubes is $\{30, 40, 50\}$ for the sample size $n\in \{1000, 2000, 4000\}$. Then the estimated identified set can be computed based on the criterion function method in chernozhukov2007.

Next, we present the performance of our approach and the HAS method in hausman1998 using different instruments.

Instrument $Z$

This section studies the identification results using instrument $Z$. The DGP is described as follows. Instrument $Z$ is uniformly distributed over the set $\{-1, -0.5, 0, 0.5, 1\}$, $\tilde{X}$ follow a uniform distribution over the interval $[-1, 1]$, and the full covariate $X$ is given as $X=[1; \tilde{X}; Z]$. The true parameter $\beta_0=[1; 1.5; -1.5]$ and the true outcome is generated by $Y^{*}= \mathbbm{1}\left\{X'\beta_0 \geq \epsilon \right\}$. We study two specifications of the error term $\epsilon$: a standard normal distribution $\mathcal{N}(0, 1)$ and a Cauchy distribution $Cauchy(0, 0.5)$.

The reported outcome $Y$ is given by $Y =M_{1}\cdot Y^*+(1-M_{0})\cdot (1-Y^*)$, where $M_y\in \{0, 1\}$ denotes the reporting variable and $M_y=0$ represents misreporting for any $y\in \{0, 1\}$. We allow for heterogeneous misreporting probabilities, depending on the value of covariate $\tilde{X}$: \[ \Pr(M_1=0\mid \tilde{x})=0.1 - 0.1 \tilde{x}, \quad \Pr(M_0=0\mid \tilde{x})=0.3 + 0.1 \tilde{x}. \]

In this DGP, it is clear that Assumptions (ref) and (ref) on the exclusion and degrees of misreporting are satisfied. The sample size is $n \in \{500, 1000, 2000\}$ and repetition number is $B = 300$.

To compare the performance of the two methods, we report the root mean-squared error (rMSE) and median of absolute deviation (MAD) for the lower bound $\hat{\beta}^{l}$ and upper bound $\hat{\beta}^u$ in this paper, along with the parametric estimator $ \hat{\beta}_{HAS}$, which assumes constant misreporting probabilities and a normal distribution of the error term. Let $\beta^k$ denote the $k$th element of the parameter $\beta$. We normalize the first element of the parameter to one for our method: $\beta_0^1=1$.

table[table omitted — 1,113 chars of source]
table[table omitted — 1,112 chars of source]

Tables (ref) and (ref) display the performance of the two methods under different specifications of the error term $\epsilon$ and different sample sizes. The results illustrate that our approach uniformly performs well across various error term specifications. When the distribution of $\epsilon$ is correctly specified, the HAS method shows reasonable performance but does not necessarily outperform our method, as the constant misreporting probabilities assumption still remains misspecified. Moreover, the HAS estimator exhibits significant bias when the distributional assumption is also misspecified (under the Cauchy design), while our approach has robust performance under different designs.

Instrument $W$

This section examines the finite sample performance of the identified set using instrument $W$. The DGP is described as follows. Covariate $X$ is given by $X=(1; \tilde{X})$, where $\tilde{X}$ follows a uniform distribution over $[-1,1]$. The true outcome $Y^{*}$ is generated by $Y^{*}=\mathbbm{1}\{\epsilon\leq X'\beta_0 \}$, where the true parameter is $\beta_0=[1; 1.5]$. Similarly, we consider two specifications of error term $\epsilon$: a standard normal distribution $\mathcal{N}(0, 1)$ and a Cauchy distribution $Cauchy(0, 0.5)$.

Instrument $W$ is uniformly distributed over the set $\{1, 2, 3, 4, 5\}$. The reported outcome $Y$ is given by $Y =M_{1}\cdot Y^*+(1-M_{0})\cdot (1-Y^*)$, where $M_y\in \{0, 1\}$ and $M_y=0$ denotes misreporting. The misreporting probability $\Pr(M_y =0\mid \tilde{x}, w)$ depends on covariate $\tilde{X}$ and instrument $W$ as follows: \[ \Pr(M_1=0\mid \tilde{x}, w)=0.1 - 0.1 \tilde{x}, \quad \Pr(M_0=0\mid \tilde{x}, w)=\frac{1}{1+0.3w^2}. \]

In this specification, it can be verified that the monotonicity condition in Assumption (ref) on instrument $W$ and the degree of misreporting probabilities in Assumption (ref) are satisfied. The sample size is $n \in \{500, 1000, 2000\}$ and repetition number is $B = 300$.

table[table omitted — 1,126 chars of source]

Table (ref) displays the performance of $\beta^2_0$ under different sample sizes and specifications of $\epsilon$. The results show that our approach consistently outperforms the HAS approach across all specifications. The HAS method exhibits a significant bias even when the distribution of $\epsilon$ is correctly specified, as the range of misreporting probabilities (under different values of $(\tilde{x}, w)$) is very large in this DGP and the constant misreporting probabilities assumption is seriously misspecified. The bias of HAS method becomes larger when the distributional assumption is also misspecified under the Cauchy design. In summary, the simulation results demonstrate the robust performance of our two approaches concerning flexible misreporting processes and distributional assumptions.

Empirical Illustration

As an empirical illustration, we apply our methods to analyze educational attainment using a binary choice model with potential misreporting. The dataset we use is drawn from the National Longitudinal Surveys in 1976 (NLSY76), which is also used in card1995 to estimate returns to education. This survey data contains 3613 individuals' self-reported information including educational experiences and family backgrounds. The objective is to explore how people's characteristics affect the probability of them attaining a college degree. However, there may be misreporting in self-reports of educational attainment in this data, which could severely bias the estimation results.

In this application, the reported outcome $Y$ is whether an individual reports attending a college which may be subject to misreporting. Instrument $Z$ is whether an individual grew up near a four-year college (college proximity). This instrument affects people's true decision of attending college, but may not affect their misreporting behaviors. We also include two other covariates in the binary choice model: parents' average education $X_1$ and whether an individual is black $X_2$. The following table shows the summary statistics of all variables.

table[table omitted — 394 chars of source]

We adopt the semiparametric binary choice model with instrument $Z$ to study how individuals' observed characteristics affect the likelihood of attending a college.\footnote{The interview-related information is only available in a “restricted-use" version of NLSY79, so there is no instrument $W$ for this application.} In addition to accommodating flexible misreporting processes, our method is robust to distributional assumptions. The full vector of covariate is $X=[1; X_1; X_2; Z]$, and the corresponding coefficient is denoted as $\beta_0=[\beta_0^0, \beta_0^1, \beta_0^2, \beta_0^3]$. For the semiparametric binary choice model, the coefficient $\beta_0$ can be only identified up to a constant. The coefficient of instrument $Z$ is normalized to one since people tend to be more likely to attend a college when they live closer to a college. For comparison, we also display the results of the HAS approach, which is described in Section (ref).

table[table omitted — 427 chars of source]

Table (ref) presents the estimation results for the coefficients in the binary choice model. The method in this paper shows a positive sign of parents' education and a negative sign of being black for educational attainment. However, the HAS method shows negative signs for all coefficients, which seems inconsistent with economic intuition. Parents' education and living closer to a college are likely to increase the chance of attending a college instead of decreasing the chance. The results show that misspecifications in misreporting processes or distributional assumptions may lead to opposite signs of the coefficients.

Extension: Two Instruments

Section (ref) and (ref) provide bounds for the true conditional probability $p^{*}(x)$ when there is only one instrument available. This section studies the joint identifying power of the two instruments $(Z, W)$. The observed variables are $(X, W, Y)=(\tilde{X}, Z, W, Y)$ when two instruments are available. We adjust previous assumptions in Section (ref) and (ref) slightly to accommodate the availability of the two instruments.

ass\begin{enumerate}[(1)] • Exclusion: for any $ \tilde{x} \in \tilde{\mathcal{X}}, z\in \mathcal{Z}, w\in \mathcal{W}$, and $y\in\{0, 1\}$, \begin{equation*} \begin{aligned} \Pr(Y^{*}=1\mid x, w)&=\Pr(Y^{*}=1\mid x)=p^{*}(x), \\ \Pr(Y=1-y\mid Y^{*}=y, x, w)&=\Pr(Y=1-y\mid Y^{*}=y, \tilde{x}, w). \end{aligned} \end{equation*} • Degree of misreporting: for any $ \tilde{x} \in \tilde{\mathcal{X}}$ and $w\in \mathcal{W}$, \begin{equation*} \Pr(Y=0\mid Y^{*}=1, \tilde{x}, w)+\Pr(Y=0\mid Y^{*}=1, \tilde{x}, w)\leq 1. \end{equation*} • Monotonicity $\&$ Relevance: for any $ \tilde{x} \in \tilde{\mathcal{X}}$, $w_1>w_2 \in \mathcal{W}$, and $y\in\{0, 1\}$, \begin{equation*} \begin{aligned} \Pr(Y=1-y\mid Y^{*}=y, \tilde{x}, w_1) \leq \Pr(Y=1-y\mid Y^{*}=y, \tilde{x}, w_2), \\ \end{aligned} \end{equation*} and there exists $k\in\{0, 1\}$ such that the above inequality is strict. • Relevance: for any $ \tilde{x} \in \tilde{\mathcal{X}}$, there exists $z_1\neq z_2\in \mathcal{Z}$ such that $ p^{*}( \tilde{x}, z_1)\neq p^{*}( \tilde{x}, z_2)$. \end{enumerate}

Assumption (ref) summarizes all assumptions for the two instruments $(Z, W)$ in previous sections. Assumption (1) states exclusion restrictions for the two instruments, which requires that instrument $Z$ does not affect misreporting probabilities and instrument $W$ does not affect the true probability. Assumption (2) requires the sum of the two-sided misreporting probabilities to be smaller than one. Assumption (3) adds one relevance restriction for instrument $W$ so that $W$ at least affects the misreporting probability for one group $Y^{*}=y$ strictly. Assumption (4) is the relevance condition for instrument $Z$, but the direction of how the true probability is affected by instrument $Z$ is not restricted so that instrument $Z$ can either increase or decrease the true probability. The relevance condition of instrument $Z$ can guarantee that the supremum and infimum of the reported probability over $Z$ are bounded away from one and zero respectively. Therefore, the boundary condition in previous sections is no longer needed in this section.

Under the above assumptions, we can use joint variation in the two instruments to derive bounds for misreporting probabilities and the true probability. The joint variation can bound misreporting probabilities through a new channel, thereby providing more informative results than simply taking intersections over the bounds derived using each instrument separately.

Let $w_{m}$ denote the maximum value of instrument $W$. Next, we establish bounds for the misreporting probabilities evaluated at $w_{m}$ by using the two instruments jointly. Under the monotonicity condition of instrument $W$ in Assumption (ref) (iii), the misreporting probability at $w_{m}$ is the smallest misreporting probability. The bounds for misreporting probabilities evaluated at other values of $W$ can be established similarly, which will lead to the same identification result for the true probability. Therefore, we focus on the results for the smallest misreporting probabilities.

The next lemma derives bounds on misreporting probabilities evaluated at $W=w_{m}$.

lemmaUnder Assumption (ref), the misreporting probability $\Pr(Y=1-y \mid Y^{*}=y, \tilde{x}, w_{m})$ can be bounded as: $[0, U_{\alpha_{y}}( \tilde{x}, w_{m})]$ for any $ \tilde{x}\in \tilde{\mathcal{X}}$ and $y\in \{0, 1\}$, where \begin{equation*} \begin{aligned} &U_{\alpha_{1}}( \tilde{x}, w_{m})=1-\sup_{z, w<w_m } \left\{ \frac{q_1( \tilde{x}, w_{m}, w)p_W( \tilde{x}, z_1, w)-p_W( \tilde{x}, z_1, w_m)}{q_1( \tilde{x}, w_{m}, w)-1}, p_W( \tilde{x}, z, w_{m}) \right\}, \\ &U_{\alpha_{0}}( \tilde{x}, w_{m})= \inf_{z, w<w_m } \left\{ \frac{q_1( \tilde{x},w_{m}, w)p_W( \tilde{x}, z_1, w)-p_W( \tilde{x}, z_1, w_{m})}{q_1( \tilde{x}, w_{m}, w)-1}, p_W( \tilde{x}, z, w_{m})\right\},\\ &q_1( \tilde{x}, w_{m}, w)=\frac{p_W( \tilde{x}, z_1, w_{m})-p_W( \tilde{x}, z_2, w_{m}) }{p_W( \tilde{x}, z_1, w)-p_W( \tilde{x}, z_2, w) }. \end{aligned} \end{equation*}

Lemma (ref) characterizes the lower and upper bound for the misreporting probabilities evaluated at $W=w_m$ by using two instruments jointly. The lower bounds for the smallest misreporting probabilities are zero, since we cannot rule out the possibility of no misreporting.

The upper bounds demonstrate the joint identifying power of the two instruments. From the definition of $U_{\alpha_{y}}( \tilde{x}, w_m)$, it uses variation from both instruments $(Z, W)$. The term $p_W( \tilde{x}, z, w_m)$ shows the identifying power of instrument $Z$, and the other term involving $q_1( \tilde{x}, w_m, w)$ uses joint information of the two instruments, providing more information compared to using only instrument $W$. The main idea is that the joint variation in the two instruments imposes additional restrictions between misreporting probabilities at different values of $W$, which can further bound the misreporting probabilities.

Given bounds on misreporting probabilities, the next proposition characterizes the identification result for the true probability $p^{*}(x)$.

propUnder Assumption (ref), the true conditional choice probability $p^{*}(x)$ is bounded as $p^{*}(x)=[L_3(x), U_3(x)]$ for any $x\in \mathcal{X}$, where \begin{equation*} L_3(x)= \frac{p_W(x, w_m)-U_{\alpha_0}( \tilde{x}, w_m)}{1-U_{\alpha_0}( \tilde{x}, w_m)}, \qquad U_3(x)= \frac{p_W(x, w_m)}{ 1-U_{\alpha_1}( \tilde{x}, w_m)}. \end{equation*} And the above bounds are sharp when instrument $W$ is binary.

Proposition (ref) establishes bounds for the true probability by using two instruments jointly, and these bounds have exhausted all possible information from assumptions and observed data when instrument $W$ is binary. From the definition of the bounds, the lower bound $L_3(x)$ decreases with respect to the bound $U_{\alpha_0}( \tilde{x}, w_m)$ on misreporting probabilities, and the upper bound $U_3(x)$ increases with respect to $U_{\alpha_1}( \tilde{x}, w_m)$. Therefore, a smaller bound $U_{\alpha_y}( \tilde{x}, w_m)$ for misreporting probabilities would imply tighter bounds for the true conditional probability. As discussed, the upper bound $U_{\alpha_y}( \tilde{x}, w_m)$ on misreporting probabilities shown in Lemma (ref) would be smaller than the one by only using one instrument. Therefore, Proposition (ref) derives more informative bounds for $p^{*}(x)$ by using the two instruments jointly.

Conclusion

This paper provides partial identification of various binary choice models with misreported dependent variables, including parametric, semiparametric, and panel binary choice models. We introduce two distinct approaches by exploiting the availability of different instrumental variables, respectively. Moreover, our approach can accommodate additional restrictions on misreporting, such as one-sided misreporting, bounded misreporting probabilities, and monotone misreporting probabilities.

Our approach allows for flexible misreporting processes in the sense that we do not impose any parametric model for misreporting processes and allow for heterogeneous misreporting probabilities. It would be interesting to explore how additional parametric structures on misreporting can tighten the bounds for the conditional expectation of the true dependent variable. Furthermore, we focus on binary choice models in this paper, while the approach for handling misreporting may be applied more broadly. It still requires substantial future work to investigate how the method can be applied in other models with potential misreporting, such as ordered and multinomial choice models with misreported dependent variables.