EconBase
← Back to paper

Nonclassical Measurement Error in the Outcome Variable

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

48,346 characters · 8 sections · 62 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonclassical Measurement Error in the Outcome Variable

abstractWe study a semi-/nonparametric regression model with a general form of nonclassical measurement error in the outcome variable. We show equivalence of this model to a generalized regression model. Our main identifying assumptions are a special regressor type restriction and monotonicity in the nonlinear relationship between the observed and unobserved true outcome. Nonparametric identification is then obtained under a normalization of the unknown link function, which is a natural extension of the classical measurement error case. We propose a novel sieve rank estimator for the regression function and establish its rate of convergence. In Monte Carlo simulations, we find that our estimator corrects for biases induced by nonclassical measurement error and provides numerically stable results. We apply our method to analyze belief formation of stock market expectations with survey data from the German Socio-Economic Panel (SOEP) and find evidence for nonclassical measurement error in subjective belief data.

\vskip .5cm

center[center omitted — 199 chars of source]

Introduction

In empirical research, measurement error is a recurring issue. In recent years, much attention has been given to various forms of measurement error in the covariates of econometric models, whereas measurement error of the dependent variable is mostly ignored. In many economic environments, measurement error of the dependent variable may be driven (in a nonlinear fashion) by the underlying variable. This nonclassical measurement error implies biased estimation results if not accounted for.

This paper is concerned with semi-/nonparametric regression models where the dependent variable of interest $Y^*$ is generally not observed and only a possibly error-contaminated measurement $Y$ is observable. Specifically, $Y^*$ satisfies

align[align omitted — 41 chars of source]

where the unknown function $g$ is of interest given observed covariates $X$ and unobservables $U$. We study the nonclassical measurement error case where $\mathop{{\mathbf E}\null}\nolimits[Y|Y^*, X] \neq Y^*$. Hence, the regression function $g$ does in general not coincide with conditional expectations of observable variables and we cannot impose $g(x)= \mathop{{\mathbf E}\null}\nolimits[Y|X=x]$.

Nonparametric identification of our model relies on the availability of covariates which do not affect the measurement error directly. We impose such type of exclusion restriction on a subset $Z$ of the vector $X=(Z,W)$, where $W$ are additional controls. Under a monotonicity condition on the measurement error mechanism, we show in this paper that model (ref) can be reformulated as a generalized regression model of the form

align*[align* omitted — 67 chars of source]

where $H(\cdot,w)$ is a nonlinear, monotonic function for $w$ in the support of $W$. Identification of the function $g$, up to strictly monotonic transformations, immediately follows, which allows us to infer on economically relevant quantities such as the direction and shape of partial effects.

Under scale and location normalization of the unknown link function $H$, nonparametric identification of the regression function $g$ is obtained. We highlight that normalization of the link function $H$ is equivalent to imposing mild shape restrictions on the measurement error mechanism. Additionally, our normalization conditions on the link function do not only naturally extend the classical measurement case but are also satisfied if there is a range of $Y^*$ where measurement error is classical. Our nonparametric identification results build thus on intuitive assumptions without relying on high-level assumptions such as completeness, see hu2008.

We consider a sieve, rank-based minimum distance estimator and establish its asymptotic properties. We derive the rate of convergence in $L^2$ sense of our estimator. We find that the sieve rank estimator generally suffers from ill-posedness in the convergence rate as the rank-based criterion function is not continuous in the usual $L^2$-norm. We develop the theory for the case where $W$ is discrete and provide an extension to allow continuous controls $W$ using kernel weights in the appendix of this paper.

We analyze the performance of the estimator in a Monte Carlo simulation study and in an empirical application using survey data. We apply our estimator to study belief formation with subjective belief data from the German Socio-Economic Panel innovation sample (SOEP-IS). Subjective belief data is known to be plagued by substantial measurement error and it is in general hard to justify that the measurement error is classical and thus not sensitive to the underlying true individual belief. We study the impact of an exogenous display of historic stock market returns provided to survey respondents prior to eliciting their belief on future returns. Applying our method, we find a monotonic and concave relationship between the historic information and stated beliefs indicating that individuals acknowledge the given information conservatively.

\paragraph{Literature} Our work ties into the literature on measurement error in observable variables of econometric models. The literature on measurement error in covariates is extensive, whereas measurement error in the outcome variable has received much less attention. For a review of models with errors in covariates, see e.g. Chen_ME_Rev and Schennach_rev. CHT_ME2005 develop a general way of accounting for measurement error in any variable of a class of semiparametric models once auxiliary data, e.g. from validation samples is available. However, this is hardly the case in most practical applications. Models focusing on nonclassical measurement error in the outcome side are rare. Chapter 3 of AbrevHausman99 considers a semiparametric model with a more simplistic measurement error mechanism. HoderleinWinter2010 and hoderlein2015 develop structural models of response error in surveys due to imperfect recall and derive testable implications for econometric analyses. The latter paper focuses on the role of rounding in individual reporting behavior which is also a more specific form of nonclassical measurement error.

NadaiLewbel allows for classical measurement error in the outcome variable that is correlated with an error in covariates. abrevaya2004 consider classical measurement error of the dependent variable in a transformation model. Given we have a precise idea on the form of measurement error, a sizeable literature is usually available providing different strategies for identification. For instance a special case of nonclassical measurement error is selective non-response in the outcome variable, see e.g. 2010Hault or breunigEndSel18 and references therein. A non-nested form of nonclassical measurement error are Berkson-type errors, see Berkson and Schennach_rev.

Our identifying assumptions lead us to the literature on generalized regression models as introduced in han1987 or the class of nonlinear index models in matzkin2007. See also the model studied in JachoLewbel. Estimation of such models often proceeds by rank-based estimation strategies, see han1987, CavSherman, khan2001, shin2010 and abrevaya2011 which all consider parametric regression models with the exception of matzkin1991 who studies a nonparametric model with additional shape restrictions on the link function. A recent contribution studying rank estimators in a high-dimensional setting is FanRank20. To the best of our knowledge, we are the first to study nonparametric M-estimation with rank-based criterion functions and to point out and illustrate the ill-posedness of the estimation problem. JureckovaBernoulli16 study a different class of rank estimators in the context of a parametric model with measurement error in both regressors and outcome. Their the outcome error may not be nonclassical as in our general notion but can at most depend on observable regressors.

The remainder of the paper is organized as follows. In Section (ref) we present our model setup and give a nonparametric identification result for features of the mean regression function when there is a form of nonclassical measurement error in the outcome variable. In Section (ref) we introduce a sieve estimator with a rank based criterion function and establish its convergence. In Section (ref) we analyze finite sample properties of the estimator in a Monte Carlo simulation study. Section (ref) contains an application of our method to belief formation of stock market expectations. Appendix (ref) provides an extension to weighted sieve rank estimation, when control variables are continuous. All proofs are postponed to the Appendix (ref).

Model Setup and Identification

We consider a nonparametric econometric model with measurement error in the outcome variable. The model we study is

equation[equation omitted — 47 chars of source]

where $Y^*$ is the scalar, outcome variable, $X$ is a $d_x$-dimensional vector of exogenous covariates, $U$ is a scalar error term, and $g$ a nonparametric function of interest. The outcome variable $Y^*$ is not observed by the researcher; only an error contaminated measurement $Y$ is available. We are primarily interested in the case where the error satisfies $\mathop{{\mathbf E}\null}\nolimits[U|X]=0$ and thus $g$ is the unknown conditional expectation function of $Y^*$ given $X$.

Throughout the paper, we assume that the regressors $X$ can be decomposed such that $X=(Z', W')'$, where $Z$ has no direct effect on the measurement error and $W$ are control variables. Also we introduce the notation $g_w(\cdot)\equiv g(\cdot, w)$ for the regression function evaluated at a fixed $w$ in the support of $W$. We now provide conditions, which allow for nonparametric identification of $g_w$ up a strictly monotonic transformation.

assA[Exclusion Restriction] The observed outcome $Y$ is conditionally mean independent of $Z$ given $Y^*$ and $W$, i.e., $\mathop{{\mathbf E}\null}\nolimits[Y|Y^*,Z, W]=\mathop{{\mathbf E}\null}\nolimits[Y|Y^*, W]$.

Assumption (ref) rules out that $Z$ has a direct effect on the measurement $Y$ in conditional expectations. Assumption (ref) is generally weaker than assuming that the conditional distribution of $Y$ given $(Y^*, Z,W)$ does not depend on $Z$, which restricts $Z$ to have no information on $Y$ that is not captured by $(Y^*,W)$. Analogues exclusion restrictions are commonly imposed in the literature on nonclassical measurement error in covariates. In hu2008, the distribution of the error-contaminated regressor is independent of instruments conditional on the latent regressor (see also Schennach_rev). Assumption (ref) is less restrictive than other exclusion restrictions found in the measurement error literature, see benmosh2017.

Conditions similar to Assumption (ref) can also be found in the literature on selective non-response, which is a special case of nonclassical measurement error in the outcome. Individuals either report the outcome truthfully (response indicator $D=1$) or not at all ($D=0$) so the observed outcome in this case is $Y=D Y^*$. See also Remark (ref) below. An identifying assumption in 2010Hault and breunigEndSel18 is that $D \perp \!\!\! \perp X \;|\; (Y^*, W)$, which is related to Assumption (ref).

In the following, we make use of the notation $h(Y^*, W)=\mathop{{\mathbf E}\null}\nolimits[Y|Y^*, W]$. Assumption (ref) implies the measurement error model

align*[align* omitted — 28 chars of source]

where $\mathop{{\mathbf E}\null}\nolimits[V|Y^*,W]=0$. Consequently, Assumption (ref) implies conditional mean independence of the measurement error $V$ given the regression error $U$, that is, $\mathop{{\mathbf E}\null}\nolimits[V | U]=0$.

assA[Monotonicity] For any $w\in\textsl{supp}(W)$, the function $h(\cdot, w)$ is weakly monotonic and non-constant over the support of $Y^*$.

Assumption (ref) imposes that the expected observed outcome $Y$ is monotonic in the latent outcome $Y^*$ given $W$. This is trivially satisfied when the measurement error is classical, i.e., when $h$ does not depend on $W$ and is the identity. A similar monotonicity condition has also been imposed in the measurement error model in AbrevHausman99.\footnote{In our notation AbrevHausman99 consider the error mechanism $Y=h(Y^*,V)$, with $\partial_y h(Y^*,V)> 0$, $\partial_v h(Y^*,V)> 0$ and $V \perp \!\!\! \perp (X, U)$. As we allow for heteroscedasticity in the measurement error model, condition $\partial_y h(Y^*,V)> 0$ may lead to one sided error restrictions.} Note that $h$ does not need to be strictly monotonic which allows to consider models with rounding error in the outcome, see hoderlein2015. We discuss the plausiblity of Assumption (ref) in the context of the application in Section (ref) in a setting with survey data.

assA[Conditional Exogeneity] The conditional independence restriction $Z \perp \!\!\! \perp U\; |\; W$ holds.

Assumption (ref) imposes a conditional independence restriction of $Z$ and the regression error $U$. This condition is also known as conditional exogeneity assumption following white2010. Independence assumptions can be restrictive, but are often required in the measurement error literature (see, e.g. Hausmanetal91, Schennach07, benmosh2017), or when accounting for endogeneity using control functions (see, e.g. Neweyetal99). We relax such restrictions by imposing independence to hold only conditional on control variables $W$. Similar conditions are often employed for identification in the econometrics literature, see e.g. CKK for nonparametric identification in a transformation model. Assumption (ref) also corresponds to the unconfoundedness assumption in the treatment effects literature and is also closely related to the special regressor assumption, see Lewbel_SpecReg for a review.

Next, we need the following set of regularity conditions. We introduce the notation $\textsl{supp}(V)$ for the support of a random vector $V$.

assAFor any $w\in \textsl{supp}(W)$: (i) the function $g_w$ is continuous; (ii) and any $z_1, z_2 \in \textsl{supp}(Z)$ such that $g_w(z_1) < g_w(z_2)$ there exists $u \in \textsl{supp}(U)$ satisfying $h(g_w(z_1)+u, w) < h(g_w(z_2)+u, w)$; (iii) there is at least one variable $Z_{(1)}$ such that $Z=(Z_{(1)},Z_{(-1)})$ with $f_{Z_{(1)}| Z_{(-1)},W}(z_1|z_{-1},w) > 0 $ for all $(z_1, z_{-1})\in\textsl{supp}(Z)$.

Assumption (ref) (ii) is a mild support condition on $U$ conditional on $W=w$. The unobservable $U$ must vary sufficiently to shift $g_w(Z)$ out of a flat region of $h$. The assumption is not required if $h$ is strictly monotonic in its first argument. Assumption (ref) (iii) requires $Z$ to contain at least one continuously distributed variable with sufficient variation. If $Z$ is scalar then Assumption (ref) (iii) may be replaced by $f_{Z|W}(z|w) > 0 $ for all $z\in\textsl{supp}(Z)$. This rules out the case of $Z$ being a discrete scalar variable.

Under the stated assumptions, now provide establish equivalence to the regression model (ref) to a generalized regression model specified by the link function $H(g_w(z), w)=\mathop{{\mathbf E}\null}\nolimits[h(g(z, W)+U, W) \;|\; W=w] $. Below, $\mathds{1}\{\cdot\}$ denotes the indicator function.

theoLet Assumptions (ref)--(ref) be satisfied, then for any $w \in \textsl{supp}(W)$ it holds \begin{align} \mathop{{\mathbf E}\null}\nolimits[Y|X=x] = H(g_w(z), w), \end{align} where $H(\cdot,w)$ is strictly monotonically increasing and $g_w(z)$ maximizes the function \begin{align} \mathcal Q(\phi,w)=\mathop{{\mathbf E}\null}\nolimits[Y_1\mathds{1}\{ \phi(X_1) > \phi(X_2)\} \;|\; W_1=W_2=w]. \end{align} In particular, the function $g_w(\cdot)$ is identified up to strictly increasing transformations.

The model (ref) falls into the class of generalized regression models studied by han1987, matzkin1991, and CavSherman. Further note that nonclassical measurement error implies heterogeneous biases for the marginal effects. When $\partial_z H(g_w(z), w)<1$ we obtain an attenuation bias for the marginal effect $\partial_z g_w(z)$ and when $\partial_z H(g_w(z), w)>1$ we get an augmentation bias for $\partial_z g_w(z)$.

Theorem (ref) implies identification of features of $g_w$ that are preserved under monotonic transformations. This includes the sign of partial effects, the ratio of two partial effects\footnote{Note that for $g(z_1, z_2)$ it holds that $\frac{\partial g}{\partial z_1}/\frac{\partial g}{\partial z_2} = \frac{\partial H(g)}{\partial z_1}/\frac{\partial H(g)}{\partial z_2}$ whenever these quantities and ratios are well-defined.} and properties such as quasi-concavity (-convexity) of the function. For the remainder of the paper we consider identification and estimation of $g_w$ in the point identified case.

We impose the following restriction on the model and the measurement error mechanism described by the function $H$.

assA(i) The function $g_w$ is additively separable such that there exists a decomposition $Z=(Z_1, Z_{-1})$ such that $g_w(Z)=m_w(Z_1) + l_w(Z_{-1})$ for some functions $m_w, l_w$. (ii) There exists $\{z_1, z_2\} \subset \textsl{supp}(Z)$ with $g_w(z_1)\neq g_w(z_2)$ and $\mathop{{\mathbf E}\null}\nolimits[Y|Z=z, W=w]=\mathop{{\mathbf E}\null}\nolimits[Y^*|Z=z, W=w ]$ for $z\in\{z_1, z_2\}$.

Assumption (ref) (i) imposes an additive separable structure on the regression function $g_w$. Following the identification statement in Theorem (ref), mere location and scale normalizations are not sufficient to point identify $g_w$. However, for any additive separable model this is the case, see also JachoLewbel. Assumption (ref) (ii) restricts the measurement error for at least to realizations of $Z$. Assumption (ref) (ii) is also in line with normalization requirements for identification under nonclassical measurement error. For instance, Assumption 5 of hu2008 requires some functional of the distribution of the measurement error conditional on the value of the true variable to be equal to the true variable itself, such as some quantile of $Y|Y^*=y^*$ to correspond to $y^*$.

Economic restrictions on the model can also be employed to sufficiently restrict the function space. We refer to the discussion in Sections 3.4 and 4.4 in matzkin2007 where several possible function spaces are discussed that can replace Assumption (ref)(i). This includes the spaces of functions that are homogeneous of degree one or so called “least-concave” functions, see also matzkin1994. matzkin2007 shows that imposing homogeneity of degree 1 and a location normalization is sufficient for Assumption (ref). Homogeneous functions are frequently encountered in microeconomics. Thus, in applications where the function $g$ has the structural interpretation of a production or cost function, homogeneity can be a reasonable restriction on the parameter space.

coroLet Assumptions (ref)-- (ref) (i) be satisfied, then the function $g_w$ is identified up to a location and scale normalization. If (ref) (ii) is additionally satisfied then the function $g_w$ is point identified.

Corollary (ref) establishes identification of the regression function under normalization imposed in Assumption (ref). The shape restrictions imposed in Assumption (ref) imply a normalization of the unknown, nonparametric link function $H$, in contrast to nonparametric generalized regression models, where normalization is typically imposed on the unknown function of interest.

We neither restrict the support of the observed outcome $Y$, nor require continuity in the function $h(\cdot, w)$. Thus, we can also cover cases where the observed outcome is categorical or has mass points. This likely occurs in survey data as respondents tend to provide rounded values. The following examples consider a generalization and special case of model ((ref)).

example[Control function approach] We can also motivate the presence of $W$ in Assumption (ref) as a control function. To this end we deviate for a moment from our previous notation and introduce the following triangular model \begin{align*} Y^*= & g(X)+U \\ X =& m(Z, \eta) \end{align*} where for simplicity $X$ is a one-dimensional endogenous covariate that may correlate with the model error $U$. The function $m$ is strictly monotonic in $\eta$ and $Z$ is an instrumental variable satisfying $ Z \perp \!\!\! \perp (U, \eta) $. Under additional regularity conditions, following ImbensNewey09 it holds that \begin{align*} X & \perp \!\!\! \perp U \;|\; W \quad with \\ W& = F_{X|Z}(X,Z)=F_\eta(\eta), \end{align*} where $F_V$ denotes the cummulative distribution function of a random variable $V$. As in Assumption (ref) we impose $\mathop{{\mathbf E}\null}\nolimits[Y | Y^*, Z, W]= \mathop{{\mathbf E}\null}\nolimits[Y | Y^*, W]$. Thus, following Theorem (ref), we obtain identification of the structural function $g$ up to a strictly monotonic transformation.
example[Selective Nonresponse] Consider a nonresponse model \begin{align*} Y &=D Y^* \\ D &= \phi(Y^*,W, V), \end{align*} for some unknown function $\phi$, where the response indicator $D\in\{0,1\}$ is always observed and $Y^*$ is only observed if $D=1$. This framework, where the response mechanism is mainly driven by the latent outcome $Y^*$ has been studied by 2010Hault and breunigEndSel18. As long as the conditional mean function $h(Y^*, W)=P(D=1|Y^*, W)Y^*$ is monotonic in its first argument, the model is in accordance to Assumption (ref). This holds e.g. when the conditional response probability function is monotonic and the support of $Y^*$ is bounded below\footnote{If $Y^*$ is bounded below, then $Y^*$ can be redefined such that without loss of generality $Y^* \geq 0$ and monotonicity of $h(Y^*, W)=P(D=1|Y^*, W)Y^*$ follows from taking the derivative.}. In this case, a completeness condition for nonparametric identification of the conditional selection probability $P(D=1|Y^*,W)$ (see 2010Hault and breunigEndSel18) via conditional moment restrictions is not required.

Estimation and Asymptotic Properties

In this section, we introduce a nonparametric sieve M-estimator with a simple, rank-based criterion function. For simplicity, we consider only the case where $W$ consists of discrete variables and defer the estimation with continuous $W$ to Appendix (ref).

The Sieve Rank Estimator

Our identification result builds on shape restrictions imposed on the measurement error mechanism, which imply identified moment conditions. Specifically, for a given $w$ we have from the identification statement in Theorem (ref) that the true $g_w$ maximizes the function

align*[align* omitted — 127 chars of source]

Based on this population criterion, we now consider a sieve rank estimator, which implicitly accounts for imposed shape restrictions required for identification.

We propose the following sieve rank estimator

align[align omitted — 275 chars of source]

for some $K=K(n)$ dimensional sieve space $\mathcal{G}_{K}$. Here, the dimension parameter $K$ grows slowly with sample size $n$. For the special case where $W$ is absent, the criterion reduces to

equation[equation omitted — 119 chars of source]

where the rank function is defined as $\text{Rank}(\phi(Z_i))=\sum_{j \neq i}^n \mathds{1}\{\phi(Z_{i}) > \phi(Z_{j})\}$. This is a nonparametric version of the criterion of CavSherman.

The specific choice of $\mathcal{G}_K$ hinges on the chosen normalization. Under a normalization of the link function $H$, see Corollary (ref), we may consider a linear sieve space $\mathcal{G}_K=\{\phi: \phi(z)=\gamma_w'p^K(z)\}$. Let $p^K=(p_{1}, \dots, p_{K})$ be a $K$- dimensional vector of known basis functions such as polynomials, splines or similar. We can in principal also apply the general sieve estimation technique of Chen07 based on the conditional moment restriction $\mathop{{\mathbf E}\null}\nolimits[Y|X=x]=H(g_w(z), w)$. This would require to estimate $H$ along with $g_w$ and nesting of two sieve spaces. Our estimation strategy constructively arises from the identification argument and provides a simple direct estimate of $g_w$. We also directly leverage the monotonicity condition on $H$ in the estimation so there is no need to introduce additional shape-constraints.

Convergence Rate

In this section, we derive a rate of convergence of the sieve rank estimator $\widehat g_w$ given in (ref). To keep notation simple, we omit the controls $W$ entirely from the following analysis. In this case, estimation amounts to maximizing the criterion in ((ref)) from the previous section over a suitable sieve space.

For the remainder of the paper we consider the centered criterion function

align[align omitted — 177 chars of source]

where $g$ is the regression function satisfying the model equation ((ref)). Centering does not change the maximizer in the optimization problem and is thus without loss of generality.

Our analysis builds on a linearization of the nonlinear criterion function $\mathcal Q(\cdot)$. The first directional derivative of $\mathcal Q$ is equal to zero for any arbitrary direction and hence, we consider the second directional derivative which can be viewed as a quadratic approximation to the criterion function $\mathcal Q(\cdot)$. Specifically, we introduce

align*[align* omitted — 112 chars of source]

denote the second directional derivative of the non-linear functional $\mathcal Q$ in the direction $\phi-g$. We assume that the functional $Q(\cdot)$ is bi-linear and continuous. Below, we denote $L^2(Z)=\{\phi:\, \|\phi\|_{L^2(Z)}<\infty\}$ where $\|\phi\|_{L^2(Z)}:=\sqrt{\mathop{{\mathbf E}\null}\nolimits\phi^2(Z)}$.

To account for the potential instability of the estimation problem, we introduce the sieve measure of ill-posedness

align*[align* omitted — 109 chars of source]

to account for the fact that the criterion function and the $L^2$-norm are generally not (locally) equivalent. If $\tau_K \to \infty$ as $K\to \infty$ the problem of estimating $g$ is ill-posed in rate and additional regularization slows down convergence in the strong $L^2$- norm. In contrast to ChenPouzo12, we rely on the second directional derivative in the denominator.

For the following assumption we introduce a local neighborhood of $g$ and define the space $\mathcal{G}_K^\delta = \{\phi \in \mathcal{G}_K: \mynorm*{\phi-g}_{L^2(Z)}< \delta\}$ with $\delta > 0$.

assA(i) A random sample $\{(Y_i, Z_i)\}_{i=1}^n$ of $ (Y,Z)$ is observed; (ii) there exists $\Pi_K g\in\mathcal{G}_K$ such that $\|\Pi_K g-g\|_{L^2(Z)}=O(K^{-\alpha/d_z})$; (iii) $\mathop{{\mathbf E}\null}\nolimits[U^2] < \infty $ and $g \in L^2(Z)$; (iv) for any $\phi$ in $\mathcal{G}_K^\delta$ there exists a constant $0<\eta< 1$ such that $|\mathcal Q(\phi)-Q(\phi-g)| \leq \eta \cdot Q(\phi -g)$; (v) the cdf of $g(Z)$ is Lipschitz continuous, i.e., $|F_{g(Z)}(a) - F_{g(Z)}(b)|\leq C|a-b|$ for some constant $C$ and any $a,b$; and (vi) $\tau_K \sqrt{K/n}=o(1)$.

Assumption (ref) (ii) imposes regularity on the regression function $g$ via a sieve approximation error, see also Chen07 for examples. Assumption (ref) (iv) is also known as the tangential cone condition and implies that $\mathcal Q(\phi)$ is locally equivalent to $Q(\phi -g)$ which is a typical condition required to derive the convergence rate for sieve estimators; see ChenPouzo12 and also Dunker2011. Assumption (ref) (v) amounts to a local continuity assumption for the kernel of an empirical process, see e.g. Chen07. Assumption (ref) (vi) restricts the growth of $K$ relative to the sieve measure of ill-posedness $\tau_K$ and is required for consistency, see Lemma (ref).

rem[Illustration of Ill-Posedness] To give an insight on the source of ill-posedness, note that \begin{align*} \mathcal Q(\phi)=\mathop{{\mathbf E}\null}\nolimits\left[Y_i\left(F_{g(Z_i)|Y_i}(g(Z_j)) - F_{\phi(Z_i)|Y_i}(\phi(Z_j)) \right) \right] \end{align*} which shows that if there is little variation in the distribution of $F_{g(Z)|Y}$ for variations of g then the ill-posed inverse problem becomes more severe. This is further illustrated by the following lemma where we study a special case for which we can derive $Q$ analytically and give sufficient conditions for Assumption (ref) (iv). \begin{lem} Consider the additive separable model $g(Z)=Z_{1} + \widetilde g(Z_{2})$ with bivariate $Z=(Z_{1}, Z_{2})$. Then Assumption (ref) (iv) is satisfied if $f^{'}_{Z_{1} | Z_{2}}$ is uniformly bounded away from zero and $f^{''}_{Z_{1} | Z_{2}}$ is uniformly bounded above. \end{lem} The special case outlined in Lemma (ref) illustrates the behavior of $\tau_K$. If the density $f_{Z_{21} | Z_{22}}$, that is the conditional density of the separable covariate, is flat in the relevant support, we may encounter the case that the criterion $\mathcal{Q}$ is close to zero for candidate functions that are arbitrarily far away from the true function in the $L^2$- sense.

We further illustrate this issue in a Monte Carlo simulation study in Section (ref), where we show that the estimation problem becomes more difficult as $f_{Z_{21} | Z_{22}}$ becomes more flat. We are now in a position to provide a general rate of convergence of our sieve rank estimator $\widehat g$.

theoLet Assumptions (ref)-(ref) be satisfied. It holds that \begin{align*} \mynorm*{\widehat g-g}_{L^2(Z)}=O_p\Big( \max \Big\{ \tau_K \sqrt{\frac{K}{n}}, \; K^{-\alpha/d_z} \Big\}\Big) \end{align*}

The proof of Theorem (ref) makes use of a representation of second-order U-processes as empirical processes following ClemLugosi08. To the best of our knowledge, this is the first convergence rate result for nonparametric M-estimators with a rank-based criterion function in the presence of ill-posedness.

The next corollary provides concrete rates of testing when the dimension parameter $K$ is chosen to level variance and square bias under classical smoothness conditions. We call our model mildly ill-posed if: $\tau_k \sim k^{\gamma/d_z}$ with $\gamma > 0$ and severely ill-posed if: $\tau_k\sim\exp(k^{\gamma/d})$, with $\gamma > 0$.\footnote{If $\{a_n\}$ and $\{b_n\}$ are sequences of positive numbers, we use the notation $a_n \lesssim b_n$ if $\limsup_{n\to\infty}a_n/b_n<\infty$ and $a_n\sim b_n$ if $a_n\lesssim b_n$ and $b_n\lesssim a_n$.}

coroLet Assumptions (ref)-(ref) be satisfied. \begin{itemize} • Mildly ill-posed case: setting $K \sim n^{d_z/d_z+2\gamma+2\alpha}$ yields \begin{align*} \mynorm*{\widehat g-g}_{L^2(Z)}=O_p(n^{-\alpha/(2\alpha+2\gamma+d_z)}). \end{align*} • Severely ill-posed case: setting $K\sim \log(n)^{d/\gamma}$ yields \begin{align*} \mynorm*{\widehat g-g}_{L^2(Z)}=O_p(\log(n)^{-\alpha/\gamma}). \end{align*} \end{itemize}

Both convergence rates are the optimal rates for ill-posed problems. As outlined in the discussion following Lemma (ref), the severity of the ill-posedness will generally depend on the chosen normalization and features of the data.

Monte Carlo Simulation Study

This section demonstrates how nonclassical measurement errors in the outcome alters mean regression results in finite samples and shows the usefulness of our approach to correct for such biases. We compare regression function estimates obtained from simply ignoring the measurement error with our estimator, which accounts for the presence of the error. Throughout this section, simulation results are based on a sample of size of $n=1000$ and 1000 Monte Carlo iterations.

We consider the following data generating process

align*[align* omitted — 52 chars of source]

where $Z_1 \sim \mathcal N(1,\sigma^2)$, $Z_2 \sim \mathcal U[-3, 3]$ independent of each other, $g(\cdot)=\sin(\cdot)$ and the error terms $(U,V)\sim\mathcal N(0,I_2)$. Here, $I_2$ is the 2-dimensional identity matrix and for the standard deviation of $Z_1$ we choose $\sigma=1$, which will be varied later. In the above model, $g$ is identified up to a location normalization. Analogously we could specify a linear or nonlinear function on $Z_1$ and impose an additional scale normalization on $g$. The function $h$ in the measurement error equation is chosen as

align*[align* omitted — 243 chars of source]

where $q_{0.3}, q_{0.7}$ denote the $30\%$- and $70\%$-quantile of $Y^*$ (determined via numerical approximation). The setup is analogous to a typical survey data setting with over- or underreporting in the tails of $Y^*$, whereas the center of the distribution is not affected. The scalars $a,b$ are chosen to vary the magnitude of measurement error.

figure[figure omitted — 281 chars of source]

Figure (ref) illustrates the effects of the measurement error for the case $a=b=0.5$. We show the realizations of $Y$ and $Y^*$ for a specific draw of the data generating process and plots the function $h$. We compare the measurement error function $h$ (depicted as red solid line) with the setup of classical measurement error, which is captured by the $45^\circ$ line (depicted as black dashed line).

We implement the sieve rank estimator $\widehat g$ given in (ref) using a linear sieve space with B-spline basis functions of order 3 with 2 interior knots that are placed according to quantiles of the empirical distribution. Thus we have $K=4$. The elements of the sieve space are normalized at the point $(0,0)$ which is the correct value of the true function $\sin(\cdot)$ at $0$. This normalization can also be rationalized as utilizing prior knowledge on the measurement error mechanism in the sense of Assumption (ref) (ii). For instance, we can expect that ignoring the measurement error results in estimates that are close to the true function $g$ in the center of the distribution of $Z_2$. Figure (ref) shows the sieve rank estimates $\widehat g$ and compares them to a nonparametric series regression that does not account for nonclassical measurement error in the outcome using the same order and the same knot placement as for $\widehat g$. For the latter estimator the same choice of basis functions and tuning parameters is adopted.

figure[figure omitted — 481 chars of source]

We study different values for $a,b$ amongst which is the severe case $a=b=0$ which essentially implies that at some point the measurements $Y$ are merely random fluctuations around a constant value\footnote{Additionally we perform Kolmogorov-Smirnov tests to test the null hypothesis that $Y$ and $Y^*$ follow the same probability distribution on every drawn sample of the MC study. In the $a=b=0.5$ setting we reject the null on a $5\%$ - level only once in 1000 samples and in the $a=b=0$ case we reject the null in 966 cases. Thus in the strong ME setting, $Y$ and $Y^*$ have different marginal distributions in contrast to the mild ME setting, where differences are virtually undetectable.}. We observe from the results in Figure (ref) that our estimation strategy results in an accurate estimate of $g$ in any of the cases, whereas ignoring the measurement error yields estimates with a sizeable bias in the tails of $Z_2$. In the severe setting depicted in the right panel, ignoring measurement error results in a rather flat estimate which is significantly different from the sieve rank estimator.

The data generating process chosen here is in line with the model in Lemma (ref) and thus allows us to study the degree of ill-posedness in the convergence rate of the estimator. As pointed out in the discussion following Lemma (ref), the behavior of the sieve measure of ill-posedness $\tau_K$ is governed by the conditional density $f_{Z_1|Z_2}$. If the density $f_{Z_1|Z_2}$ is flat over the relevant support, $\tau_K$ diverges faster and the ill-posedness is more severe.

Table (ref) below shows mean squared errors of function estimates across different standard deviations of the separable covariate $Z_1$ which affects the slope of the density $f_{Z_1|Z_2}$. For small standard deviations, the conditional density $f_{Z_1|Z_2}$, i.e., here $f_{Z_1}$ by full independence, will be rather flat over most of the support. For small standard deviations of $Z_1$, the MSE increases more severely with $K$ as compared to large standard deviations. This illustrates that the degree of ill-posedness of the estimation problem is more severe whenever the slope of the density $f_{Z_1|Z_2}$ is small.

table[table omitted — 795 chars of source]

Additionally we see that this is not the case when the distribution of $Z_1$ is fixed and the dispersion of $Z_2$ is varied. This confirms that the ill-posedness in this setting is not driven by the distribution of $Z_2$ in this setting.

Application: Beliefs on Stock Market Returns

Subjective beliefs on stock market returns are an important determinant in economic models that seek to explain stock market participation and portfolio choice, see e.g. HuckWeiz19 and the references therein. Subjective belief data, however, is known to be prone to a large degree of measurement error, see the discussion and references in GaudeckerME.

We study the impact of historic return information on subjective beliefs of future stock market returns. We account for nonclassical measurement error in the outcome variable by applying our sieve rank method and contrast the results to a model where we simply ignore measurement error in the outcome.

We use novel data from the innovation sample of the 2017 wave of the German Socio Economic Panel (SOEP-IS), which contains survey questions on individual beliefs on future stock market returns. In the interviews, respondents are asked their expectations on the DAX, Germany’s prime blue chip stock market index, in one, two, ten and thirty years with respect to the current level. They are asked to provide a direction of the change (increase or decrease) as well as a percentage change.

Prior to elicitation of their beliefs, individuals obtain information about historical DAX returns. Two observations of the time series of yearly DAX returns from $1951$ to $2016$ are randomly drawn and presented to the respondent. Afterwards they are asked to report their beliefs on how the DAX changes in the next year (in percentage points).

table[table omitted — 433 chars of source]

In this application, we are interested in the effect of the historical DAX information on the individuals expected DAX return in one year. Let $Y^*$ denote the individual true belief on the DAX return in one year and let $Z_1,Z_2$ be the two treatment variables, i.e., the randomly drawn historical returns. The reported belief is denoted by $Y$. We consider the following flexible additively separable model

align[align omitted — 126 chars of source]

It is difficult to rationalize a classical measurement error assumption a priori. Various forms of nonclassical measurement error may occur in this setting: (i) Respondents may tend to provide rounded values instead of precise beliefs, (ii) respondents may systematically over- or underreport their beliefs, e.g., individuals with extreme beliefs may resort to reporting more modest values, or (iii) the reporting may additionally depend on variables $W$ such as certain cognitive skills or personality traits like patience or perseverance. Note that by the experimental design $Z_1,Z_2$ and $W$ are credibly fully independent so there is no need to specify the variables in $W$ or to apply our weighted sieve rank estimator.

figure[figure omitted — 397 chars of source]

We now discuss the plausibility of Assumptions (ref)-(ref) required for identification. Assumptions (ref) posits that given true beliefs $Y^*$ and relevant individual characteristics $W$, the historic return information $Z_1, Z_2$ have no impact on the mean reported belief. Assumption (ref) imposes a mild restriction on the measurement error mechanism in that it requires monotonicty in the reporting of beliefs (in the conditional mean). Assumption (ref) is satisfied as $Z_1, Z_2$ are by the experimental setup credibly fully independent of unobservables $U$. The data consists of 1084 interviewed persons but 306 people do not respond to the question on beliefs. We removed missing values and report the summary statistics in Table (ref).

We estimate functions $g_1$, $g_2$, $g_3$ with our method outlined in ((ref)) and contrast the results to estimates obtained from assuming classical measurement error, i.e., from a standard additive-separable, nonparametric regression of $Y$ on $Z_1$ and $Z_2$ with the respective interaction term. We choose a B-Spline basis of degree two without interior knots for each function estimate. This choice is motivated by a 10-fold cross-validation on the model ignoring the measurement error.

The results are presented in Figure (ref). Accounting for the measurement error leads to a concave, symmetric effect of both treatments on the individual beliefs. When ignoring the possibility of measurement error, results are much more asymmetric, including convex marginals for the first treatment and flat parts in the surface. In contrast, our method yields that individuals learn conservatively from both treatments which is in line with the a priori economic intuition. Note that on the z-axis that estimates in both columns have been normalized to move through coordinates (-20,-20,0) and (50,50,1). Functions are evaluated on a grid ranging from -20 to 50 which corresponds to the $10\%$- and $90\%$-quantile of the marginal distributions of the treatment variables. Summarizing, accounting for possible nonclassical measurement error in the outcome variable delivers function estimates of belief formation that are more in line with economic intuition.

Conclusion

This paper provides new insights on the analysis of regression models with nonclassical measurement error in the outcome variable. Our nonparametric identification result is based on intuitive assumptions involving shape restrictions on measurement error functions. This novel result builds on the equivalence of nonclassical measurement models and generalized regression models. We consider a sieve rank estimator which constructively arises from our identification result and implicitly accounts for the required shape restrictions. We establish the rate of convergence of the sieve rank estimator which is affected by a potentially ill-posed inverse problem. The proposed estimation method is easy to implement and provides numerically stable results as demonstrated in a finite sample analysis. Finally, we demonstrate the usefulness of our method in an empirical application on belief elicitation, where we find measurement error in subjective belief data to be of a nonclassical form.