EconBase
← Back to paper

Specification Testing in Nonparametric Instrumental Quantile Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,861 characters · 16 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Specification Testing in Nonparametric Instrumental Quantile Regression

abstractThere are many environments in econometrics which require nonseparable modeling of a structural disturbance. In a nonseparable model with endogenous regressors, key conditions are validity of instrumental variables and monotonicity of the model in a scalar unobservable variable. Under these conditions the nonseparable model is equivalent to an instrumental quantile regression model. A failure of the key conditions, however, makes instrumental quantile regression potentially inconsistent. This paper develops a methodology for testing the hypothesis whether the instrumental quantile regression model is correctly specified. Our test statistic is asymptotically normally distributed under correct specification and consistent against any alternative model. In addition, test statistics to justify the model simplification are established. Finite sample properties are examined in a Monte Carlo study and an empirical illustration is provided.
tabbingKeywords: \=Nonparametric quantile regression, instrumental variable,\\ \> specification test, local alternative, nonlinear inverse problem.\\[.2ex] \>

Introduction

Regression models that involve instrumental variables are widely used in economics to overcome endogeneity problems. In these models, assuming the structural disturbances to be additively separable implies that marginal effects do not depend on unobserved characteristics, which may be difficult to justify. This is why their nonseparable extension has received a lot of attention recently. Under certain key conditions, the nonseparable model is equivalent to an instrumental quantile regression model. These conditions are validity of instruments and monotonicity of the model in a scalar unobservable. If one of these conditions is violated, however, the quantile regression representation is misspecified.

In this paper, we propose a specification test of the instrumental quantile regression model

equation[equation omitted — 99 chars of source]

for each $0<q<1$, where $Y$ is a scalar dependent variable, $Z$ a vector of potentially endogenous regressors, $W$ a vector of instruments, and $U(q)$ an unobservable disturbance.\footnote{Since conditional expectations are defined only up to equality almost surely, all (in)equalities with conditional expectations and/or random variables are understood as (in)equalities almost surely.} This quantile regression model is equivalent to a nonseparable model (cf. HL07econometrica) given by

equation[equation omitted — 52 chars of source]

where

itemize• the instrumental variable $W$ is independent of $V$, • the function $\sol$ is strictly monotonic increasing in the scalar disturbance $V$, and • $V \sim \cU (0,1)$.

Condition (a.3) can be assumed without loss of generality if $V$ is continuously distributed with positive density on its support which we assume to hold throughout the paper. The quantile regression model (ref) for all $0<q<1$ is thus misspecified if in its nonseparable version (ref) the instrument is not valid, that is, $W$ is not independent of $V$, or the function $\sol$ is not monotonic in $V$.

Specification testing in instrumental variable models is a subject of considerable literature. In the context of nonparametric instrumental mean regression $Y=g(Z)+U$ with $\Ex[U|W]=0$, tests for correct specification have been proposed by GS07, Horowitz2009, and Breunig2012. These tests are, however, not robust against potential nonseparability of the structural disturbance. On the other hand, by considering the nonseparable model (ref) with conditions (a.1)--(a.3) a failure of the exclusion restriction of the instruments might only be one source of misspecification. Indeed, as argued by hoderlein2007, in certain applications, such as consumer demand, the monotonicity restriction (a.2) might be highly unrealistic. As such, providing a specification test of model (ref) together with conditions (a.1)--(a.3) seems paramount but, as far as we know, has not yet been addressed in the literature.

Research on identification and estimation in nonparametric instrumental quantile regression has been active in the last decade. 2003Chesher establishes nonparametric identification of derivatives of the unknown functions in a triangular array structure. Chern2005 and Chern2007 give identification conditions and develop a nonparametric minimum distance estimator. Sufficient conditions for local identification are given by chen2012. HL07econometrica propose an estimator based on Tikhonov regularization, chen08 study penalized sieve minimum distance estimation, and Dunker2011 consider an iteratively regularized Gau\ss-Newton method. Further, Gag2012 obtain asymptotic distribution results of a Tikhonov regularized estimator. There is also a large literature on testing quantile regression models with exogenous covariates. In this context particularly relevant is quantile regression testing using an infinite number of quantiles for parametric functions, see escanciano2010 and, in the nonparametric context, escanciano2014.

In instrumental quantile regression (ref) for a fixed quantile $0<q<1$, HorLee2009 established a test of parametric specification of $\sol$. chen2015 consider functionals of semi/nonparametric conditional moment restrictions with possibly nonsmooth generalized residuals. A test of monotonicity in unobservables of $\sol$ has been proposed by hoderlein2011 but requires conditional exogeneity of $Z$ and hence, is not related to instrumental variables methodology. Recently and independently of this paper, feve2013 developed a test of whether $Z$ is independent of the nonseparable disturbance $V$ in the model (ref).

Our test statistic is based on the $L^2$--norm of the empirical conditional quantile restriction and involves sieve methodology. The sieve approach makes the statistic easy to implement and further, is convenient to impose additional constraints on the structural function $\sol$. As an example, we discuss a test of additivity of $\sol$ with respect to the vector of regressors $Z$. In addition, we establish a test statistic for testing exogeneity which is robust against nonseparability. More precisely, we establish a test of exogeneity of the regressors $Z$ at some quantile $0<q<1$, that is, whether $\PP(Y\leq \sol(Z,q)|Z)=q$. This extends the results on nonparametric tests of exogeneity in mean regression suggested by Blundell2007 and Breunig2012 to the quantile regression case.

It should also be noted that the test proposed in this paper is a joint test of monotonicity and instrument validity. This is the nature of many nonparametric tests, see, for instance, chiappori2015 or lewbel2015. On the other hand, we show in this paper how the sign of $\PP(Y\leq \sol(Z,q)|W)-q$ can be exploited to make inferences on the validity of the instrumental variables. As such, in many cases it is possible to detect the cause of a rejection of our test.

We establish the asymptotic distribution of our test statistic under the null hypothesis and its consistency against fixed alternatives. We study the power of our test against a sequence of local alternatives. By Monte Carlo simulations we demonstrate the power properties of our test in finite samples. As an empirical illustration, we study a nonseparable model of the effects of class size on test scores of 4th grade students in Israel. We reject the hypothesis of exogeneity of class size but fail to reject the instrumental variable model.

The remainder of this work is organized as follows. In Section (ref), we propose a test statistic and obtain its asymptotic distribution. We further establish consistency of our test. The power of the test is judged by considering a sequence of local alternatives. Section (ref) gives several extensions of the previous results. In Section (ref) and (ref) we study the finite sample properties of our test and give an empirical illustration. All proofs can be found in the appendix.

The test statistic and its asymptotic properties

This section begins with the definition of the test statistic and states assumptions required to obtain its asymptotic distribution under the null hypothesis. Moreover, we study power and consistency properties of our test.

Definition of the test statistic

The quantile regression model (ref) leads to a nonlinear operator equation, as we see in the following. Let $\varPhi$ be a Banach space endowed with the norm $\|\phi\|_{Z,p}:=(\Ex|\phi(Z)|^p)^{1/p}$ for some integer $p>0$ and if $p=\infty$ then $\|\phi\|_{Z,\infty}:=\sup_{z}|\phi(z)|$. For simplicity let $\|\phi\|_Z:=\|\phi\|_{Z,2}$. Further, let us introduce the Hilbert space $L_W^2:=\{\psi:\,\|\psi\|_W^2:=\Ex|\psi(W)|^2<\infty\}$. We define a nonlinear operator $\cT:\varPhi\to L_W^2$ with

equation[equation omitted — 76 chars of source]

for any $\phi\in\varPhi$ where $\1$ denotes the indicator function. Thereby, model (ref) can be rewritten as the operator equation $\cT\solq=q$ with $\sol_q(\cdot):=\sol(\cdot,q)$ for all $0<q<1$.

In many economic applications, for instance when estimating a demand function or Engel curves, the structural function of interest may be assumed to be smooth. This a priori knowledge is captured by a set $\cB\subset \varPhi$ which we introduce below. The set $\cB$ may also contain constraints on the function $\solq$ such as monotonicity, concavity/convexity or additivity (see also Section (ref)) and can also ensure uniqueness of $\solq$ (see Example (ref) below). Let us introduce the set $\cB^{(0,1)}=\set{\phi:\, \phi(\cdot,q)\in\cB \text{ for all }q\in(0,1)}$. We consider the null hypothesis

equation[equation omitted — 148 chars of source]

The alternative is that there exists no function $\sol\in\cB^{(0,1)}$ solving $\cT\solq=q$ for all $q\in(0,1)$.

We construct in the following a test statistic based on the $L^2$--distance. Throughout the paper, we assume that an independent and identically distributed $n$-sample of $(Y,Z,W)$ is available. Let $\{f_j\}_{j\geq 1}$ be a sequence of approximating functions in $L_W^2$. Then, for any integer $k\geq 1$ we denote $f_\uk(\cdot)=(f_1(\cdot),\dots,f_k(\cdot))^t$ and $\bold W_k=\big(f_\uk(W_1),\dots,f_\uk(W_n)\big)^t$ which is a $n\times k$ matrix. A series least square estimator of $\Ex[\1\set{Y\leq \phi(Z)}-q|W=\cdot]$ then writes

equation*[equation* omitted — 112 chars of source]

where $(\cdot)^-$ denotes a generalized inverse. Further, we define the sieve least square estimator of $\solq$ by

equation[equation omitted — 208 chars of source]

where $\cB_\k$ is a $\k$--dimensional sieve space that becomes dense in $\cB$ as the sample size $n$ tends to infinity. If $\cB$ contains additional constraints then these are imposed in $\cB_\k$ on the finite dimensional functions. Here, $\k$ and $\l$ grow with sample size $n$. Clearly, $\k\leq \l$ for each $n$ is required and in our simulations we choose $\l = C\k$ for some constant $C>1$ (see also chen2015optimal in the case of nonparametric instrumental mean regression). The estimator $\hsol_{qn}$ is a simplified version of the penalized sieve minimum distance estimator suggested by chen08.

The test statistic is then given by

equation[equation omitted — 186 chars of source]

where $\m$ grows with sample size $n$. As the test is one sided, we reject the null hypothesis at level $\alpha$ when the standardized version of $S_n$, namely $3\sqrt{5/\m}\big(S_n-\m/6\big)$, is larger than the $(1-\alpha)$--quantile of $\cN(0, 1)$. The asymptotic distribution of $S_n$ is derived below under mild restrictions on the dimension parameters $\k$, $\l$, and $\m$. We require that the number of unconditional moment restrictions determined by $\m$ is asymptotically larger than the dimension of the sieve space $\cB_\k$. This corresponds to the test of overidentifying restrictions in parametric models. In contrast to the parametric setting, however, also the number of unconditional moment restrictions used to construct the estimator (determined by $\l$) must be asymptotically smaller than the number of moment restrictions used for the test statistic. This ensures that the estimation error in the test statistic becomes asymptotically negligible as we see below.

Our test statistic builds on the nonparametric specification test in instrumental mean regression suggested by Breunig2012. Testing in instrumental quantile regression, on the other hand, requires a different methodology. First, the test statistic is a discontinuous function of the unknown structural effect $\solq$. Second, instrumental quantile regression leads to a nonlinear inverse problem and hence deriving asymptotic results is more challenging. Third, to verify the conditional moment restrictions for all quantiles we need to integrate over them. In the appendix, we show that the mapping $q\mapsto \solq$ is continuous under mild assumptions. This justifies the use of our $L^2$--type rather than a sup norm statistic.

Assumptions and notation

In order to obtain our asymptotic result we state the following assumptions. Our first assumption gathers conditions which we require for the basis functions $\{f_j\}_{j\geq1}$. In the following, the supports $\cZ$ of $Z$ and $\cW$ of $W$ are assumed to be bounded (see also Assumption (ref)). The probability density function (p.d.f.) of $W$, denoted by $p_W$, is assumed to be uniformly bounded from above and away from zero.

assA(i) There exists a constant $C>0$ and a sequence of positive integers $(m_n)_{n\geq 1}$ satisfying $\sup_{w\in\cW}\|f_\umn(w)\|^2\leqslant C\m$. (ii) The smallest eigenvalue of the matrix $\Ex[f_\um(W)f_\um(W)^t]$ is bounded away from zero uniformly in $m$.

Assumption (ref) $(i)$ holds for sufficiently large $C$ if $\{f_j\}_{j\geq1}$ are trigonometric basis functions, B-splines, or wavelets. Assumption (ref) $(ii)$ is satisfied if the marginal density of $W$ is uniformly bounded away from zero on $\mathcal W$ and $f_{\umn}$ forms a vector of orthonormal basis functions. For any $\phi\in\cB^{(0,1)}$ we write $\phi_q(\cdot):=\phi(\cdot,q)$ for all $0<q<1$. We denote the Fr\'echet derivative of $\cT$ at $\solq$ by

equation*[equation* omitted — 86 chars of source]

where $p_{Y|Z,W}$ denotes the density of $Y$ conditional on $(Z,W)$. We introduce the notation $\interleave\phi\interleave_{Z,p}=\big(\int_0^1\|\phi(\cdot,q)\|_{Z,p}^pdq\big)^{1/p}$ and $\interleave\psi\interleave_W=\big(\int_0^1\|\psi(\cdot,q)\|_W^2dq\big)^{1/2}$ for functions $\phi(\cdot,q)\in\varPhi$ and $\psi(\cdot,q)\in L_W^2$ for all $q\in(0,1)$.

assA(i) If $\interleave\cT\phi-\cT\sol\interleave_W^2=0$ for some function $\phi\in\cB^{(0,1)}$ then it holds $\interleave\phi-\sol\interleave_{Z,p}^2=0$. (ii) There exists some constant $0<\eta<1$ such that for all $0<q<1$ and all functions $\phi\in\set{\phi\in\cB:\,\|\phi-\sol_q\|_{Z,p}\leq \varepsilon}$ for some $\varepsilon>0$ it holds \begin{equation} \|\cT\phi-\cT\solq-\op_q(\phi-\solq)\|_W\leq \eta\|T_q(\phi-\solq)\|_W. \end{equation}

Assumption (ref) $(i)$ ensures identification of $\solq$ for almost all $0<q<1$ on the set $\cB$ which we introduce below. Assumption (ref) $(ii)$ specifies an upper bound on the Taylor remainder of $\cT$ in a small neighborhood around $\solq$. It is also known as the tangential cone condition and frequently used in the analysis of nonlinear operator equations (cf. Hanke1995 or Dunker2011 in case of instrumental variable estimation). We provide sufficient conditions for the tangential cone condition in Example (ref) below and refer to chen2012 for further discussions.

assAThere exists a sequence $(r_n)_{n\geq 1}$ with $r_n=o(1)$ such that for constants $C>0$ and $\kappa\in(0,1]$ it holds \begin{equation} \max_{1\leq j\leq \m}\Ex\Big[\int_0^1\sup_{\phi\in\cB_{n}}\big|\1\{Y\leq\phi(Z,q)\}-\1\{Y\leq \sol(Z,q)\}\big|^2dq\,f_j^2(W)\Big] \leq Cr_n^{2\kappa} \end{equation} where $\cB_{n}:=\{\phi\in\cB^{(0,1)}:\,\interleave\phi-\sol\interleave_{Z,p}^2\leq r_n^2\}$.

Assumption (ref) states that the function $\solq\mapsto(\1\{Y\leq\sol(Z,q)\}-q)f_j(W)$, $1\leq j\leq \m$, is locally uniformly $L_W^2$ continuous for almost all $0<q<1$. This condition has also been exploited by CLVK03econometrics (Theorem 3), Chen07 (Lemma 4.2 (i)) or chen08 (Remark c.1). Example (ref) below gives primitive conditions under which Assumption (ref) holds true.

Let $\cZ\subset \mathbb R^{d_z}$ and for any vector of nonnegative integers $k=(k_1,\dots,k_{d_z})$ define $|k|=\sum_{j=1}^{d_z}k_j$ and $D^k=\delta^{|k|}/(\delta z_1^{k_1}\dots\delta z_{d_z}^{k_{d_z}})$. For some integer $p>0$ we define the norms

equation*[equation* omitted — 205 chars of source]

where $\alpha$ and $\alpha_0$ are positive integers. We denote the Sobolev spaces associated with the norm $ \|\cdot\|_{\alpha,p}$ by

equation[equation omitted — 87 chars of source]

For some constant $\rho>0$, define $\cB$ as the Sobolev ellipsoid of radius $\rho$ given by

equation[equation omitted — 93 chars of source]

On the other hand, our sieve space $\cB_\k$ used to approximate $\cB$ is compact under $\|\cdot\|_Z$ and thus, penalization is not necessary for consistent estimation (see also chen08). Also additional constraints such as monotonicity can be imposed by $\cB=\set{\phi\in W^{\alpha,p}:\,\|\phi\|_{\alpha,p}\leq \rho, \,\inf_{z\in\cZ}\phi'(z)>0}$ for scalar $z$. Such a monotonicty constraint does not necessarily lead to faster rates of convergence, in contrast to an additivity restriction on $\solq$. Consequently, we do not treat shape restrictions like monotonicty explicitly but only discuss a test of additivity in Section (ref). In this context, we also refer to chetverikov2015 for using shape restriction for sieve estimation in instrumental mean regression. The following assumption gathers regularity conditions imposed on the structural functions $\sol$ and the supports $\cZ$ and $\cW$.

assA(i) Let $\alpha_0>{d_z}/p$ and $\alpha>{d_z}/\kappa$. (ii) $\cZ$ is bounded, convex and satisfies a uniform cone property. (iii) $\cW$ is bounded. (iv) The marginal density of $W,$ denoted by $p_W$, is bounded from above and uniformly bounded away from zero on $\cW$. (v) $p_{Y|Z,W}(\cdot,Z,W)$ is bounded from above.

Assumption (ref) $(i)$ requires $\alpha$ to be large if (ref) holds only for small $\kappa>0$ or the dimension ${d_z}$ is large. Assumption (ref) $(ii)$ imposes a weak regularity condition on the shape of $\cZ$. For the uniform cone property see, for instance, Paragraph 4.4 in 2003Adams. This property was also used by Santos12. Assumption (ref) $(v)$ ensures that $\|T_q\phi\|_W\leq C\|\phi\|_Z$ for all $\phi\in L_Z^2$ and some constant $C>0$.

example[Primitive Conditions for Assumption (ref)] Let $\varPhi$ coincide with the Hilbert space $L_Z^2:=\set{\phi:\,\|\phi\|_Z<\infty}$. If for any $0<q<1$ the operator $T_q$ is compact then there exists an orthonormal basis in $L_Z^2$ denoted by $\{e_j\}_{j\geq 1}$ satisfying $\|T_q\phi\|_W^2=\sum_{j=1}^\infty s_{qj}^2\Ex[\phi(Z)e_j(Z)]^2$ where $(s_{qj})_{j\geq 1}$ are the singular values of $T_q$. If \begin{equation*} \cB\subset\cB_{source,q}:=\set{\phi\in L_Z^2:\,\sum_{j=1}^\infty s_{qj}^{-2}\Ex[(\phi(Z)-\sol(Z,q))e_j(Z)]^2<c_0} \end{equation*} for some constant $c_0>0$ then, under mild assumptions on the joint distribution of $(Y,Z,W)$, the function $\solq$ is identified on $\cB$ (cf. Theorem 6 of chen2012). A similar restriction was also imposed by HL07econometrica. If $\cB\subset \bigcap_{q\in(0,1)}\cB_{source,q}$ then Assumption (ref) $(i)$ holds true. Under further assumptions, imposing bounds on the generalized Fourier coefficients is equivalent to imposing smoothness restrictions. To illustrate this relation let $Z$ be a scalar uniformly distributed random variable and assume $s_{qj}=j^{-\zeta}$, $j\geq 1$, for some constant $\zeta>0$. In this case, if $\set{e_j}_{j\geq 1}$ are the usual trigonometric basis functions then $\cB_{source,q}$ coincides with the Sobolev space of $\zeta$--times differentiable functions with periodic boundary conditions, while if $s_{qj}^2=\exp(-j^{2\zeta})$, $j\geq 1$ and $\zeta>1$, $\cB_{source,q}$ contains only analytic functions (see also Kr89). In this sense, $\cB_{source,q}$ links the smoothness of $\phi-\sol_q$ to the degree of ill-posedness determined by the degree of decay of $(s_{qj})_{j\geq 1}$, which is also known as a so-called source condition (cf. ChenReiss2008 or Dunker2011 for a further discussion). Under the singular value decomposition of $T_q$ it is also possible to provide primitive conditions for the tangential cone condition (ref). Assume that the conditional p.d.f. of $Y$ given $(Z,W)$, denoted by $p_{Y|Z,W}$, is continuously differentiable with $|\partial p_{Y|Z,W}(\cdot,Z,W)/\partial y|\leq c_1$ and the conditional p.d.f. of $Z$ given $W$ satisfies $p_{Z|W}(\cdot,W)\leq c_2p_{Z}(\cdot)$, for some constants $c_1,c_2>0$. Then by Theorem 6 of chen2012 it holds \begin{equation} \|\cT\phi-\cT\solq-\op_q(\phi-\solq)\|_W\leq c_1\, c_2\,\|\phi-\solq\|_Z^2. \end{equation} We further obtain for all $\phi\in\cB_{source,q}$ by making use of the Cauchy-Schwarz inequality \begin{align*} \|\phi-&\solq\|_Z^2=\sum_{j=1}^\infty \frac{s_{qj}}{s_{qj}}\Ex[(\phi(Z)-\sol(Z,q))e_j(Z)]^2\\ &\leq \Big(\sum_{j=1}^\infty s_{qj}^{-2}\Ex[(\phi(Z)-\sol(Z,q))e_j(Z)]^2\Big)^{1/2} \Big(\sum_{j=1}^\infty s_{qj}^{2}\Ex[(\phi(Z)-\sol(Z,q))e_j(Z)]^2\Big)^{1/2}\\ &\leq c_0^{1/2}\,\|T_q(\phi-\solq)\|_W. \end{align*} Consequently, the tangential cone condition (ref) is satisfied if we assume $c_0^{1/2}\,c_1\, c_2<1$. We also note that for our test of exogeneity in Section (ref) only the weaker condition (ref) is required. $\hfill\square$
example[Primitive Conditions for Assumption (ref)] Let $F_{Y|ZW}$ denote the cumulative distribution function of $Y$ given $(Z,W)$ and assume that it is Lipschitz continuous with constant $C_L>0$, that is, $|F_{Y|ZW}(y)-F_{Y|ZW}(y')|\leq C_L|y-y'|$ for all $(y,y')$. Due to Assumption (ref) the Sobolev space $W^{\alpha,p}$ can be embedded in $W^{\alpha, \infty}$ (cf. Theorem 6 of 2003Adams). In particular, the supremum norm is bounded on $\cB$ and moreover, Assumption (ref) holds true. Indeed, $\int_0^1\|\phi_q-\solq\|_\infty^2dq\leq r_n^2$ implies $\|\phi_q-\solq\|_\infty\leq c \,r_n$ for almost all $0<q<1$ and some constant $c>0$. Hence, $\sol(Z,q)- c \,r_n\leq \phi(Z,q)\leq \sol(Z,q)+ c \,r_n$ for almost all $0<q<1$ and following CLVK03econometrics (page 1599 -- 1600) we observe \begin{align*} \Ex\Big[\int_0^1\max_{\phi\in\cB_{n}}\big(\1\{Y\leq\phi&(Z,q)\}-\1\{Y\leq \sol(Z,q)\}\big)^2dq\,\Big|W\Big]\\ \leq& \int_0^1\Ex\Big[\1\big\{Y\leq\sol(Z,q)+ c \,r_n\big\}-\1\big\{Y\leq \sol(Z,q)- c \,r_n\big\}\Big|W\Big]dq\\ =&\int_0^1\Ex\Big[F_{Y|ZW}\big(\sol(Z,q)+ c \,r_n\big)-F_{Y|ZW}\big(\sol(Z,q)- c \,r_n\big)\Big|W\Big]dq\\ \leq& C_L\, c \,r_n \end{align*} which implies Assumption (ref) with $\kappa=1/2$. $\hfill\square$
rem[Local Overidentification] In this remark, we discuss local overidentification restrictions in nonparametric instrumental quantile regression for some $0<q<1$. As chen2015overidentification point out in their Example 5.2, the range of the Fr\'echet derivative $T_q$, is given by \begin{align*} \mathcal R_q =\set{\psi\in L_W^2:\, \psi=T_q\phi for some \phi\in L_Z^2}. \end{align*} Local overidentification corresponds to the case where the closure of the range $\mathcal R_q$ is a strict subset of $L_W^2$. In this paper, the class of structural functions $\sol$ is restricted to belong to an ellipsoid $\cB$ and thus, we consider for each $q$: \begin{align*} \mathcal R_q(\cB) =\set{\psi\in L_W^2:\, \psi=T_q\phi for some \phi\in \cB}. \end{align*} Mild restrictions on the ellipsoid $\cB$ imply local overidentification and hence, the class of functions in the alternative model is not empty. $\hfill\square$

The next result formalizes the discussion of the previous remark and shows that the regularity conditions imposed on the function set $\mathcal B$ ensure overidentification.

propLet $\varPhi$ coincide with the Hilbert space $L_Z^2$ and let Assumption (ref) $(v)$ be satisfied. Then we have local identification, i.e., for any $0<q<1$ the closure of $\mathcal R_q(\cB)$ is a strict subset of $L_W^2$.

The proof of Proposition (ref) relies on the fact that the functions in $\mathcal B$ are bounded by some constant $\rho>0$ and, in particular, no smoothness restrictions are employed here to achieve overidentification.\footnote{I thank an anonymous referee for suggesting this argumentation.} It is also possible to achieve overidentification for classes containing unbounded functions, as long as they satisfy minimal smoothness conditions.

The following result is due to chen2015overidentification and gives a condition for local overidentification without imposing a priori restrictions on the set of functions $\mathcal B$.

lem[chen2015overidentification] The model is locally overidentified if and only if \begin{align*} \left\{\psi\in L_W^2 : \EE[p_{Y|Z,W} (\varphi(Z,q),Z,W)\psi(W)|Z] = 0\right\} \neq \{0\}. \end{align*}

Lemma (ref) provides a necessary and sufficient condition for local overidentification without imposing regularity or other shape restrictions. This result involves the adjoint of the Fr\'echet derivative $T_q$ and can be characterized more explicitly in different cases. For instance, assume that the vector of instruments can be decomposed such that $W=(W_1,W_2)$ with $p_{Y|Z,W} =p_{Y|Z,W_1}$, i.e., $W_2$ has no additional information on $Y$ which is not contained in $(Z,W_1)$. In this case, we have

align*[align* omitted — 130 chars of source]

and hence, the model is locally overidentified when there exists a nontrivial function $\psi$ such that $\EE[\psi(W)|Z,W_1]=0$. The last criterion is satisfied, for instance, if $W_2$ is independent of $(Z,W_1)$ for all $\psi$ which only depend on $W_2$ and $\EE[\psi(W_2)]=0$.

\paragraph{Notation} For any $\phi\in\cB$ we introduce $\varPi_\k \phi\in\cB_\k$ satisfying $\|\varPi_\k\phi-\phi\|_{Z,p} = o(1)$. Further, we define

equation*[equation* omitted — 164 chars of source]

The rate $\omega_n$ captures the variance and bias part for estimating $\mathcal T \phi$ for a fixed function $\phi$ and also contains the bias for approximating the structural function $\sol$ in the weak norm induced by the Fr\'echet derivative of $\mathcal T$. Following chen08 we introduce the sieve measure of local ill-posedness by

equation*[equation* omitted — 152 chars of source]

where $\cA_\k=\set{\phi\in\cB_\k^{(0,1)}:\interleave T_{\cdot}(\phi-\sol)\interleave_W^2>0}$. We write $a_n\sim b_n$ when there exist constants $c,c' > 0$ such that $c b_n\leq a_n\leq c' b_n$ for sufficiently large $n$.

Asymptotic distribution under the null hypothesis

The following theorem establishes asymptotic normality of the test statistic $S_n$ after standardization under the null hypothesis $H_0$.

theoLet Assumptions (ref)--(ref) be satisfied. Assume that \begin{equation} m_n^{-1}=o(1),\quad \m=o(n^{1/2}) \end{equation} and in addition \begin{equation} n\omega_n=o(\sqrt{\m}) and \interleave\varPi_\k\sol-\sol\interleave_{Z,p}^2+\tau_\k\omega_n=o\big(m_n^{-(1+\epsilon)/\kappa}\big) \end{equation} for some $\epsilon>0$. Then we have under $H_0$ \begin{equation*} 3\sqrt{5/\m}\big(S_n-\m/6\big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

To motivate the constants in the sieve mean and variance, respectively, we observe

equation*[equation* omitted — 84 chars of source]

and

multline*[multline* omitted — 168 chars of source]

see also the proof of Lemma (ref). The required rate imposed in (ref) on $\m$ is milder than the rate requirement $\m=o(n^{1/3})$ imposed by Breunig2012 in case of nonparametric instrumental mean regression. This is due to the fact that in the latter case we do not have a lower bound for the sieve standard deviation in general, while in case of quantile regression the sieve standard deviation is $\sqrt\m$ within a positive constant. This can be exploited to weaken rate restrictions on $\m$. Further, note that restriction (ref) implies $\k=o(\sqrt\m)$ (by using that $\l\leq \k$). This requirement essentially determines the degree of overidentification required for inference.

The rate restriction $\tau_\k\omega_n=o\big(m_n^{-(1+\epsilon)/\kappa}\big)$ imposed in condition (ref) implies that the dimension parameter $\m$ dominates the effect of estimation of the structural function. Consequently, the asymptotic behavior of our test statistic is not affected by the estimation of $\sol$, regardless of the underlying degree of ill-posedness. Note that this rate restriction can be ensured by choosing $\k$ relative to decay of the sieve measure of local ill-posedness, which is described in more detail in Example (ref) below. We illustrate below that condition (ref) is satisfied under common smoothness restrictions on $\sol$ and mapping requirements of the Fr\'echet derivative $T_q$.

remConsider the Hilbert space case $\varPhi=L_Z^2$ and let $\{e_j\}_{j\geq 1}$ be an orthonormal basis in $L_Z^2$. In this case, $\varPi_\k\phi=\sum_{j=1}^\k\Ex[\phi(Z)e_j(Z)]e_j$. Let us assume the following two conditions. \begin{enumerate} • Sieve approximation error: $\|\varPi_\k\phi-\phi\|_Z=O(k_n^{-\alpha/d_z})$ for all $\phi\in\cB$. • Link condition: $\int_0^1\|T_q (\varPi_\k\phi-\phi)\|_W^2dq\leq \sum_{j\geq1}\opw_j\Ex[(\varPi_\k\phi-\phi)(Z)e_j(Z)]^2$ for all $\phi\in\cB$ and some positive nonincreasing sequence $(\opw_j)_{j\geq 1}$. \end{enumerate} If the p.d.f. $p_Z$ of $Z\in[0,1]^{d_z}$ is bounded then it is well known that the sieve approximation error condition holds for splines, wavelets, and Fourier series bases. Due to Assumption (ref) $(v)$ the link condition is always satisfied with $\opw_j=1$ for all $j\geq 1$. The link condition implies an upper bound for the sieve measure of ill-posedness; that is, $\tau_\k\leq C\opw_\k$ for some constant $C>0$ and all $n\geq 1$ (cf. Lemma B.2 of chen08). Consequently, the first part of condition (ref) simplifies to \begin{equation*} \max\big(\l,n\,l_n^{-2\beta/d_w}, n\opw_\k k_n^{-2\alpha/d_z}\big)=o(\sqrt{\m}) \end{equation*} if $\{\cT\phi:\,\phi\in\cB_\k\}$ belongs to a H\"older space with H\"older parameter $\beta$. In addition, in the setting of Example (ref), the second part of condition (ref) simplifies to \begin{equation*} m_n^{1+\epsilon}\max\big(n^{-1}\l,\,l_n^{-2\beta/d_w},k_n^{-2\alpha/d_z}\big)=o(1) \end{equation*} for some $\epsilon>0$. $\hfill\square$

In the next example, we illustrate different mapping properties of the operator $T_q$ which are usually studied in the literature.

exampleConsider the Hilbert space setting of Remark (ref) with conditions $(i)$ and $(ii)$. In addition assume that the reverse link condition $\int_0^1\|T_q \phi\|_W^2dq\geq c\sum_{j\geq1}\opw_j\Ex[\phi(Z)e_j(Z)]^2$ for $\phi\in\cB$ and some constant $c>0$ is satisfied. In the setting of Example (ref), we have $\int_0^1s_{qj}^2dq>\opw_j$ for all $j\geq 1$ implying that $T_q$ is nonsingular for almost all $0<q<1$ (since any countable union of null sets is null). For simplicity, let $Z$ and $W$ be scalars. Further, let $\max\big(n^{-1}\l, l_n^{-2\beta}\big)\sim n^{-1}\k$ and $\k\sim n^\chi$ for some constant $\chi>0$ which is specified in the following two cases. \begin{enumerate} • Mildly ill-posed case: If $\opw_\k\sim k_n^{-2\zeta}$ for some $\zeta\geq 0$ then in order for (ref) to hold we require $\m\sim n^\iota$ with $0<\iota<1/3$ and \begin{equation*} (1-\iota/2)/(2\alpha+2\zeta)<\chi<\iota/2. \end{equation*} Further, $\int_0^1\|\varPi_\k\solq-\solq\|_{Z}^2dq+\tau_\k\omega_n=O(k_n^{-2\alpha}+k_n^{2\zeta+1}n^{-1})$ which is $o(m_n^{-2/\kappa})$ if $\iota/(\alpha\kappa)<\chi<(1-2\iota/\kappa)/(2\zeta+1)$. Thus, condition (ref) is satisfied if \begin{equation*} \max\Big((1-\iota/2)/(2\alpha+2\zeta),\iota/(\alpha\kappa)\Big) <\chi<\min\Big(\iota/2, (1-2\iota/\kappa)/(2\zeta+1)\Big). \end{equation*} • Severely ill-posed case: If $\opw_\k\sim \exp\big(-k_n^{2\zeta}\big)$ for some $\zeta>0$ then $\int_0^1\|\varPi_\k\solq-\solq\|_{Z}^2dq+\tau_\k\omega_n=O(k_n^{-2\alpha}+\exp(k_n^{2\zeta})k_nn^{-1})$. Thereby, condition (ref) is satisfied if, for example, $\m=o\big((\log n)^{\alpha\kappa/\zeta}\big)$ and $\k\sim (\log n)^{1/\zeta}$. \end{enumerate} In both situations we conclude that the dimension parameter $\m$ is required to be larger than the dimension $\k$ of the sieve space for $n$ sufficiently large. Roughly speaking we require more moment restrictions implied by the instrument than the number of parameters we want to estimate. This corresponds to the test of overidentification in the parametric framework. $\square$

In contrast to a test integrated over all quantiles, one might be interested to check model (ref) for one specific quantile. In this case, we consider the test statistic

equation[equation omitted — 183 chars of source]

If $S_n(q)$ becomes too large then we reject the null hypothesis $H_0$. The derivation of the asymptotic behavior of $S_n(q)$ is similar as in Theorem (ref). Indeed, only the Lebesgue measure over $(0,1)$ has to be replaced by the Dirac measure which has its mass at the quantile of interest.

coroLet Assumptions (ref) and (ref) be satisfied. For a fixed quantile $q\in(0,1)$, let Assumptions (ref), (ref), and conditions (ref) and (ref) hold. If there exists a function $\sol_q\in\cB$ with $\cT\solq=q$ then \begin{equation*} (2\m)^{-1/2}\Big(\frac{1}{q(1-q)}\, S_n(q)-\m\Big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

In addition, one might be interested in certain regions of quantile functions. Let $\mu$ denote any measure on $(0,1)$. Again, the next result is a direct implication of Theorem (ref) and hence we omit its proof.

coroLet Assumptions (ref) and (ref) be satisfied. For all $q$ in the support of $\mu$, let Assumptions (ref), (ref), and conditions (ref) and (ref) hold. If there exists a function $\sol\in\cB$ with $\int|\cT\solq-q|d\mu(q)=0$ then \begin{equation*} \Big(2\m\int_0^1 (\min(q,q')-qq')^2d\mu(q)d\mu(q')\Big)^{-1/2}\Big( \int_0^1 S_n(q)d\mu(q)-\m\int_0^1 q(1-q)d\mu(q)\Big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

As mentioned in the introduction, our test is a joint test of instrument validity and monotonicity of $\sol$ in its second entry. The following remark illustrates how the test statistic $S_n(q)$ integrated over a subset of $(0,1)$ can be useful to detect which kind of deviation exists.

rem[Detecting the kind of deviation] Suppose that the structural function is strictly monotonically increasing in its second entry for values $q\in(0,q')$ given some $q'\in(0,1)$ (can be checked using Corollary (ref)). Further, let $q\mapsto\sol(\cdot,q)$ be either nonincreasing or decreasing on $(q',q'')$. This can be assured by letting $q''$ close to $q'$ and assuming that $\sol$ does not oscillate for $q\geq q'$. If $W$ is a valid instrument, employing model equation (ref) and $V\sim\cU(0,1)$ yields \begin{align*} \PP(Y\leq \sol(Z,q)|W)&=\PP(\sol(Z,V)\leq \sol(Z,q)|W)\\ &\leq\PP(V\leq q|W)\\ &=q \end{align*} for all $q\leq q''$ and $q''$ sufficiently close to $q'$. The last inequality holds regardless whether the function $q\mapsto\sol(\cdot,q)$ is strictly monotone or not. Consequently, if $\inf_{w\in\cW}\PP(Y\leq \sol(Z,q)|W=w)>q$ for some $q\in(q',q'')$ we may conclude that $W$ is not a valid instrument. The analysis of a one sided test based on this inequality is beyond the scope of this paper. On the other hand, we can check the kind of deviation by using the estimator $\inf_{w\in\cW} f_\umn(w)^t\big[n^{-1}\sum_{i=1}^n(\1\set{Y_i\leq \hsolq(Z_i)}-q)f_\umn(W_i)\big]$. Further, confidence statements can be achieved by using resampling methods. $\square$
rem[Implementation of the test statistic] This remark provides some details on the implementation of our test. First, discretize the $(0,1)$--integral by using the grid $1/N,2/N,\dots,(N-1)/N$ for some integer $N$. In different simulations, we found that a grid size of $N=20$ was sufficiently large. Also note that by the choice of the grid we avoid evaluation at boundary points zero or one. Second, for any integer $\m\leq n^{1/2}$ estimate the structural effect $\sol_q$ given in (ref) for each grid point $q$, each parameter $\k$ with $k_n^2\leq \m$ and $\l=2\k$. Third, compute the standardized test statistic $S_n$ such that it is maximized w.r.t. $\m$ and minimized w.r.t. $\k$. That is, we choose $\k$ to provide a good model fit and $\m$ to increase the power of the test. The choice of the dimension parameters capture essential rate requirements imposed to achieve asymptotic normality and is also motivated by simulation results. This leads to a so-called minimum-maximum principle, see also Subsection (ref) for more details. $\square$

Consistency against a fixed alternative

Let us first establish consistency when $H_0$ does not hold, that is, there exists no function $\sol$ belonging to $\cB^{(0,1)}$ which solves $\cT\solq=q$ for all $0<q<1$. The following proposition shows that our test has the ability to reject a false null hypothesis with probability $1$ as the sample size grows to infinity. In the following analysis of the asymptotic power of our testing procedure we let $\solq=\argmin_{\phi\in\cB}\|\cT\phi-q\|_W$. So if $H_0$ is false then $\int_0^1\|\cT\solq-q\|_W^2dq>0$ since $p_W$ is uniformly bounded from below.

propAssume that $H_0$ does not hold. Let Assumptions (ref)--(ref) be satisfied. Consider a sequence $(\gamma_n)_{n\geq 1}$ satisfying $\gamma_n=o(n/\sqrt\m)$. If conditions (ref) and (ref) hold we have \begin{gather*} \PP\Big(3\sqrt{5/\m}\big(S_n-\m/6\big)>\gamma_n\Big)=1+o(1). \end{gather*}

Limiting behavior under local alternatives

In the following, we study the power of the test, that is, the probability to reject a false hypothesis against a sequence of linear local alternatives that tends to zero as the sample size tends to infinity. We proceed similarly as Ait2001 (Section 3.3). More precisely, let $(\sol_{qn})_{n\geq 1}$ be a sequence of (nonstochastic) functions satisfying $n\int_0^1\|\cT\sol_{qn}-\cT\solq\|_W^2dq=o(\sqrt\m)$ where $\solq=\argmin_{\phi\in\cB}\|\cT\phi-q\|_W$. Then we consider alternative models defined by $\sol_{qn}$ with

equation[equation omitted — 168 chars of source]

Here, $\xi_q\in L_W^2$ is a function satisfying $\int_0^1\|\xi_q\|_W^2dq>0$. The next result establishes asymptotic normality for the standardized test statistic $S_n$.

propLet Assumptions (ref)--(ref) be satisfied. Assume that $(\sol_{qn})_{n\geq 1}$ satisfies (ref) and $n\int_0^1\|\cT\sol_{qn}-\cT\solq\|_W^2dq=o(\sqrt\m)$. If conditions (ref) and (ref) hold we have \begin{equation*} 3\sqrt{5/\m}\big( S_n-\m/6\big)\stackrel{d}{\rightarrow}\cN\Big(\sum_{j= 1}^\infty\int_0^1\Ex[\xi_q(W)f_j(W)]^2dq,1\Big). \end{equation*}

From Proposition (ref) we see that our test can detect local linear alternatives at the rate $\delta_n$. If $\{f_j\}_{j\geq 1}$ forms an orthonormal basis in $L_W^2$ then $\delta_n$ coincides with $m_n^{1/4}n^{-1/2}$ within a constant. Hence, our test has the same power against local linear alternatives as the test of Hong95 who consider parametric specification testing.

Inference based on bootstrap

Nonparametric tests that rely on the asymptotic normal approximation may perform poorly in finite samples. An alternative approach is to use bootstrap approximation. It is known that bootstrap based procedures could approximate finite sample distributions more accurately. In the following, we propose a bootstrap version of our test statistic $S_n$.

The bootstrap procedure is based on a sequence of independent and identically distributed random variables $\varepsilon_i$, $1\leq i\leq n$, drawn independently of the original data $(Y_i,X_i,W_i)$, $1\leq i\leq n$. Following chen2015 we then consider the bootstrap residual function

align*[align* omitted — 68 chars of source]

Let $\widehat \sol_{qn}^*$ be the bootstrap version of the sieve least squares estimator (ref), which is computed in the same way but where only $\big(\1{\{Y_i\leq \phi(Z_i)\}}-q\big)$ is replaced by $\varepsilon_{i}\big(\1{\{Y_i\leq \phi(Z_i)\}}-q\big)$. The bootstrap version $S_n^*$ of our test statistic $S_n$ given in (ref) builds on $\widehat \sol_{qn}^*$. More precisely, $S_n^*$ is computed as the test statistic $S_n$ but where only $\big(\1{\{Y_i\leq \hsolq(Z_i)\}}-q\big)$ is replaced by $\varepsilon_{i}\big(\1{\{Y_i\leq \widehat \sol_{qn}^*(Z_i)\}}-q\big)$.

assALet $(\varepsilon_i)_{i\geq 1}$ be an independent and identically distributed sequence of random variables drawn independently of $(Y,Z,W)$ such that $\Ex[\varepsilon]=1$, $\Var(\varepsilon)=:\sigma_\varepsilon^2\in(0,\infty)$ and $\Ex[|\varepsilon-1|^4]<\infty$

Assumption (ref) corresponds to Assumption Boot.1 of chen2015. We slightly strengthen their assumption by imposing a fourth moment restriction, which we require to derive asymptotic validity of the bootstrap procedure. Due to the bootstrap innovations $\varepsilon_i$ the constants in the sieve mean and sieve standard deviation change. For the bootstrap test $S_n^*$ we obtain the sieve mean constant

equation*[equation* omitted — 102 chars of source]

and the sieve standard deviation constant

equation*[equation* omitted — 178 chars of source]

chen2015 show that the bootstrap version of the sieve estimator $\widehat \sol_{qn}^*$ converges at the same rate as $\hsolq$. Thus, following line by line the proof of Theorem (ref) and using the imposed restrictions on the weights $\varepsilon_i$ we obtain the following result.

coroLet the assumptions of Theorem (ref) be satisfied. Under Assumption (ref) and null hypothesis $H_0$ we have \begin{equation*} 3\sqrt{5/(\m(\sigma_\varepsilon^2+1))}\big(S_n^*-\m(\sigma_\varepsilon^2+1)/6\big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

It should be emphasized the asymptotic validity of the bootstrap procedure is, in particular, due to the rate condition (ref), which ensures that the asymptotic distribution of $S_n^*$ is not affected by the estimation of the structural function. The next result establishes consistency of the bootstrap test against fixed alternatives.

coroAssume that $H_0$ does not hold and that the assumptions of Proposition (ref) are satisfied. Under Assumption (ref) we have \begin{gather*} \PP\Big(3\sqrt{5/(\m(\sigma_\varepsilon^2+1))}\big(S_n^*-\m(\sigma_\varepsilon^2+1)/6\big)>\gamma_n\Big)=1+o(1). \end{gather*}

Extensions

As we see in this section, our testing procedure can potentially be applied to a much wider range of situations. We now discuss corollaries that generalize the previous results in different ways. For the following analysis we focus on a fixed quantile $q\in(0,1)$.

Testing exogeneity

Falsely assuming exogeneity of the regressors leads to inconsistent estimators while on the other hand treating exogenous regressors as if they were endogenous can lower the rate of convergence dramatically. In this subsection, we develop a nonparametric test of exogeneity that is robust against possible nonseparability of unobservables. The test statistic is similar to the statistic $S_n(q)$ given in (ref) but where $\hsolq$ is replaced by an estimator of the conditional quantile function.

In contrast to the previous section, we assume here that there exists a unique function $\sol_q$ satisfying $Y= \solq(Z)+U_q$ with $\PP(U_q\leq 0|W)=q$ and for some $q\in(0,1)$. The relation between $Z$ and $W$ is thus restricted through this maintained hypothesis. Under the maintained hypothesis, we propose a test whether the vector of regressors $Z$ is exogenous at a quantile $q\in(0,1)$, that is,

equation*[equation* omitted — 54 chars of source]

In the following, we denote the conditional quantile function by $\sol^{\textsl e}_q$ which satisfies $\PP(Y\leq \sol^{\textsl e}_q(Z)|Z)=q$. The null hypothesis $H_0^{\textsl e}$ is satisfied if and only if the structural function $\solq$ coincides with the conditional quantile function $\sol_q^{\textsl e}$. Further, under nonsingularity of the operator $\cT$, hypothesis $H_0^{\textsl e}$ is equivalent to

equation[equation omitted — 41 chars of source]

Our test of exogeneity, which we propose below, is based on this equation or equivalently on $\PP(Y\leq \sol^{\textsl e}_q(Z)|W)=q$. More precisely, to test exogeneity we replace in the statistic $S_n(q)$ given in (ref) the estimator of $\solq$ by an estimator of $\sol^{\textsl e}_q$.

In the following, $\hsol^{\textsl e}_{qn}$ denotes an estimator for the conditional quantile function $\sol^{\textsl e}_q$. For instance, an estimator of $\sol^{\textsl e}_q$ is given by

equation[equation omitted — 123 chars of source]

where $\varrho_q(u)=|u|-(2q-1)u$ is the check function and here, $\cB_\k=\big\{\phi\in\cB:\,\phi(\cdot)=\sum_{j=1}^\k\beta_je_j(\cdot)\big\}$. For B-spline basis functions and an additional penalty this estimator was proposed by koenker1994. In the following, let $p_Z$ and $p_{Z|W}$ denote the marginal density of $Z$ and the conditional density of $Z$ given $W$, respectively.

assA(i) There exists a function $\solq\in\cB$ such that $\cT\solq=q$. (ii) $p_{Y|Z,W}(\cdot,Z,W)$ is continuously differentiable, $|\partial p_{Y|Z,W}(\cdot,Z,W)/\partial y|\leq C$ and $p_{Z|W}(\cdot,W)\leq Cp_{Z}(\cdot)$ for some constant $C>0$. (iii) There exists a sequence $(R_n^{\textsl e})_{n\geq 1}$ with $R_n^{\textsl e}=o(1)$ such that $\|\hsol_{qn}^{\textsl e}-\sol_q^e\|_Z^2=O_p(R_n^{\textsl e})$.

Assumption (ref) $(i)$ formalizes the maintained hypothesis of a correctly specified nonparametric instrumental quantile moment equation. Section (ref) provides a test for it. Due to Assumption (ref) $(ii) $ we do not require Assumption (ref) $(ii)$ but can rather rely on an upper bound of the Taylor reminder of $\cT$ obtained by chen2012. In this sense, the test of exogeneity presented below requires weaker restrictions on the local curvature of $\cT$ than in the case of specification testing. Assumption (ref) specifies a rate requirement for the $L_Z^2$ distance of the estimator $\hsol_{qn}^{\textsl e}$. For instance, under $H_0^{\textsl e}$, Assumption (ref) $(iii)$ is satisfied with $R_n^{\textsl e}=\k/n+k_n^{-2r}$ when $\hsol_{qn}^{\textsl e}$ is given by the estimator (ref) with the B-splines basis functions $\{e_j\}_{j\geq 1}$ and $Z$ is scalar, see he1994. The same rate is obtained by horowitz2005 in the case of multivariate $Z$ in an additive quantile regression model.

For a test of the null hypothesis $H_0^{\textsl e}$ we replace in the definition of $S_n(q)$ given in (ref) the estimator $\hsolq$ by $\hsol^{\textsl e}_{qn}$. That is,

equation*[equation* omitted — 203 chars of source]

We reject the hypothesis $H_0^{\textsl e}$ if $S_n^{\textsl e}(q)$ becomes too large. The next result establishes asymptotic normality of our test statistic $S_n^{\textsl e}(q)$ under the null hypothesis.

coroLet Assumptions (ref), (ref) $(i)$, (ref), (ref) and (ref) hold. Let $\m$ satisfy condition (ref). Consider the estimator $\hsolq^{\textsl e}$ given in (ref) where $\k$ satisfies \begin{equation} n R_n^{\textsl e}=o(\sqrt{\m})\quad and\quad R_n^{\textsl e}=o\big(m_n^{-(1+\epsilon)/\kappa}\big) \end{equation} for some $\epsilon>0$. Then we have under $H_0^{\textsl e}$ \begin{equation*} \big(2\m\big)^{-1/2}\Big(\frac{1}{q(1-q)}\,S_n^{\textsl e}(q)-\m\Big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}
exampleLet us illustrate when condition (ref) holds true. Let $\m\sim n^\iota$ with $0<\iota<1/3$. Then for (ref) to hold let $\k\sim n^\chi$ where $\chi>0$ satisfies \begin{equation*} \max\Big(\frac{1-\iota/2}{2r},\,\frac{\iota}{r\kappa}\Big)<\chi<\min\Big(\frac{\iota}{2}, 1-\frac{2\iota}{\kappa}\Big). \end{equation*} Hence, we require $r>2/\kappa$ which is a slightly stronger restriction than Assumption (ref) $(i)$. $\hfill\square$

In the following, we study the power of the test, that is, the probability to reject a false hypothesis against a sequence of linear local alternatives that tends to zero as the sample size tends to infinity. More precisely, let $(\sol_{qn}^{\textsl e})_{n\geq 1}$ be a sequence of (nonstochastic) functions satisfying

equation[equation omitted — 173 chars of source]

Here, $\xi_q^{\textsl e}\in L_W^2$ is a function satisfying $\|\xi_q^{\textsl e}\|_W^2>0$. The next result establishes asymptotic normality for the standardized test statistic $S_n^{\textsl e}(q)$.

coroLet Assumptions (ref), (ref) $(i)$, (ref), (ref), and (ref) be satisfied. Assume that $(\sol_{qn}^e)_{n\geq 1}$ satisfies (ref). If condition (ref) holds true we have \begin{equation*} \big(2\m\big)^{-1/2}\Big(\frac{1}{q(1-q)}\,S_n^{\textsl e}(q)-\m\Big)\stackrel{d}{\rightarrow}\cN\Big(\sum_{j= 1}^\infty\int_0^1\Ex[\xi_q^{\textsl e}(W)f_j(W)]^2dq,1\Big). \end{equation*}

Testing additivity

The test statistic given in (ref) is also convenient to check additional restrictions on the structural effect $\solq$ for $0<q<1$. These additional restrictions can be easily imposed by constraints on the functions of the sieve space $\cB_\k$. For instance, one may impose an additive structure of the quantile structural effects.

By assuming an additive structure of $\solq$ one might reduce the effect of dimensionality of the regressors on the convergence rate of an estimator (cf. chen08 in case of instrumental quantile regression). Applying this structure leads, however, to inconsistent estimators in general if the function $\solq$ does not obey an additive form. Our aim in the following is to test whether

equation*[equation* omitted — 156 chars of source]

Similarly as above we obtain the test statistic

equation*[equation* omitted — 202 chars of source]

Here the estimator $\hsolq^{\textsl add}=(\hsol_{qn}^1,\hsol_{qn}^2)$ of $\solq=(\sol^1_q,\sol^2_q)$ is given by (ref) where the sieve basis is a tensor product of basis functions that depend either on $Z'$ or $Z''$. For a more detailed discussion we refer to Section 6 of chen08. The next asymptotic normality result is a direct consequence of Corollary (ref) and hence its proof is omitted.

coroGiven the conditions of Corollary (ref) we have under $H_0^{add}$ \begin{equation*} \big(2\m\big)^{-1/2}\Big(\frac{1}{q(1-q)}\,S_n^{add}(q)-\m\Big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

Monte Carlo simulation

In this section, we study the finite sample performance of our test by presenting the results of a Monte Carlo investigation. There are $1000 $ Monte Carlo replications in each experiment. Results are presented for the nominal levels $0.05$. Let $\Phi$ denote the cumulative standard normal distribution function. Throughout this simulation study, realizations $(Z,W)$ were generated by $Z = \Phi\big(\zeta\omega +\sqrt{1-\zeta^2}\,\varepsilon\big)$ and $W =\Phi(\omega)$ where $\omega$ is independent of $\varepsilon$ and $\omega,\,\varepsilon\sim \cN(0, 1)$. Here, the constant $\zeta>0$ determines the degree of correlation between $Z$ and $W$ and is varied in the experiments.

Testing a Nonparametric Specification

We begin with the finite sample analysis of our test statistics in case of nonparametric specification testing. To analyze the finite sample power we distinguish in the following between a failure of the null hypothesis caused either by a lack of instrument validity or by non-monotonicity of the structural function in unobservables. \paragraph{Failure of instrument validity.} We first generate realizations of $Y$ under the null hypothesis $H_0$. Recall that under $H_0$ there exists a function $\sol\in\cB^{(0,1)} \text{ such that } \PP(Y\leq \sol(Z,q)|W)=q \text{ for all }q\in(0,1)$. In the following finite sample analysis, we restrict $\cB^{(0,1)}$ to contain continuously differentiable functions only. Under $H_0$ we generate realizations of $Y$ from the nonseparable model

equation[equation omitted — 57 chars of source]

where $V = \vartheta\,\varepsilon + \sqrt{1-\vartheta^2}\,\epsilon$ with $\epsilon\sim \cN(0, 1)$ independent of $(\omega,\varepsilon)$ and $\vartheta=0.7$. We consider the function $\phi(z)=\sum_{j=1}^\infty \,j^{-4} \cos(j\pi z)$. For computational reasons we truncate the infinite sum at $100$. The resulting function is displayed in Figure (ref). Since $\phi$ is continuously differentiable the null hypothesis $H_0$ is satisfied with $\sol(z,q)=\phi(z)\big(1+F_V^{-1}(q)/6\big) + F_V^{-1}(q)/2$, where $F_V^{-1}$ denotes the quantile function of $V$.

figure[figure omitted — 154 chars of source]
table[table omitted — 2,555 chars of source]

When $H_0$ is false we generate realizations of $Y$ from

equation[equation omitted — 73 chars of source]

where $\rho_j(z)=10\,j\,(z\1\{z\leq 0.25\}+(z-1)\1\{z>0.25\})$ for $j=1,2$ and $\rho_j(z)=(z/2c_j)\1\{0.5-c_j\leq z<0.5+c_j\}$ for $j=3,4$, with $c_3=0.1$ and $c_4=0.05$. Here, the variable $V$ is generated as in (ref). Under (ref), the structural function $\sol$ satisfying the quantile restriction $\PP(Y\leq \sol(Z,q)|W)=q$ is given by $\sol(z,q)=(\phi(z)+\rho_j(z))(1+F_V^{-1}(q)/6) + F_V^{-1}(q)/2$. So $\sol(\cdot,q)$ is not continuously differentiable and thus, $H_0$ is false. Due to the ill-posed inverse problem estimation of $\sol(\cdot,q)$ we cannot choose $\k$ sufficiently large to capture such irregularities which implies finite sample power of our test against those alternatives. This corresponds to the analysis of Horowitz2011 in the instrumental mean regression case.

For each quantile $0<q<1$, we estimate the structural function using the estimator $\hsol_{qn}$ given in (ref) with B-splines as approximation basis functions. More precisely, for the sieve space $\cB_\k$ we use B-splines of order 2 with 1 knot or 2 knots (hence $\k=4$ or $\k=5$) and for the criterion function we use B-splines of order 2 with 5 knots or 7 knots (hence $\l=2\k$), respectively. We thus follow chen2015optimal and choose $\l$ to be a constant multiple of $\k$. Also for the vector of basis functions $f_\umn$, used to construct the test statistic, we use B-spline basis of order 2 with knots varying between 17, 22 or 27 (hence $\m=20$, $\m=25$ or $\m=30$).

The empirical rejection probabilities of our standardized test statistic $3\sqrt{5/\m}\big(S_n-\m/6\big)$ at nominal level $0.05$ are shown in Table (ref). We approximate the integral over the quantiles on $(0,1)$ by the mean of a random sample from the uniform $(0,1)$ distribution. As we see from Table (ref), our test is less sensitive with respect to the choice of $\m$ than to the choice of $\k$, which is not surprising and well known from nonparametric instrumental variable estimation problems, see also chen2015. Table (ref) shows the empirical rejection probabilities for the sample sizes $500$ and $1000$. We see that as the sample size increases the finite sample rejection probabilities become larger in the alternative models. For $\k=4$ we see that the finite sample coverage improves slightly as the sample size increases. This is not the case for $\k=5$ which appears to be an inappropriate choice implying a large variance.

In Table (ref) we also compare our testing procedure to a bootstrap version of it. We consider the generalized residual bootstrap as proposed in Subsection (ref). We generate the bootstrap weights by $\varepsilon\sim\mathcal N(1,\sigma_\varepsilon^2)$, independently of $(Y,X,W)$, where $\sigma_\varepsilon=0.5$. We run $200$ bootstrap evaluations per Monte Carlo replication. We see from Table (ref) that the bootstrap leads to an improvement in the finite sample coverage in the true model. In this sense, the bootstrap test statistic is less sensitive to the choice of $\k$ under the true model. Similar to chen2015 (see p. 1059), we see only a minor improvement of the bootstrap test in the alternative models but we expect that it improves further as the number of bootstrap runs is increased.

As we fix the dimension parameter $\l=2\k$, two dimension parameters remain to be chosen by the econometrician, namely, $\k$ and $\m$. While proposing an adaptive testing procedure is beyond the scope of this paper, we want to provide an heuristic argument for the parameter choice. Intuitively, we want to choose $\k$ such that we have a good model fit, i.e., a small value of the test statistic, and $\m$ to have good power properties, i.e., a large value of the test statistics. Moreover, the choice should reflect the rate requirement from our theory, that is, $\k\leq \l =o(m_n^{1/2})$ and $m_n=o(n^{1/2})$. We implement such a heuristic parameter choice criterion via the following minimum-maximum principle. That is, if $\set{s(\k,\m)}$ denotes the standardized value of our test $S_n$ with dimension parameters $\k$ and $\m$ then we choose these parameters such that

align*[align* omitted — 78 chars of source]

The values of this minimum-maximum principle (over the range $\m\in\set{20,25,30}$ and $\k\in\set{4,5}$) are shown in bold in Table (ref). Note that the requirement $\k<n^{1/4}$ implies $\k\leq 4$ when $n=500$ and $\k\leq 5$ when $n=1000$. Further, $m_n<n^{1/2}$ implies $m_n\leq 22$ for $n=500$ and $m_n\leq 31$ for $n=1000$. We see that this criterion helps to avoid choosing the dimension parameter $\k$ too large which would yield inaccurate coverage. Such a rule, however, does not account for ill-posedness of the estimation problem and hence, $\k$ might still be chosen too large. We thus could calculate the sieve measure of ill-posedness by estimating the first $\k$ minimal eigenvalues of $T_q$ (see also chen2015). \paragraph{Failure of monotonicity in unobservables.} We study the finite sample power of our test when $\sol$ is not strictly monotonic in the structural disturbance $V$. Realizations of $Y$ were generated from

equation[equation omitted — 51 chars of source]

where $V = \Phi\big((\vartheta\,\varepsilon + \sqrt{1-\vartheta^2}\,\epsilon)/4\big)$ with $\epsilon\sim \cN(0, 1)$ and where $\vartheta=0.8$.

table[table omitted — 2,913 chars of source]

When $H_0$ is false we generate

equation[equation omitted — 66 chars of source]

or

equation[equation omitted — 67 chars of source]

for $j=1,2$. In the alternative models, the structural disturbance enters the model in a nonmonotonic way. We construct the statistic $S_n$ and its bootstrap counterpart $S_n^*$ as described in the previous paragraph.

Table (ref) depicts the empirical rejection probabilities of our test against the alternative models (ref) and (ref). Again we observe that our test is not very sensitive to the choice of the dimension parameter $\m$. Our test becomes somewhat less powerful for large $\k$. But in contrast to the alternatives involving discontinuous functions in the previous paragraph, the choice of $\k$ is not as sensitive. For each choice of parameter $\k$, our test becomes more powerful as the sample size increases from $500$ to $1000$. For $n=1000$ we see that the parameter choice $\k=5$ leads to a more accurate finite sample coverage. This is captured by the minimum-maximum principle as introduced above. Again, the resulting values of the test statistic using this criterion over the range $\m\in\set{20,25,30}$ and $\k\in\set{4,5}$ are shown in bold. Again we observe that the boostrap version of the test statistic behaves similarly as the statistic $S_n$.

Testing exogeneity

Realizations $Y$ were generated by

equation*[equation* omitted — 45 chars of source]

where $V$ is generated as described in model (ref), that is, $V = \vartheta\,\varepsilon + \sqrt{1-\vartheta^2}\,\epsilon$ with $\epsilon\sim \cN(0, 1)$ independent of $(\omega,\varepsilon)$. The function $\sol^{\textsl e}$ is given by $\sol^{\textsl e}(z)=\sum_{j=1}^\infty (-1)^{j+1}\,j^{-2} \sin(j\pi z)$. Again, for computational reasons we truncate the infinite sum at $100$. The resulting function is displayed in Figure (ref). Note that $\vartheta$ determines the degree of endogeneity of $Z$ and is varied among the experiments. The null hypothesis $H_0:\PP(Y\leq \sol^{\textsl e}(Z)|Z)=q$ holds true if $\vartheta=0$ and is false otherwise. In the following, we perform a test at the median $q=0.5$. As our test relies on the equation $\PP(Y\leq \sol^{\textsl e}(Z)|W)=q$ we expect our test to have more power as the correlation between $W$ and $Z$ increases.

The test statistic is implemented as described in Section (ref). To estimate the structural effect we make use of the estimator $\hsol^{\textsl e}_{qn}$ of he1994 given in (ref). Here, we use B-splines of order 2 with 1 knot (hence $\k=4$) or 2 knots (hence $\k=5$). In contrast to the previous section, the choice of the dimension parameter $\k$ is not affected by the ill-posedness of the underlying inverse problem. As above, the vector of basis functions $f_\umn$ is also constructed with B-spline basis of order 2 with knots varying between 17, 22 or 27 (hence $\m=20$, $\m=25$ or $\m=30$).

Table (ref) depicts the empirical rejection probabilities with varying number of basis functions. As we see from Table (ref), our test becomes more powerful for larger $\zeta$; that is, for instruments with a stronger correlation to the covariates $Z$. From Table (ref) we see that the test of exogeneity becomes somewhat less powerful for larger values of $\m$. On the other hand, the test seems not to be too sensitive with respect to the choice of the dimension parameters $\k$ and $\m$. We also see from Table (ref) that the finite sample coverage and power properties of the test improve as the sample size increases from $500$ to $1000$.

Similarly as above, a guideline for smoothing parameter choice in practice is given by the following minimum-maximum principle. That is, if $\set{s_q^{\textsl e}(\k,\m)}$ denotes the standardized value of our test $S_n^{\textsl e}(q)$ with dimension parameters $\k$ and $\m$ then choose these parameters such that

align*[align* omitted — 92 chars of source]

Again this criterion takes the rate condition for the asymptotic theory into account. In Table (ref) the resulting values of the test statistic using this criterion over the range $\m\in\set{20,25,30}$ and $\k\in\set{4,5}$ are shown in bold.

table[table omitted — 2,351 chars of source]

An empirical illustration

To illustrate our testing procedure, we present an empirical application concerning estimation of the effects of class size on students' performance on standardized tests. Angrist1999 studied the effects of class size on test scores of 4th and 5th grade students in Israel. In this empirical illustration, we focus on 4th grade reading comprehension a feature that was also considered by Horowitz2011.

In this empirical example we study the model

equation[equation omitted — 89 chars of source]

where $Y_{sc}$ is the average reading comprehension test score of 4th grade students in class $c$ of school $s$, $Z_{sc}$ is the number of students in class $c$ of school $s$, $D_{sc}$ is the fraction of disadvantaged students in class $c$ of school $s$ with unknown scalar function $\beta$, $V_{sc}=U_{s}+\varepsilon_{sc}$ where $U_{s}$ is an unobserved school-specific random effect, and $\varepsilon_{sc}$ is an unobserved, independently over classes and schools distributed random variable.

The class size $Z_{sc}$ may be endogenous, for instance, due to the socioeconomic background of the students. To identify the causal effect of class size on scholar achievement Angrist1999 use Maimonides' rule as instruments. According to this administrative rule, maximum class size is given by 40 pupils and will be split if the number of enrolled students exceeds this number. More precisely, assuming that cohorts are divided into classes of equal size, Maimonides’ rule is described by

equation*[equation* omitted — 59 chars of source]

where $E_s$ denotes enrollment in school $s$ and $\lceil x\rceil$ denotes the largest integer less or equal to $x$. Note that Horowitz2011 could show that a linear relation between class size and scholar achievement as used by Angrist1999 is misspecified. To apply our tests, we consider a subsample where only one representative class per school is considered. By doing so, we avoid that rejection of a hypothesis may be caused by within class correlation. Moreover, only schools with at least two classes are considered which leads to a sample size of 707.

In the following, we want to test nonparametrically whether class size is endogenous at the $0.5$--quantile. The null hypothesis is that $\PP(Y_{sc}\leq \varphi(Z_{sc},q)+D_{sc}\,\beta(q)|Z_{sc})= q$ where $q=0.5$. The value of our test statistic $S_n^{\textsl e}(0.5)=(2\m)^{-1/2}\big(4\, S_n^{\textsl e}(0.5)-\m\big)$ is given by $1.885$. For the choice of smoothing parameters $\k$ and $\m$ we applied the minimum-maximum principle as described in Section (ref). The resulting dimension parameters are $\k= 4$ and $\m=23$.\footnote{The value of the test for other choices of $\k$ is $2.254$ for $\k=3$ and $2.182$ for $\k=5$ where $\l=2\k$ and $\m$ is maximized over the range $k_n^2$ to $26$ (being the largest integer smaller than $\sqrt{707}$). } We thus reject the hypothesis of exogeneity at the $0.05$ nominal level. In particular, in model (ref) under conditions (a.1)--(a.3) we conclude that $Z_{sc}$ is not independent of $V_{sc}$.

We now test whether the model (ref) with conditions (a.1)--(a.3) is correctly specified. We construct our test statistic using B-splines as described in Section (ref). For the choice of smoothing parameters $\k$ and $\m$ we applied the minimum-maximum principle as described in Section (ref). As in the Monte Carlo section we choose $\l=2\k$. Our test statistic attains the value $ 1.4152$ and thus fails to reject the nonseparable model (ref) with conditions (a.1)--(a.3) at the $0.05$ nominal level. This value of the test statistic is obtained when $\k=4$ and $\m=26$. For the fixed quantile $q=0.5$, we also performed a test of $\PP(Y_{sc}\leq \varphi(Z_{sc},q)+D_{sc}\,\beta(q)|W_{sc})= q$. In this case, our test statistic attains the value $0.981$ and again fails to reject the hypothesis.\footnote{This is not the case if $\k$ is chosen too small or too large. For instance if $\k=4$ or $\k=9$, respectively, then the value of the test statistic is $2.064$ or $3.420$ (as above maximized of $\m$ and $\l=2\k$). }

figure[figure omitted — 243 chars of source]

For the full sample, Figure (ref) depicts estimators of the structural effect $\solq$ for the quantiles $q\in\{0.75,\,0.5,\,0.25\}$ where the number of disadvantaged students is restricted to be smaller than 15% (which implies $n=688$). The solid lines are the estimators and the dashed lines are the 90% pointwise bootstrap confidence intervals using 1000 bootstrap iterations (we account for within school correlation by using schools as the bootstrap sampling units, see also Horowitz2011). We can see that the confidence intervals are tight enough to reject the hypothesis that the quantile structural effects are overall upward sloping. In particular, we see that the effect of class size variation on test scores is more severe for lower performing classes.

Conclusion

In this paper, we develop a nonparametric specification test for the quantile regression model (ref). The power of the test derives either from violations of regularity conditions imposed on the structural function, such as bounds or smoothness requirements, or a failure of monotonicity in the nonseparable unobservable variable. The test statistic is easy to implement and a natural extension of specification testing in a parametric framework. As the test builds on the sieve methodology, it allows to incorporate restrictions under the null hypothesis directly on the sieve space. As examples of tests of constraint hypotheses we consider in detail a test of exogeneity and a test of additivity of the structural function. We establish the large sample behavior of our test statistics and show that our tests work well in finite sample experiments. We also obtain reasonable results in an empirical illustration concerning the analysis of class size on students' performance. While we provide some heuristic guideline how to choose the sieve dimension in finite samples, an interesting future research area remains to provide asymptotic justification for it via adaptive testing.