EconBase
← Back to paper

Goodness-of-Fit Tests based on Series Estimators in Nonparametric Instrumental Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,073 characters · 18 sections · 51 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Goodness-of-Fit Tests based on Series Estimators in Nonparametric Instrumental Regression

\vskip 1cm

abstractThis paper proposes several tests of restricted specification in nonparametric instrumental regression. Based on series estimators, test statistics are established that allow for tests of the general model against a parametric or nonparametric specification as well as a test of exogeneity of the vector of regressors. The tests' asymptotic distributions under correct specification are derived and their consistency against any alternative model is shown. Under a sequence of local alternative hypotheses, the asymptotic distributions of the tests is derived. Moreover, uniform consistency is established over a class of alternatives whose distance to the null hypothesis shrinks appropriately as the sample size increases. A Monte Carlo study examines finite sample performance of the test statistics.
tabbingKeywords: \=Nonparametric regression, instrument, linear operator, orthogonal series\\ \> estimation, hypothesis testing, local alternative, uniform consistency.\\[.1ex] \noindentJEL classification: C12, C14.

Introduction

While parametric instrumental variables estimators are widely used in econometrics, its nonparametric extension has not been introduced until the last decade. The study of nonparametric instrumental regression models was initiated by F03eswc and NP03econometrica. In these models, given a scalar dependent variable $Y$, a vector of regressors $Z$, and a vector of instrumental variables $W$, the structural function $\sol$ satisfies

equation[equation omitted — 90 chars of source]

for an error term $U$. Here, $Z$ contains potentially endogenous entries, that is, $\Ex[U|Z]$ may not be zero. Model (ref) does not involve the a priori assumption that the structural function is known up to finitely many parameters. By considering this nonparametric model, we minimize the likelihood of misspecification. On the other hand, implementing the nonparametric instrumental regression model can be challenging.

Nonparametric instrumental regression models have attracted increasing attention in the econometric literature. For example, AC03econometrica, BCK07econometrica, ChenReiss2008, NP03econometrica or Johannes2009 consider sieve minimum distance estimators of $\sol$, while DFR02, HallHorowitz2005, GS06 or FJVB07 study penalized least squares estimators. When the methods of analysis are widened to include nonparametric techniques, one must confront two mayor challenges. First, identification in model (ref) requires far stronger assumptions about the instrumental variables than for the parametric case (cf. NP03econometrica). Second, the accuracy of any estimator of $\sol$ can be low, even for large sample sizes. More precisely, ChenReiss2008 showed that for a large class of joint distributions of $(Z,W)$ only logarithmic rates of convergence can be obtained. The reason for this slow convergence is that model (ref) leads to an inverse problem which is ill posed in general, that is, the solution does not depend continuously on the data.

In light of the difficulties of estimating the nonparametric function $\sol$ in model (ref), the need for statistically justified model simplifications is paramount. We do not face an ill posed inverse problem if a parametric structure of $\sol$ or exogeneity of $Z$ can be justified. If these model simplifications are not supported by the data, one might still be interested in whether a smooth solution to model (ref) exists and if some regressors could be omitted from the structural function $\sol$. These model simplifications have important potential since they might increase the accuracy of estimators of $\sol$ or lower the required conditions imposed on the instrumental variables to ensure identification.

In this work we present a new family of goodness-of-fit statistics which allows for several restricted specification tests of the model (ref). Our method can be used for testing either a parametric or nonparametric specification. In addition, we perform a test of exogeneity and of dimension reduction of the vector of regressors $Z$, that is, whether certain regressors can be omitted from the structural function $\sol$. By a withdrawal of regressors which are independent of the instrument, identification in the restricted model might be possible although $\sol$ is not identified in the original model (ref).

There is a large literature concerning hypothesis testing of restricted specification of regression. In the context of conditional moment equation, Donald2003 and Tripathi2003 make use of empirical likelihood methods to test parametric restrictions of the structural function. In addition, Santos12 allows for different hypothesis tests, such as a test of homogeneity. Based on kernel techniques, Horowitz2006, Blundell2007, and Horowitz2011 propose test statistics in which an additional smoothing step (on the exogenous entries of $Z$) is carried out. Horowitz2006 considers a parametric specification test. Blundell2007 establish a consistent test of exogeneity of the vector of regressors $Z$, whereas Horowitz2011 tests whether the endogenous part of $Z$ can be omitted from $\sol$. GS07 and Horowitz2009 develop nonparametric specification tests in an instrumental regression model. We like to emphasize that their test cannot be applied to model (ref) where some entries of $Z$ might be exogenous.

Our testing procedure is entirely based on series estimation and hence is easy to implement. We use approximating functions to estimate the conditional moment restriction implied by the model (ref) where $\sol$ is replaced by an estimator under each conjectured hypothesis. It is worth noting that by our methodology we can omit some assumptions typically found in related literature, such as smoothness conditions on the joint distribution of $(Z,W)$. In addition, a Monte Carlo indicates that the finite sample power of our tests exceed that of existing tests.

The paper is organized as follows. In Section 2, we start with a simple hypothesis test, that is, whether $\sol$ coincides with a known function $\sol_0$. We obtain the test's asymptotic distribution under the null hypothesis and its consistency against any fixed alternative model. Moreover, we judge its power by considering linear local alternatives and establish uniform consistency over a class of functions. In Sections 3--5 we consider a parametric specification test, a test of exogeneity, and a nonparametric specification test. The goodness-of-fit statistics are obtained by replacing $\sol_0$ in the statistic of Section 2 by an appropriate estimator. In each case, the asymptotic distribution under correct specification and power statements against alternative models are derived. In Section 6, we investigate the finite sample properties of our tests by Monte Carlo simulations. All proofs can be found in the appendix.

A simple hypothesis test

In this section, we propose a goodness-of-fit statistic for testing the hypothesis $ H_0:\,\sol=\sol_0$, where $\sol_0$ is a known function, against the alternative $\sol\neq\sol_0$. We develop a test statistic based on $\cL^2$ distance. As we will see in the following chapters, it is sufficient to replace $\sol_0$ by an appropriate estimator to allow for tests of the general model against other specifications. We first give basic assumptions, then obtain the asymptotic distribution of the proposed statistic, and further discuss its power and consistency properties.

Assumptions and notation.

\paragraph{The model revisited} The nonparametric instrumental regression model (ref) leads to a linear operator equation. To be more precise, let us introduce the conditional expectation operator $T\phi:=\Ex[\phi(Z)|W]$ mapping $\cL^2_Z=\{\phi:\,\Ex|\phi(Z)|^2<\infty\}$ to $\cL^2_W=\{\psi:\, \Ex|\psi(W)|^2<\infty\}$. Consequently, model (ref) can be written as

equation[equation omitted — 46 chars of source]

where the function $g:=\Ex[Y | W]$ belongs to $\cL^2_W$. Throughout the paper we assume that an iid. $n$-sample of $(Y,Z,W)$ from the model (ref) is available. \paragraph{Assumptions.} Our test statistic based on a sequence of approximating functions $\{\basW_l\}_{l\geq 1}$ in $\cL^2_W$. Let $\cW$ denote the support of $W$ and the marginal density of $W$ by $p_W$. Let $\nu$ be a probability density function that is strictly positive on $\cW$. We assume throughout the paper that $\{\basW_l\}_{l\geq 1}$ forms an orthonormal basis in $\cL_\nu^2(\mathbb R^{d_w}):=\{\phi:\,\int \phi^2(s)\nu(s)ds<\infty\}$ where $d_w$ denotes the dimension of $W$. For instance, if $\cW\subset[a,b]$ then a natural choice of $\nu$ would be $\nu(w)=1/(b-a)$ for $w\in[a,b]$ and zero otherwise.

assAThere exist constants $\eta_f, \eta_p\geqslant 1$ such that (i) $ \sup_{l\geq 1}\int|f_l(s)|^4\nu(s)ds\leqslant \eta_f$ and (ii) $\sup_{w\in\cW}\big\{ p_W(w)/\nu(w)\big\}\leq \eta_p$ with $\nu$ being strictly positive on $\cW$.

Assumption (ref) $(i)$ restricts the magnitude of the approximating functions $\{f_j\}_{j\geq1}$ which is necessary for our proof to determine the asymptotic behavior of our test statistic. This assumption holds for sufficiently large $\eta_f$ if the basis $\{f_l\}_{l\geq1}$ is uniformly bounded, such as trigonometric bases. Moreover, Assumption (ref) $(i)$ is satisfied by Hermite polynomials. Assumption (ref) $(ii)$ is satisfied if, for instance, $p_W/\nu$ is continuous and $\cW$ is compact.

The results derived below involve assumptions on the conditional moments of the random variables $U$ given $W$ gathered in the following assumption.

assAThere exists a constant $\sigma>0$ such that $\Ex[U^4|W]\leq \sigma^4$.

The conditional moment condition on the error term $U$ helps to establish the asymptotic distribution of our test statistics. The following assumption ensures identification of $\sol$ in the model (ref).

assAThe conditional expectation operator $T$ is nonsingular.

Under Assumption (ref), the hypothesis $H_0$ is equivalent to $g=T\sol_0$ which is used to construct our test statistic below. Note that the asymptotic results under null hypotheses considered in Sections 2--4 hold true even if $T$ is singular. If Assumption (ref) fails, however, our test has no power against alternative models whose structural function satisfies $\sol=\sol_0+\delta$ with $\delta$ belonging to the null space of $T$.

We will see below that the power of our test can be increased by carrying out an additional smoothing step. Therefore, we introduce a smoothing operator $L$ mapping $\cL_W^2$ to $\cL_W^2$. In contrast to the unknown conditional expectation operator $\op$, which has to be estimated, the operator $L$ can be chosen by the econometrician. Let $L$ have an eigenvalue decomposition given by $\{\tau_j^{1/2},f_j\}_{j\geq 1}$. We allow in this paper for a wide range of smoothing operators. In particular, $L$ may be the identity operator, that is, no smoothing step is carried out. We only require the following condition on the operator $L$ determined by the sequence of eigenvalues $\tau=(\tau_j)_{j\geq 1}$.

assAThe weighting sequence $\tau$ is positive, nonincreasing, and satisfies $\tau_1=1$.

Assumption (ref) ensures that the operator $L$ is nonsingular.

remHorowitz2006, Blundell2007, and Horowitz2011 consider as a smoothing operator a Fredholm integral operator, that is, $L\phi(s)=\int_0^1\ell(s,t)\phi(t)dt$ for some function $\phi\in \cL^2[0,1]=\{\phi:\,\int_0^1\phi^2(s)ds<\infty\}$ and some kernel function $\ell:[0,1]^2\to\mathbb{R}$. In order to ensure $L\phi\in \cL^2[0,1]$ it is sufficient to assume $\int_0^1\int_0^1|\ell(s,t)|^2dsdt<\infty$. Let $\{\tau_j^{1/2},f_j\}_{j\geq 1}$ be the eigenvalue decomposition of $L$. By Parseval's identity \begin{equation*} \int_0^1\int_0^1|\ell(s,t)|^2dsdt=\int_0^1\sum_{j=1}^\infty\tau_j|f_j(s)|^2ds=\sum_{j=1}^\infty\tau_j \end{equation*} where the right hand side is only finite if the sequence $\tau$ decays sufficiently fast. In our case, if we apply a smoothing operator $L$ with $\sum_{j=1}^\infty\tau_j<\infty$ then our test statistics converges also to a weighted series of chi-squared random variables. In addition, we allow for a milder degree of smoothing or no smoothing at all and show below that then asymptotic normality of our test statistics can be obtained. $\square$

\paragraph{Notation.} For a matrix $A$ we denote its transposed by $A^t$, its inverse by $A^{-1}$, and its generalized inverse by $A^{-}$. The euclidean norm is denoted by $\|\cdot\|$ which in case of a matrix denotes the spectral norm, that is $\|A\|=(\text{trace}(A^tA))^{1/2}$. The norms on $L_Z^2$ and $L_W^2$ are denoted by $\|\phi\|_Z^2:=\Ex|\phi(Z)|^2$ for $\phi\in L_Z^2$ and $\|\psi\|_W^2:=\Ex|\psi(W)|^2$ for $\psi\in L_W^2$. The $k\times k$ identity matrix is denoted by $I_k$. For a vector $V$ we write diag$(V)$ for the diagonal matrix with diagonal elements being the values of $V$. Moreover, $\basZ_\um(Z)$ and $\basW_\um(W)$ denote random vectors with entries $\basZ_j(Z)$ and $\basW_j(W)$, $1\leq j\leq m$, respectively. For any weighting sequence $w$ we introduce vectors $\basZ_\um^w(Z)$ and $\basW_\um^w(W)$ with entries $\basZ_j^w(Z)=\sqrt{w_j}e_j(Z)$ and $\basW_j^w(W)=\sqrt{w_j}f_j(W)$, $1\leq j\leq m$. We write $a_n\sim b_n$ when there exist constants $c,c' > 0$ such that $c b_n\leq a_n\leq c' b_n$ for all sufficiently large $n$.

The test statistic and its asymptotic distribution

Nonsingularity of the conditional expectation operator $T$ and the smoothing operator $L$ implies that the null hypothesis $H_0$ is equivalent to $L(g-T\sol_0)=0$. Note that $\|L(g-T\sol_0)\|_W=0$ if and only if $\int\big|L(g-T\sol_0)(w)p_W(w)/\nu(w)\big|^2\nu(w)dw=0$ since the Lebesgue measure $\nu$ is strictly positive on $\cW$. Moreover, since $\{f_j\}_{j\geq 1}$ is an orthonormal basis with respect to $\nu$ we obtain by Parseval's identity

equation[equation omitted — 122 chars of source]

Now we truncate the infinite sum at some integer $\m$ which grows with the sample size $n$. This ensures consistency of our testing procedure. Further, replacing the expectation by sample mean we obtain our test statistic

equation[equation omitted — 107 chars of source]

We reject the hypothesis $H_0$ if $n\,S_n$ becomes too large. When no additional smoothing is carried out, that is, $L$ is the identity operator, then $\tau_j=1$ for all $j\geq 1$. To achieve asymptotic normality we need to standardize our test statistic $S_n$ by appropriate mean and variance, which we introduce in the following definition.

definFor all $m\geq 1$ let $\varSigma_m$ be the covariance matrix of the random vector $Uf_\um^\tau(W)$ with entries ${\textsl s}_{jl}=\Ex\big[U^2f_j^\tau(W)f_{l}^\tau(W)\big]$, $1\leq j,l\leq m$. Then the trace and the Frobenius norm of $\varSigma_m$ are respectively denoted by \begin{equation*} \mu_m:=\sum_{j=1}^m {\textsl s}_{jj}\quadand\quad\varsigma_m:=\Big(\sum_{j,\,l=1}^m {\textsl s}_{jl}^2\Big)^{1/2}. \end{equation*}

Indeed the next result shows that $nS_n$ after standardization is asymptotically normally distributed if $\m$ increases appropriately as the sample size $n$ tends to infinity.

theoLet Assumptions (ref)--(ref) hold true. If $\m$ satisfies \begin{equation} \varsigma_\m^{-1}=o(1)\quad and\quad \Big(\sum_{j=1}^\m\tau_j\Big)^3=o(n) \end{equation} then under $H_0$ \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n-\mu_\m\big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}
remSince $\varsigma_\m^2\leq \eta_p\,\sigma^4\big(\sum_{j=1}^\m\tau_j\big)^2$ (cf. proof of Theorem (ref)) condition $\varsigma_\m^{-1}=o(1)$ implies that $\sum_{j=1}^\m\tau_j$ tends to infinity as $n$ increases. Moreover, from condition (ref) we see that by choosing a stronger decaying sequence $\tau$ the parameter $\m$ may be chosen larger. From the following theorem we see that if $\sum_{j=1}^\m\tau_j=O(1)$ only $m_n^{-1}=o(1)$ is required. $\square$

In the following result, we establish the asymptotic distribution of our test when the sequence of weights $\tau$ may have a stronger decay than in Theorem (ref), that is, we consider the case where $\tau$ satisfies $\sum_{j=1}^\m\tau_j=O(1)$. This holds, for instance, if the sequence $\tau$ satisfies $\tau_j\sim j^{-(1+\varepsilon)}$ for any $\varepsilon>0$. In this case, the asymptotic distribution changes and additional definitions have to be made. Let $\varSigma$ be the covariance matrix of the infinite dimensional centered vector $\big(Uf_j^\tau(W)\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma$ are denoted by $(\lambda_j)_{j\geq1}$. Below, we introduce a sequence $\{\chi_{1j}^2\}_{j\geq 1}$ of independent random variables that are distributed as chi-square with one degree of freedom.

theoLet Assumptions (ref)--(ref) hold true. If $\m$ satisfies \begin{equation} \sum_{j=1}^\m\tau_j=O(1)\quadand\quad m_n^{-1}=o(1) \end{equation} then under $H_0$ \begin{equation*} n\,S_n\stackrel{d}{\rightarrow} \sum_{j=1}^\infty \lambda_j\,\chi_{1j}^2. \end{equation*}
rem[Estimation of Critical Values] The asymptotic results of Theorem (ref) and (ref) depend on unknown population quantities. As we see in the following, the critical values can be easily estimated. Let $\bold W_m (\tau)$ denote a $n\times m$ matrix with entries $f_j^\tau(W_i)$ for $1\leq i\leq n$ and $1\leq j\leq m$. Moreover, $\bold U_n=(Y_1-\sol_0(Z_1),\dots,Y_n-\sol_0(Z_n))^t$. In the setting of Theorem (ref), we replace $\varSigma_m$ by \begin{equation*} \widehat\varSigma_m:=n^{-1}\bold W_m(\tau)^t\,diag(\bold U_n)^2\,\bold W_m(\tau). \end{equation*} Now the asymptotic result of Theorem (ref) continues to hold if we replace $\varsigma_\m$ by the Frobenius norm of $\widehat\varSigma_\m$ and $\mu_\m$ by the trace of $\widehat\varSigma_\m$. In the setting of Theorem (ref), the asymptotic distribution is not pivotal and has to approximated. First, the difference of critical values between $\sum_{j=1}^\infty \lambda_j\chi_{1j}^2$ and the truncated sum $\sum_{j=1}^{M_n}\lambda_j\,\chi_{1j}^2$ converges to zero if the integer $M_n>0$ tends to infinity (depending on $n$). Second, replace $(\lambda_j)_{1\leq j\leq M_n}$ by $(\widehat\lambda_j)_{1\leq j\leq M_n}$ which are the ordered eigenvalues of $\widehat\varSigma_{M_n}$. Observe $\max_{1\leq j\leq M_n}|\widehat\lambda_j-\lambda_j|=\|\widehat\varSigma_{M_n}-\varSigma_{M_n}\|=O(M_n n^{-1/2})$ almost surely. Hence, the critical values of $\sum_{j=1}^{M_n}\widehat\lambda_j\,\chi_{1j}^2$ converge in probability to the ones of the limiting distribution of $n\, S_n$ if $M_n=o(\sqrt{n})$. $\square$

Limiting behavior under local alternatives.

Let us study the power of the test statistic $S_n$, that is, the probability to reject a false hypothesis, against a sequence of linear local alternatives that tends to zero as $n\to\infty$. It is shown that the power of our tests essentially relies on the choice of the weighting sequence $\tau$.

Let us start with the case $\varsigma_\m^{-1}=o(1)$. We consider the following sequence of linear local alternatives

equation[equation omitted — 86 chars of source]

for some function $\delta\in \cL_Z^4:=\{\phi:\Ex|\phi(Z)|^4<\infty\}$. The next result establishes asymptotic normality for the standardized test statistic $S_n$. Let us denote $\delta_j:=\sqrt{\tau_j}\,\Ex[\delta(Z)f_j(W)]$.

propGiven the conditions of Theorem (ref) it holds under (ref) \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n-\mu_\m\big)\stackrel{d}{\rightarrow}\cN\Big(2^{-1/2}\sum_{j= 1}^\infty\delta_j^2,1\Big). \end{equation*}

As we see below the test statistic $S_n$ has power advantages if $\sum_{j=1}^\m\tau_j=O(1)$. Let us consider the sequence of linear local alternatives

equation[equation omitted — 69 chars of source]

for some function $\delta\in \cL_Z^4$. For the next result, the sequence $\{\chi_{1j}^2(\delta_j/\lambda_j)\}_{j\geq 1}$ denotes independent random variables that are distributed as non-central chi-square with one degree of freedom and non-centrality parameters $\delta_j/\lambda_j$.

propGiven the conditions of Theorem (ref) it holds under (ref) \begin{equation*} n\, S_n\stackrel{d}{\rightarrow}\sum_{j=1}^\infty\lambda_j\,\chi_{1j}^2(\delta_j/\lambda_j). \end{equation*}
remWe see from Proposition (ref) that our test can detect linear alternatives at a rate $\varsigma_\m^{1/2} n^{-1/2}$. On the other hand, if $\sum_{j=1}^\m\tau_j=O(1)$ then $S_n$ can detect local linear alternatives at the faster rate $n^{-1/2}$. But still our test with $L=\text{Id}$ can have better power against certain smooth classes of alternatives as illustrated by Hong95 and Horowitz2001. Indeed, the next subsection shows that additional smoothing changes the class of alternatives over which uniform consistency can be obtained. $\square$

Consistency

In this subsection, we establish consistency against a fixed alternative and uniform consistency of our test over appropriate function classes. Let us first consider the case of a fixed alternative. We assume that $H_0$ does not hold, that is, $\PP(\sol=\sol_0)<1$. The following proposition shows that our test has the ability to reject a false null hypothesis with probability $1$ as the sample size grows to infinity.

The consistency properties require the following additional assumption.

assA(i) The function $p_W/\nu$ is uniformly bounded away from zero. (ii) There exists a constant $\sigma_o>0$ such that $\Ex[U^2|W]\geq \sigma_o^2$.

Assumption (ref) $(i)$ implies that $\|LT(\sol-\sol_0)\|_W>0$ for any structural function $\sol$ in the alternative. Further, Assumption (ref) implies that $\sum_{j=1}^\m\tau_j^2=O(\varsigma_\m^2)$.

propAssume that $H_0$ does not hold. Let $\Ex|Y-\sol_0(Z)|^4<\infty$ and let Assumption (ref) (i) hold true. Consider the sequence $(\alpha_n)_{n\geq 1}$ satisfying $\alpha_n=o(n\varsigma_\m^{-1})$. Under the conditions of Theorem (ref) we have \begin{gather*} \PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big( nS_n-\mu_\m\big)>\alpha_n\Big)=1+o(1). \end{gather*} Under the conditions of Theorem (ref) we have $\alpha_n=o(n)$ and \begin{gather*} \PP\big(nS_n>\alpha_n\big)=1+o(1). \end{gather*}

In the following, we specify a class of functions over which our test $S_n$ is uniformly consistent. This essentially implies that there are no alternative functions in this class over which our test has low power. We show that our test is consistent uniformly over the class

equation*[equation* omitted — 175 chars of source]

where $C>0$ is a finite constant. Clearly, if $H_0$ is false then $\|LT(\sol-\sol_0)\|_W^2\geq \sr\,\varsigma_\m n^{-1}$ for all sufficiently large $n$ and some $\rho>0$. By Assumption (ref) the sequence $\tau$ is nonincreasing sequence with $\tau_1=1$ and hence, $\|LT(\sol-\sol_0)\|_W^2\leq \|T(\sol-\sol_0)\|_W^2\leq \|\sol-\sol_0\|_Z^2$ by Jensen's inequality. We conclude that $\cG_n^\sr$ contains all alternative functions whose $\cL_Z^2$-distance to the structural function $\sol_0$ is at least $n^{-1}\varsigma_\m$ within a constant. If the coefficients $\Ex[(\sol-\sol_0)(Z)f_j(W)]$ fluctuate for large $j$ then $\sol$ does not belong to $\cG_n^\sr$ if the decay of $\tau$ is too strong. On the other hand, if $\Ex[(\sol-\sol_0)(Z)f_j(W)]$ is sufficiently small for $j$ up to a finite constant then $\sol$ does not necessarily belong to $\cG_{n}^\sr$ with $\tau$ having a slow decay. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j\,\chi_{1j}^2$, respectively.

propLet Assumption (ref) be satisfied. For any $\varepsilon>0$, any $0<\alpha<1$, and any sufficiently large constant $\rho>0$ we have under the conditions of Theorem (ref) that \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cG_n^\sr}\PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big( nS_n-\mu_\m\big)>q_{1\alpha}\Big)\geq 1-\varepsilon, \end{equation*} while under the conditions of Theorem (ref) \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cG_n^\sr}\PP\Big(n\,S_n>q_{2\alpha}\Big)\geq 1-\varepsilon. \end{equation*}

A parametric specification test

In this section, we present a test whether the structural function $\sol$ is known up to a finite dimensional parameter. Let $\varTheta$ be a compact subspace of $\mathbb R^k$ then we consider the null hypothesis $H_{\text{p}}:$ there exists some $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$ for a known function $\phi$. The alternative hypothesis is that there exists no $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$ holds true.

The test statistic and its asymptotic distribution

Under Assumptions (ref) and (ref), the null hypothesis $H_{\text{p}}$ is equivalent to $L(g-T\phi(\cdot,\vartheta))=0$ for some $\vartheta\in\varTheta$. Thereby, to verify $H_{\text{p}}$ we make use of the test statistic $S_n$ given in (ref) where $\sol_0$ is replaced by $\phi(\cdot,\hthet)$ with $\hthet$ being an estimator of $\vartheta$. Hence, our test statistic for a parametric specification is given by

equation*[equation* omitted — 123 chars of source]

If the test statistic $S_n^{\text{p}}$ becomes too large then $H_{\text{p}}$ has to be rejected. To obtain asymptotic results for the statistic $S_n^{\text{p}}$ we require smoothness conditions of the function $\phi$ with respect to its second argument. Below we denote the vector of partial derivatives of $\phi$ with respect to $\vartheta=(\vartheta_1,\dots,\vartheta_k)^t$ by $\phi_\vartheta=(\phi_{\vartheta_l})_{1\leq l\leq k}$ and the matrix of second-order partial derivatives by $\phi_{\vartheta\vartheta}=(\phi_{\vartheta_j\vartheta_l})_{1\leq j,l\leq k}$.

assA(i) Let $\hthet$ be an estimator satisfying $\|\hthet-\thet\|=O_p(n^{-1/2})$ for some $\thet\in{\text{int}}(\varTheta)$ with $\sol(\cdot)=\phi(\cdot,\vartheta_0)$ if $H_{\text{p}}$ holds true. (ii) The function $\phi$ is twice partial differentiable with respect to its second argument. There exists some constant $\eta_\phi\geqslant 1$ such that \begin{equation*} \sup_{1\leq l\leq k}\Ex |\phi_{\vartheta_l}(Z,\thet)|^4\leqslant \eta_\phi\quadand \quad \sup_{1\leq j,l\leq k}\Ex \sup_{\theta\in\Theta}|\phi_{\vartheta_j\vartheta_l}(Z,\theta)|^4\leqslant \eta_\phi. \end{equation*}

The following proposition establishes asymptotic normality of $S_n^{\text{p}}$ after standardization.

theoLet Assumptions (ref)--(ref) and (ref) hold true. If $\m$ satisfies (ref), then under $H_{\textsl{p}}$ \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\, S_n^{\textsl{p}}-\mu_\m\big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}

In the following theorem, we state the asymptotic distribution of $nS_n^{\textsl{p}}$ when $\sum_{j=1}^\m\tau_j=O(1)$. In this case, we assume that $\hthet$ satisfies under $H_{\textsl{p}}$

equation[equation omitted — 100 chars of source]

where $V_{i}:=(Y_i,Z_i,W_i,\thet)$ and $h_\uk(V_{i})=(h_1(V_i),\dots, h_k(V_i))^t$ where $h_j$, $1\leq j\leq k$, are real valued functions. It is well known that this representation holds if $\hthet$ is the generalized method of moments estimator. Let $\varSigma^\textsl{p}$ be the covariance matrix of the infinite dimensional centered vector $\big(Uf_j^\tau(W)-\Ex[f_j^\tau(W)\phi_\vartheta(Z,\thet)^t]h_\uk(V)\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{p}$ are denoted by $(\lambda_j^\textsl{p})_{j\geq1}$.

theoLet Assumptions(ref)--(ref) and (ref) hold true. Assume that $H_{\textsl{p}}$ holds true and $\hthet$ satisfies condition (ref) with $\Ex h_{j}(V)=0$ and $\Ex |h_{j}(V)|^4<\infty$, $1\leq j\leq k$. If $\m$ satisfies (ref), then \begin{equation*} n\,S_n^{\textsl{p}}\,\stackrel{d}{\rightarrow} \sum_{j=1}^\infty\lambda_j^\textsl{p}\,\chi_{1j}^2. \end{equation*}
rem[Estimation of Critical Values] For the estimation of critical values of Theorem (ref) and (ref), let us define $\bold U_n^{\textsl{p}}=\big(Y_1-\phi(Z_1,\hthet),\dots,Y_n-\phi(Z_n,\hthet)\big)^t$. We estimate the covariance matrix $\varSigma_m$ by \begin{equation*} \widehat\varSigma_m:=n^{-1}\,\bold W_m(\tau)^t\,diag(\bold U_n^{\textsl{p}})^2\,\bold W_m(\tau). \end{equation*} Now the asymptotic result of Theorem (ref) continues to hold if we replace $\varsigma_\m$ by the Frobenius norm of $\widehat\varSigma_\m$ and $\mu_\m$ by the trace of $\widehat\varSigma_\m$. In the setting of Theorem (ref), we replace $\varSigma^\textsl{p}$ by a finite dimensional matrix. Let $\bold A_k$ be a $n\times k$ matrix with entries $\phi_{\vartheta_l}(Z_i,\hthet)$ for $1\leq i\leq n$, $1\leq l\leq k$ and $\bold h_k(V)=\big(h_\uk(V_{1}),\dots,h_\uk(V_{n})\big)^t$. Then define $\bold V_k:=n^{-1}\bold h_k(V)\bold A_k^t $. Given a sufficiently large integer $M>0$ we estimate $\varSigma^\textsl{p}$ by \begin{equation*} \widehat\varSigma_M^{\textsl{p}}:=n^{-1}\bold W_M(\tau)^t \Big(diag(\bold U_n^{\textsl{p}})-\bold V_k\Big)^t \Big(diag(\bold U_n^{\textsl{p}})-\bold V_k\Big)\bold W_M(\tau). \end{equation*} Hence, we approximate $\sum_{j=1}^\infty \lambda_j\chi_{1j}^2$ by the finite sum $\sum_{j=1}^{M_n}\widehat\lambda_j^{\textsl{p}}\,\chi_{1j}^2$ where $(\widehat\lambda_j^{\textsl{p}})_{1\leq j\leq M_n}$ are the ordered eigenvalues of $\widehat\varSigma_{M_n}^{\textsl{p}}$. We have $\max_{1\leq j\leq M_n}|\widehat\lambda_j^{\textsl{p}}-\lambda_j^{\textsl{p}}|=o_p(1)$ if $M_n=o(\sqrt n)$. $\square$

Limiting behavior under local alternatives and consistency.

In the following, we study the power and consistency properties of the test statistic $S_n^{\text{p}}$. In the following, we consider a sequence of linear local alternatives (ref) or (ref) with $\sol_0=\phi(\thet,\cdot)$. Further, let $\delta_\perp$ denote the projection of $\delta$ onto the orthogonal complement of $\phi(\cdot,\thet)$; that is, $\Ex[\phi_\vartheta(Z,\thet)\delta_\perp(Z)]=0$. Let us denote $\delta_{j\perp}:=\sqrt{\tau_j}\,\Ex[\delta_\perp(Z)f_j(W)]$.

propLet the conditions of Theorem (ref) be satisfied. Then under (ref) with $\sol_0=\phi(\cdot,\thet)$ it holds \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{p}}-\mu_\m\big)\stackrel{d}{\rightarrow}\cN\Big(2^{-1/2}\sum_{j= 1}^\infty\delta_{j\perp}^2,1\Big). \end{equation*} Let the conditions of Theorem (ref) be satisfied. Then under (ref) with $\sol_0=\phi(\cdot,\thet)$ it holds \begin{equation*} n\,S_n^{\textsl{p}}\stackrel{d}{\rightarrow}\sum_{j=1}^\infty\lambda_j^{\textsl{p}}\,\chi_{1j}^2(\delta_{j\perp}/\lambda_j^{p}). \end{equation*}
remUnder homoscedasticity, that is, $\Ex[U^2|W]=\sigma_o^2$, $W\sim\cU[0,1]$, and $L=\text{Id}$ we see from Proposition (ref) that our test has the same power properties as the test of Hong95. On the other hand, if $\sum_{j=1}^\m\tau_j=O(1)$ then our test can detect local linear alternatives at a rate $n^{-1/2}$ as in Horowitz2006, which decreases more quickly than the rate obtained by Tripathi2003. $\square$

The next proposition establishes consistency of our test against a fixed alternative model. It is assumed that $H_\text p$ is false, that is, there exists no $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$. In this situation, $\thet$ denotes the probability limit of the estimator $\hthet$.

propAssume that $H_{\textsl{p}}$ does not hold. Let $\Ex|Y-\phi(Z,\thet)|^4<\infty$ and Assumption (ref) (i) hold true. Let $(\alpha_n)_{n\geq 1}$ as in Proposition (ref). Under the conditions of Theorem (ref) we have \begin{equation*} \PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big( nS_n^{\textsl{p}}-\mu_\m\big)>\alpha_n\Big)=1+o(1). \end{equation*} Given the conditions of Theorem (ref) it holds \begin{equation*} \PP\big(n\,S_n^{\textsl{p}}>\alpha_n\big)=1+o(1). \end{equation*}

In the following, we show that $S_n^{\textsl{p}}$ is consistent uniformly over the function class

multline*[multline* omitted — 202 chars of source]

for some constant $C>0$ and $\thet$ denotes the probability limit of $\hthet$. Similarly as in the previous section, it can be seen that $\cH_{n}^\sr$ only contains functions whose $\cL_Z^2$ distance to $\phi(\cdot,\vartheta_0)$ is at least $n^{-1}\varsigma_\m$ within a constant. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{p}}\,\chi_{1j}^2$, respectively.

propLet Assumption (ref) be satisfied. For any $\varepsilon>0$, any $0<\alpha<1$, and any sufficiently large constant $\rho>0$ we have under the conditions of Theorem (ref) that \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cH_n^\sr}\PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{p}}-\mu_\m\big)>q_{1\alpha}\Big)\geq 1-\varepsilon, \end{equation*} whereas under the conditions of Theorem (ref) it holds \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cH_{n}^\sr}\PP\big(n\,S_n^{\textsl{p}}>q_{2\alpha}\big)\geq 1-\varepsilon. \end{equation*}

A nonparametric test of exogeneity

Endogeneity of regressors is a common problem in econometric applications. Falsely assuming exogeneity of the regressors leads to inconsistent estimators. On the other hand, treating exogenous regressors as if they were endogenous can lower the accuracy of estimation dramatically. In this section, we propose a test whether the vector of regressors $Z$ is exogenous, that is, $\Ex[U|Z]=0$ or equivalently $\sol(Z)=\Ex[Y|Z]$. In this section, let $\sol_0(Z)=\Ex[Y|Z]$ then the hypothesis under consideration is given by $H_{\textsl{e}}:\, \sol=\sol_0$. The alternative hypothesis is that $\sol\neq \sol_0$.

The test statistic and its asymptotic distribution

To establish a test of exogeneity, let us first introduce an estimator of the conditional mean of $Y$ given $Z$. This estimator is based on a sequence of approximating functions $\{e_j\}_{j\geq 1}$ belonging to $\cL_Z^2$. Further, let $\bold Z_k$ denote a $n\times k$ matrix with entries $e_j(Z_i)$ for $1\leq i\leq n$ and $1\leq j\leq k$. Moreover, let $\bold Y_n=(Y_1,\dots,Y_n)^t$. Then we define the estimator

equation[equation omitted — 166 chars of source]

In contrast to the parametric case we need to allow for $k$ tending to infinity as $n\to\infty$ in order to ensure consistency of the estimator $\osol_k$. Under conditions given below $\bold Z_\k^t\bold Z_\k$ will be nonsingular with probability approaching one and hence its generalized inverse will be the standard inverse. Note that the asymptotic behavior of the estimator $\osol_k$ was studied, for example, by Newey1997.

Under Assumptions (ref) and (ref), the null hypothesis $H_{\textsl{e}}$ is equivalent to $L(g-T\sol_0)=0$. Consequently, our test of exogeneity of $Z$ is based on the goodness-of-fit statistic $S_n$ introduced in (ref) but where $\sol_0$ is replaced by the series estimator $\osol_\k$. The proposed test statistic for $H_{\textsl{e}}$ is now given by

equation*[equation* omitted — 119 chars of source]

where $\k$ and $\m$ tend to infinity as $n\to\infty$. The hypothesis of exogeneity of $Z$ has to be rejected if $S_n^{\textsl{e}}$ becomes too large.

For controlling the bias of the estimator $\osol_\k$ we specify in the following a rate of approximation (cf. Newey1997). Let $\sw=(\sw_j)_{j\geq 1}$ be a nondecreasing sequence with $\sw_1=1$. We assume that $\sol_0$ belongs to

equation*[equation* omitted — 162 chars of source]

Here, the sequence of weights $\sw$ measures the approximation error of $\sol_0$ with respect to the functions $\{e_j\}_{j\geq1}$.

assA(i) Let $\sol_0\in\cF_\sw$ with nondecreasing sequence $\sw$ satisfying $j^2=o(\sw_j)$. (ii) There exists some constant $\eta_e\geqslant 1$ such that $\sup_{z\in\cZ}\|e_\ukn(z)\|^2\leqslant \eta_e\k$. (iii) The smallest eigenvalue of $\Ex[e_\uk(Z)e_\uk(Z)^t]$ is bounded away from zero uniformly in $k$. (iv) $\Ex[U^2|Z]$ is bounded.

Assumption (ref) $(i)$ determines the required asymptotic behavior of the rate $\sw$. For splines and power series this assumption is satisfied if the number of continuous derivatives of $\sol_0$ divided by the dimension of $Z$ equals two. Assumption (ref) $(ii)$ and $(iii)$ restrict the magnitude of the approximating functions $\{e_j\}_{j\geq1}$ and impose nonsingularity of their second moment matrix.

We are now in the position to proof the following asymptotic result for the standardized test statistic $S_n^\text e$. Here, a key requirement is that $\k=o(\varsigma_\m)$ implying that $\k=o(\sum_{j=1}^\m\tau_j)$ and, in particular, $\k=o(\m)$ if the smoothing operator $L$ is the identity.

theoLet Assumptions (ref)--(ref) and (ref) be satisfied. If \begin{equation} n=o(\sw_\k\varsigma_\m), \,\k=o(\varsigma_\m),\quad and\quad \Big(\sum_{j=1}^\m\tau_j\Big)^3=o(n) \end{equation}then under $H_{\textsl{e}}$ it holds \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{e}}-\mu_\m\big)\stackrel{d}{\rightarrow} \cN(0,1). \end{equation*}
exampleLet $Z$ be continuously distributed with $\dim(Z)=r$ and set $L=\text{Id}$. Consider the polynomial case where $\sw_j\sim j^{2p/r}$ with $p>1$ and let $\m\sim n^\nu$ with $0<\nu<1/3$. Let Assumption (ref) hold true then $\sqrt \m=O(\varsigma_\m)$. Hence, condition (ref) is satisfied if $\k\sim n^\kappa$ with \begin{equation} r(1-\nu/2)/(2p)<\kappa<\nu/2. \end{equation} This ensures that the bias of this estimator in the statistic $S_n^{\textsl{e}}$ is asymptotically negligible. Note that condition (ref) requires $2p>r\,(2/\nu-1)$. Hence, with a larger dimension $r$ also the smoothness of $\sol_0$ has to increase, reflecting the curse of dimensionality. $\hfill\square$

The next result states an asymptotic distribution result for the statistic $S_n^{\textsl{e}}$ if $\sum_{j=1}^\m\tau_j=O(1)$. Let $\varSigma^\textsl{e}$ be the covariance matrix of the infinite dimensional centered vector $\big(U(f_j^\tau(W)-\sum_{l\geq 1}\Ex[f_j^\tau(W)e_l(Z)]e_l(Z))\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{e}$ are denoted by $(\lambda_j^\textsl{e})_{j\geq1}$.

theoLet Assumptions (ref)--(ref) and (ref) be satisfied. If \begin{equation} \sum_{j=1}^\m\tau_j=O(1),\quad n=O(\sw_\k),\quad k_n^3=o(n), and\quad m_n^{-1}=o(1) \end{equation} then under $H_{\textsl{e}}$ it holds \begin{equation*} n\,S_n^{\textsl{e}}\stackrel{d}{\rightarrow} \sum_{j=1}^\infty\lambda_j^\textsl{e}\,\chi_{1j}^2. \end{equation*}
exampleConsider the setting of Example (ref) but where the eigenvalues of $L$ satisfy $\tau_j\sim j^{-2}$. Condition (ref) is satisfied if $\m\sim n^\nu$ for some $\nu>0$ and $\k\sim n^\kappa$ with $r/(2p)<\kappa<1/3$. Here, the required smoothness of $\sol_0$ is $p>3r/2$. In contrast to the setting of Theorem (ref), the estimator of $\sol_0$ needs to be undersmoothed. This ensures that the bias of this estimator in the statistic $S_n^{\textsl{e}}$ is asymptotically negligible. $\hfill\square$
remIn contrast to Blundell2007 no smoothness assumptions on the joint distribution of $(Z,W)$ is required here. In addition, we do not need any assumption that links the smoothness of the regression function $\sol_0$ to the smoothness of the joint density of $(Z,W)$. $\square$
rem[Estimation of Critical Values] For the estimation of critical values of Theorem (ref) and (ref), let us define $\bold U_n^{\textsl{e}}=\big(Y_1-\osol_\k(Z_1),\dots,Y_n-\osol_\k(Z_n)\big)^t$. For any $m\geq 1$ we estimate the covariance matrix $\varSigma_m$ by \begin{equation*} \widehat\varSigma_m:=n^{-1}\,\bold W_m(\tau)^t\,diag(\bold U_n^{\textsl{e}})^2\,\bold W_m(\tau). \end{equation*} Now the asymptotic result of Theorem (ref) continues to hold if we replace $\varsigma_\m$ by the Frobenius norm of $\widehat\varSigma_\m$ and $\mu_\m$ by the trace of $\widehat\varSigma_\m$. This consistency is shown in Lemma (ref). In the setting of Theorem (ref), we replace $\varSigma^\textsl{e}$ by a finite dimensional matrix \begin{equation*} \widehat\varSigma_M^{\textsl{e}}:=n^{-1} \bold W_M(\tau)^t\Big(I_n-n^{-1}\bold Z_\k\bold Z_\k^t\Big) \,diag(\bold U_n^{\textsl{e}})^2\,\Big(I_n-n^{-1}\bold Z_\k\bold Z_\k^t\Big)\bold W_M(\tau) \end{equation*} where $M>0$ is a sufficiently large integer. Let $(\widehat\lambda_j^{\textsl{e}})_{1\leq j\leq M_n}$ denote the ordered eigenvalues of $\widehat\varSigma_{M_n}^{\textsl{e}}$. Hence, we approximate $\sum_{j=1}^\infty \lambda_j^\textsl{e}\chi_{1j}^2$ by the finite sum $\sum_{j=1}^{M_n}\widehat\lambda_j^{\textsl{e}}\,\chi_{1j}^2$ where $\max_{1\leq j\leq M_n}|\widehat\lambda_j^{\textsl{e}}-\lambda_j^{\textsl{e}}|=o_p(1)$ if $M_n=o(\sqrt n)$. $\square$
lemConsider $\widehat\varSigma_\m$ as defined in Remark (ref). Under conditions of Theorem (ref) or Theorem (ref) the difference of its Frobenius norm to $\varsigma_\m$ and its trace to $\mu_\m$ converge in probability to zero.

Limiting behavior under local alternatives and consistency.

Similar to the previous sections we study the power and consistency properties of our test. Let us study the power of our test of exogeneity under linear local alternatives (ref) or (ref). In these cases, it holds $\Ex[U|W]=0$ but $\Ex[U|Z]= -\varsigma_\m^{1/2}n^{-1/2} \delta(Z)$ under (ref) or $\Ex[U|Z]=-n^{-1/2}\delta(Z)$ under (ref).

propGiven the conditions of Theorem (ref) and Assumption (ref) (ii) it holds under (ref) \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{e}}-\mu_\m\big)\stackrel{d}{\rightarrow}\cN\Big(2^{-1/2}\sum_{j= 1}^\infty\delta_j^2,1\Big). \end{equation*} Given the conditions of Theorem (ref) it holds under (ref) \begin{equation*} n\, S_n^{\textsl{e}}\stackrel{d}{\rightarrow}\sum_{j=1}^\infty\lambda_j^\textsl{e}\,\chi_{1j}^2(\delta_j/\lambda_j^\textsl{e}). \end{equation*}

Let us now establish consistency of our tests when $H_\text e$ does not hold, that is, $\PP\big(\sol=\sol_0\big)<1$.

propAssume that $H_{\textsl{e}}$ does not hold. Let $\Ex|Y-\sol_0(Z)|^4<\infty$ and Assumption (ref) (i) hold true. Let $(\alpha_n)_{n\geq 1}$ as in Proposition (ref). Under the conditions of Theorem (ref) we have \begin{equation*} \PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big( nS_n^{\textsl{e}}-\mu_\m\big)>\alpha_n\Big)=1+o(1), \end{equation*} whereas in the setting of Theorem (ref) \begin{equation*} \PP\big(n\,S_n^{\textsl{e}}>\alpha_n\big)=1+o(1). \end{equation*}

In the following we show that our tests are consistent uniformly over the function class

equation*[equation* omitted — 162 chars of source]

form some constant $C>0$. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{e}}\,\chi_{1j}^2$, respectively.

propLet Assumption (ref) be satisfied. Under the conditions of Theorem (ref) we have for any $\varepsilon>0$, any $0<\alpha<1$, and any sufficiently large constant $\rho>0$ that \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cI_n^\sr}\PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{e}}-\mu_\m\big)>q_{1\alpha}\Big)\geq 1-\varepsilon,\\ \end{equation*} whereas under the conditions of Theorem (ref) it holds \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cI_{n}^\sr}\PP\big(n\,S_n^{\textsl{e}}>q_{2\alpha}\big)\geq 1-\varepsilon. \end{equation*}

A nonparametric specification test

A solution to the linear operator equation (ref) only exists if $g$ belongs to the range of $\op$. This might be violated if, for instance, the instrument is not valid, that is, $\Ex[U|W]\neq 0$. In many economic applications a priori smoothness restriction on the unknown function can be justified which we capture by a set of functions $\cF$. We consider the hypothesis $H_{\textsl{np}}$: there exists a solution $\sol_0\in\cF$ to (ref). The alternative hypothesis is that there exists a solution (ref) which does not belong to $\cF$. Under the alternative only unsmooth functions solve the conditional moment restriction which can be interpreted as a failure of validity of the instrument $W$. We see in this section that our results allow also for a test of dimension reduction of the vector of regressors $Z$, that is, whether some regressors can be omitted from the structural function $\sol_0$.

Nonparametric estimation method

\paragraph{The nonparametric estimator.} In the following, we derive an estimator of $\sol_0$ under the null hypothesis $H_{\textsl{np}}$. For simplicity, assume that $\cZ=\cW$ and consider a sequence $\{e_j\}_{j\geq 1}$ of approximating functions which are orthonormal on $\cZ$ with respect to the Lebesque measure $\nu$. Under conditions given below, $\sol_0$ has the expansion $\sol_0(\cdot)=\sum_{l=1}^\infty \int\sol_0(z)e_l(z)\nu(z)dz\,e_l(\cdot)$. Thereby, the conditional moment restriction under $H_{\textsl{np}}$ leads to the following unconditional moment restrictions

equation[equation omitted — 93 chars of source]

for $j \geq 1$. This motivates the following orthogonal series type estimator. Let $\bold Z_k$ and $\bold Y_n$ be as in the previous section and let $\bold X_k$ denote a $n\times k$ matrix with entries $e_j(W_i)$ for $1\leq i\leq n$ and $1\leq j\leq k$. Then for each $k\geq 1$ we consider the estimator

equation[equation omitted — 174 chars of source]

Under conditions given below $\bold X_\k^t\bold Z_\k$ will be nonsingular with probability approaching one and hence its generalized inverse will be the standard inverse. The nonparametric estimator $\hsol_k$ given in (ref) was studied by Johannes2009, Horowitz2011, and Horowitz2009.

\paragraph{Additional assumptions.} In the following, we specify a priori smoothness assumptions captured by the set $\cF$. As noted by Horowitz2009, uniformly consistent testing of $H_{\textsl{np}}$ is only possible if the null is restricted that any solution to (ref) is smooth. Here, we assume that under the null hypothesis $\sol_0$ belongs to the ellipsoid $\cF:=\cF_{\sw}^\sr := \big\{\phi\in \cL_Z^2: \sum_{j= 1}^\infty\sw_j\Ex[\phi(Z)e_j(Z)]^2\leq \sr\big\}$. As in the previous section, $\sw=(\sw_j)_{j\geq 1}$ measures the approximation error of $\sol_0$ with respect to the basis $\{e_j\}_{j\geq 1}$.

Further, as usual in the context of nonparametric instrumental regression, we specify some mapping properties of the conditional expectation operator $T$. Denote by $\cT$ the set of all nonsingular operators mapping $\cL_Z^2$ to $\cL^2_W$. Given a sequence of weights $\Opw:=(\Opw_j)_{j\geqslant 1}$ and $\Opd\geqslant 1$ we define the subset $\Opwd$ of $\cT$ by

equation*[equation* omitted — 201 chars of source]

If $p_Z/\nu$ is bounded from above and $p_W/\nu$ is uniformly bounded away from zero then the conditional expectation operator $T$ belongs to $\Opwd$ with $\opw_j=1$, $j\geq 1$, due to Jensen's inequality. Notice that for all $T\in\Opwd$ it follows that $\normV{T\basZ_j}_W^2\leq d\,\eta_p \Opw_j$ and thereby, the condition $T\in\Opwd$ links the operator $T$ to the basis $\{e_j\}_{j\geq 1}$. In the following, we denote $[\op]_\uk=\Ex[e_\uk(W)e_\uk(Z)^t]$ which is assumed to be a nonsingular matrix. In what follows, we introduce a stronger condition on the basis $\{e_l\}_{l\geq1}$. We denote by $\cTdDw$ for some $\tD\geq \td$ the subset of $\cTdw $ given by

equation*[equation* omitted — 184 chars of source]

The class $\cTdDw$ only contains operators $T$ whose off-diagonal elements of $[T]_\uk^{-1}$ are sufficiently small for all $k\geq 1$. A similar diagonality restriction has been used by HallHorowitz2005 or BJ2011. Besides the mapping properties for the operator $T$ we need a stronger assumption for the basis under consideration. The following condition gathers conditions on the sequences $\sw$ and $\opw$.

assA(i) Under $H_{\textsl{np}}$, let $\sol_0\in\cF_\sw^\sr$ with nondecreasing sequence $\sw$ satisfying $j^3=o(\sw_j)$. (ii) The sequence $\{e_j\}_{j\geq 1}$ is an orthogonal basis on $\cZ=\cW$ with respect to $\nu$. (iii) There exists some constant $\eta_e\geqslant 1$ such that $\sup_{j\geq 1}\sup_{z\in\cZ}|e_j(z)|\leqslant \eta_e$. (iv) Let $T\in\cTdDw$ with $\Opw$ being a strictly positive sequences such that $\opw$ and $(\opw_{j}/\tau_{j})_{j\geq 1}$ are nonincreasing. (v) $p_Z/\nu$ is bounded from above and $p_W/\nu$ is uniformly bounded away from zero.

Due to Assumption (ref) $(iv)$ the degree of additional smoothing for our testing procedure must not be stronger than the degree of ill-posedness implied by the conditional expectation operator $T$. Under similar assumptions as above, Johannes2009 show that mean integrated squared error loss of $\hsol_{\k}$ attains the optimal rate of convergence $\cR_n:=\max\big(\sw_{\k}^{-1}, \sum_{j=1}^{\k}(n\opw_j)^{-1}\big)$. Due to Assumption (ref) $(v)$ we do not require orthonormal bases with respect to the unknown distribution $(Z,W)$ (cf. Remark 3.2 of BJ2011).

The test statistic and its asymptotic distribution

As in the previous sections, our test is based on the observation that the null hypothesis $H_{\textsl{np}}$ is equivalent to $L(g-T\sol_0)=0$. Our goodness-of-fit statistic for testing nonparametric specifications is given by $S_n$ where $\sol_0$ is replaced by the nonparametric estimator $\hsol_{\k}$ given in (ref), that is,

equation*[equation* omitted — 124 chars of source]

If $S_n^{\textsl{np}}$ becomes too large then there exists no function in $\cF_\sw^\sr$ solving (ref). The next result establishes asymptotic normality of $S_n^{\textsl{np}}$ after standardization. Again, a key requirement to obtain this asymptotic distribution is that $\k=o(\varsigma_\m)$ implying that $\k=o(\m)$ if the smoothing operator $L$ is the identity. This corresponds to the test of overidentification in the parametric framework where more orthogonality restrictions than parameters are required.

theoLet Assumptions (ref)--(ref) and (ref) be satisfied. If \begin{equation} n\opw_\k=o(\sw_\k\varsigma_{\m}), \,\k=o(\varsigma_{\m}),\,\k\Big(\sum_{j=1}^\m\tau_j\Big)^2=O(n\opw_\k),\,and \Big(\sum_{j=1}^\m\tau_j\Big)^3=o(n) \end{equation} then it holds under $H_{\textsl{np}}$ \begin{equation*} (\sqrt 2\varsigma_{\m})^{-1} \Big(nS_n^{\textsl{np}}-\mu_{\m}\Big)\stackrel{d}{\rightarrow}\mathcal N(0,1). \end{equation*}
exampleConsider the setting of Example (ref). In the mildly ill posed case where $\opw_j\sim j^{-2a/r}$ for some $a\geq0$ condition (ref) holds true if $\k\sim n^\kappa$ with $\kappa<\nu/2$ and \begin{equation*} {r(1-\nu/2)/(2a+2p)}<\kappa<r(1-2\nu)/(2a+r). \end{equation*} In the severely ill posed case, that is, $\opw_j\sim \exp(-j^{2a/r})$ for some $a>0$, condition (ref) is satisfied if, for example, $\m$ satisfies $\m=o(k_n^p)$ and $k_n=o(\sqrt\m)$ where $\k\sim\big(\log n-\log (m_n^{3/2})\big)^{r/(2a)}$. $\hfill\square$

The next result states an asymptotic distribution of our test if $\sum_{j=1}^\m\tau_j=O(1)$. Let $\varSigma^\textsl{np}$ be the covariance matrix of the infinite dimensional centered vector $\big(U(f_j^\tau(W)-e_j^\tau(W))\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{np}$ are denoted by $(\lambda_j^\textsl{np})_{j\geq1}$.

theoLet Assumptions (ref)--(ref) and (ref) be satisfied. If \begin{equation} \sum_{j=1}^\m\tau_j=O(1), \quad n\opw_\k=o(\sw_\k),\quad k_n^3=o(n\opw_\k), and m_n^{-1}=o(1) \end{equation} then it holds under $H_{\textsl{np}}$ \begin{equation*} n\,S_n^{\textsl{np}}\stackrel{d}{\rightarrow} \sum_{j=1}^\infty\lambda_j^\textsl{np}\,\chi_{1j}^2. \end{equation*}
exampleConsider the setting of Example (ref). In the mildly ill posed case, that is, $\opw_j\sim j^{-2a/r}$ for some $a\geq0$, condition (ref) is satisfied if $\m\sim n^\nu$ for some $\nu>0$ and $\k\sim n^\kappa$ with \begin{equation*} r/(2a+2p)<\kappa<r/(2a+3r). \end{equation*} In the severely ill posed case, that is, $\opw_j\sim \exp(-j^{2a/r})$ for some $a>0$, condition (ref) is satisfied if $\k\sim\big(\log (n^{1+\varepsilon})\big)^{r/(2a)}$ for any $\varepsilon>0$. In contrast to Theorem (ref), we require undersmoothing of the estimator $\hsol_\k$. $\hfill\square$
remIf the basis $\{e_j\}_{j\geq 1}$ coincides with $\{f_j\}_{j\geq 1}$ then $n\,S_n^{\textsl{np}}$ is asymptotically degenerate. To avoid this degeneracy problem we choose different bases functions and hence, sample splitting as used by Horowitz2009 is not necessary here. $\hfill\square$
remLet $Z'$ be a vector containing only entries of $Z$ with $\dim(Z')<\dim(Z)$. It is easy to generalize our previous result for a test of $H_{\textsl{np}}'$: there exists a solution $\sol_0\in\cF_\sw^\sr$ to (ref) only depending on $Z'$. To be more precise consider the test statistic \begin{equation*} S_n^{'\textsl{np}}:=\big\|n^{-1}\sum_{i=1}^n\big(Y_i-\hsol_\k(Z_i')\big)f_\umn^\tau(W_i)\|^2 \end{equation*} where $\hsol_\k$ is the estimator (ref) based on an iid. sample $(Y_1,Z_1',W_1),\dotsc,(Y_n,Z_n',W_n)$ of $(Y,Z',W)$. Under $H_{\textsl{np}}'$ we consider the conditional expectation operator $T':\cL_{Z'}^2\to \cL_W^2$ with $T'\phi:=\Ex[\phi(Z')|W]$. It is interesting to note that if $T$ is nonsingular then also $T'$ is. Hence, for a test of $H_{\textsl{np}}'$ we may replace Assumption (ref) by the weaker condition that $T'$ is nonsingular. Moreover, under $H_{\textsl{np}}'$ the results of Theorem (ref) and (ref) still hold true if we replace $Z$ by $Z'$. $\hfill\square$

In the mildly ill-posed case, the estimation precision suffers from the curse of dimensionality. Hence, by the test of dimension reduction of $Z$ we can increase the accuracy of estimation of $\sol_0$. On the other hand, in the severely ill-posed case the rate of convergence is independent of the dimension of $Z$ (cf. ChenReiss2008). As the next example illustrates, a dimension reduction test can also weaken the required restrictions on the instrument to obtain identification of $\sol$ in the restricted model.

exampleLet $Z=(Z^{(1)},Z^{(2)})$ where both, $Z^{(1)}$ and $Z^{(2)}$ are endogenous vectors of regressors. But only $Z^{(1)}$ satisfies a sufficiently strong relationship with the instrument $W$ in the sense that for all $\phi\in \cL_{Z^{(1)}}^2$ condition $\Ex[\phi(Z^{(1)})|W]=0$ implies $\phi=0$. In this example, we do not assume that this completeness condition is fulfilled for the joint distribution of $(Z^{(2)},W)$. Thereby only the operator $T^{(1)}:\cL_{Z^{(1)}}^2\to \cL_W^2$ with $T^{(1)}\phi:=\Ex[\phi(Z^{(1)})|W]$ is nonsingular but $T$ is singular. If our dimension reduction test of $Z$ indicates that $Z^{(2)}$ can be omitted from the structural function $\sol_0$ then we obtain identification in the restricted model. $\hfill\square$
rem[Estimation of Critical Values] For the estimation of critical values of Theorem (ref) and (ref), let us define $\bold U_n^{\textsl{np}}=\big(Y_1-\hsol_\k(Z_1),\dots,Y_n-\hsol_\k(Z_n)\big)^t$. For all $m\geq 1$, we estimate the covariance matrix $\varSigma_m$ by \begin{equation*} \widehat\varSigma_m:=n^{-1}\,\bold W_m(\tau)^tdiag(\bold U_n^{\textsl{np}})^2\,\bold W_m(\tau). \end{equation*} Now the asymptotic result of Theorem (ref) continues to hold if we replace $\varsigma_\m$ by the Frobenius norm of $\widehat\varSigma_\m$ and $\mu_\m$ by the trace of $\widehat\varSigma_\m$ (this is easily seen from the proof of Lemma (ref) assuming that $\{f_j\}_{j\geq 1}$ is uniformly bounded). In the setting of Theorem (ref), we replace $\varSigma^\textsl{np}$ by a finite dimensional matrix. Let $\bold V_{k}:=\bold W_k \big(\bold Z_k^t\bold W_k)^{-1}\bold Z_k^t$ for $k\geq 1$. Then for a sufficiently large integer $M>0$ we estimate $\varSigma^\textsl{np}$ by \begin{equation*} \widehat\varSigma_M^{\textsl{np}}:=n^{-1} \bold W_M(\tau)\big(I_n-\bold V_{\k}\big)^t\, diag(\bold U_n^{\textsl{np}})^2\,\big(I_n-\bold V_{\k}\big)\bold W_M(\tau). \end{equation*} Hence, we approximate $\sum_{j=1}^\infty \lambda_j^{\textsl{np}}\chi_{1j}^2$ by the finite sum $\sum_{j=1}^{M_n}\widehat\lambda_j^{\textsl{np}}\,\chi_{1j}^2$ where $(\widehat\lambda_j^{\textsl{np}})_{1\leq j\leq M_n}$ are the ordered eigenvalues of $\widehat\varSigma_{M_n}^{\textsl{np}}$ where $\max_{1\leq j\leq M_n}|\widehat\lambda_j^{\textsl{np}}-\lambda_j^{\textsl{np}}|=o_p(1)$ if $M_n=o(\sqrt n)$. $\square$

Limiting behavior under local alternatives and consistency.

Similar to the previous sections we study the power and consistency properties of our test. To study the power against local alternatives of the statistic $S_n^{\textsl{np}}$ we consider alternative models with the function $\sol_\k(\cdot)=e_\k(\cdot)^t[T]_\k^{-1}\Ex[Yf_\ukn(W)]$. We consider alternative models

equation[equation omitted — 84 chars of source]

for some function $\delta\in L_Z^4$ and where $\Ex[U|W]=0$. Let $\sol$ be such that $\Ex[Y-\sol(Z)|W]=0$. Due to (ref) $\sol$ does not belong $\cF_\sw^\rho$ and hence $H_{\textsl{np}}$ fails. Indeed, if $\sol\in \cF_\sw^\rho$ then we show in the appendix that $\|T(\sol-\sol_\k)\|_W^2=O(\opw_\k\sw_\k^{-1})=o(\varsigma_\m n^{-1})$ due to condition (ref) (or (ref)), which is in contrast to (ref).

propLet Assumption (ref) (ii) be satisfied. Given the conditions of Theorem (ref) it holds under (ref) \begin{equation*} (\sqrt 2\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{np}}-\mu_\m\big)\stackrel{d}{\rightarrow}\cN\Big(2^{-1/2}\sum_{j= 1}^\infty\xi _j^2,1\Big). \end{equation*} Given the conditions of Theorem (ref) it holds under (ref) where $\varsigma_\m$ is replaced by $1$ that \begin{equation*} n\, S_n^{\textsl{np}}\stackrel{d}{\rightarrow}\sum_{j=1}^\infty\lambda_j^\textsl{np}\,\chi_{1j}^2(\delta_j/\lambda_j^\textsl{np}). \end{equation*}

In the next proposition, we establish consistency of our test when $H_{\textsl{np}}$ does not hold, that is, the solution to (ref) does not belong to $\cF_\sw^\sr$ for any sequence $\sw$ satisfying Assumption (ref) and any sufficiently large constant $0<\sr<\infty$.

propAssume that $H_{\textsl{np}}$ does not hold. Let Assumption (ref) (i) hold true. Let $(\alpha_n)_{n\geq 1}$ be as in Proposition (ref). Under the conditions of Theorem (ref) and (ref), respectively, we have \begin{gather*} \PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big( nS_n^{\textsl{np}}-\mu_\m\big)>\alpha_n\Big)=1+o(1),\\ \PP\big(nS_n^{\textsl{np}}>\alpha_n\big)=1+o(1). \end{gather*}

In the following we show that our tests are consistent uniformly over the function class

equation*[equation* omitted — 158 chars of source]

where $\sol_0\in\cF_\sw^\rho$ solves (ref) and $C>0$ is a finite constant. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{np}}\,\chi_{1j}^2$, respectively.

propLet Assumption (ref) be satisfied. For any $\varepsilon>0$, any $0<\alpha<1$, and any sufficiently large constant $\rho>0$ we have under the conditions of Theorem (ref) \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cJ_n^\sr}\PP\Big((\sqrt 2\,\varsigma_\m)^{-1}\big(n\,S_n^{\textsl{np}}-\mu_\m\big)>q_{1\alpha}\Big)\geq 1-\varepsilon,\\ \end{equation*} whereas under the conditions of Theorem (ref) it holds \begin{equation*} \lim_{n\to\infty}\inf_{\sol\in\cJ_{n}^\sr}\PP\big(n\,S_n^{\textsl{np}}>q_{2\alpha}\big)\geq 1-\varepsilon. \end{equation*}

Monte Carlo simulation

In this section, we study the finite-sample performance of our test by presenting the results of Monte Carlo experiments. There are $1000 $ Monte Carlo replications in each experiment. Results are presented for the nominal level $0.05$. Realizations of Y were generated from

equation[equation omitted — 50 chars of source]

for some constant $c_U>0$ specified below. The structural function $\sol$ and the joint distribution of $(Z,W,U)$ varies in the experiments below. As basis $\{f_j\}_{j\geq 1}$ we choose cosine basis functions given by $f_{j}(t)=\sqrt{2}\cos(\pi j t)$ for $j=1,2,\dots$ throughout this simulation study. \paragraph{Parametric Specification} Let us investigate the finite sample performance of our tests in the case of parametric specifications. Realizations $(Z,W)$ were generated by $ W \sim \cU[0,1]$, $Z = (\xi\, W +(1-\xi)\, \varepsilon)^2$ where $\xi=0.8$ and $\varepsilon\sim\cN(0.5, 0.1)$. Moreover, let $U = \kappa\,\varepsilon + \sqrt{1-\kappa^2}\,\epsilon$ with $\kappa=0.3$ and $\epsilon\sim N(0, 1)$. Then realizations of $Y$ where generated by (ref) with $c_U=0.2$ by an either linear function

equation[equation omitted — 44 chars of source]

a polynomial of second degree

equation[equation omitted — 49 chars of source]

or a polynomial of third degree

equation[equation omitted — 63 chars of source]

Given (ref) is the correct model, then $\theta_3=1.5$ if (ref) is the null model and $\theta_3=3$ if (ref) is the null model. In Table 1 we depict the empirical rejection probabilities when using $S_n^\textsl{p}$ with additional smoothing where either $\tau_j= j^{-1}$ or $\tau_j= j^{-2}$, $j\geq 1$, which we denote by $S_n^{1\textsl{p}}$ or $S_n^{2\textsl{p}}$, respectively.

table[table omitted — 1,700 chars of source]

When $\tau_j=j^{-1}$ then the number of basis functions used is $m=200$ while in the case of $\tau_j=j^{-2}$ a choice of $m=100$ is sufficient. The critical values are estimated as described in Remark (ref) where $M=150$ if $\tau_j=j^{-1}$ and $M=100$ if $\tau_j=j^{-2}$. This choice of $M$ ensures that the estimated eigenvalues $\widehat\lambda_j$ are sufficiently close to zero for all $j\geq M$. We compare our test statistic with the test of Horowitz2006. We follow his implementation using biweight kernels. The bandwidth used to estimate the joint density of $(Z,W)$ was also selected by cross validation. As Table 1 illustrates, the results for $S_n^{1\textsl{p}}$ and $S_n^{2\textsl{p}}$ are quite similar. In both situations, our test is more powerful than the test of Horowitz2006 when testing (ref) against (ref). In this simulation study, we observed that the estimated coefficients of $T(\sol-\phi(\thet,\cdot))$ have a fast decay. Consequently, the test statistic $S_n$ with no weighting has less power, as we discussed in Subsection 2.4. In contrast, we will demonstrate by the end of this section that using weights can be inappropriate. \paragraph{Testing Exogeneity} We now turn to the test of exogeneity where the realizations $(Z,W)$ are generated by $W \sim\cU[0,1]$ and $Z = \xi\,W +\sqrt{1-\xi^2}\,\varepsilon$ with $\xi=0.7$, and $\varepsilon\sim \cU[0,1]$. Moreover, let $U = \kappa\,\varepsilon + \sqrt{1-\kappa^2}\,\epsilon$ with $\epsilon\sim \cU[0,1]$. Here, $\kappa$ measures the degree of endogeneity of $Z$ and is varied among the experiments. The null hypothesis $H_0$ holds true if $\kappa=0$ and is false otherwise. Now realizations of $Y$ where generated by (ref) with $c_U=1$ and the nonparametric structural function $\sol_1(z)=\sum_{j=1}^\infty (-1)^{j+1}\,j^{-1} \sin(j\pi z)$. For computational reasons we truncate the infinite sum at $K=100$. The resulting function is displayed in Figure (ref). We estimate the structural relationship using Lagrange polynomials. Indeed, only a few basis functions are necessary to accurately approximate the true function. If we choose $\k$ too small or too large then the estimator will be a poor approximate of the true structural function and hence, the test statistic will reject $H_{\textsl{np}}$. In this experiment we set $\k=4$ for $n=250$ and $n=500$.

table[table omitted — 1,071 chars of source]

In Table 2 we depict the empirical rejection probabilities when using $S_n^\textsl{e}$ with additional smoothing where either ${\tau_j} =j^{-1}$ or ${\tau_j}=j^{-2}$, $j\geq 1$, which we denote by $S_n^{1\textsl{e}}$ or $S_n^{2\textsl{e}}$, respectively. The critical values of these statistics are estimated as described in Remark (ref) with $M=50$ in case of $\tau_j=j^{-1}$ and $M=40$ in case of $\tau_j=j^{-2}$. We compare our results with the test of Blundell2007. We follow their approach by choosing the bandwidth of the joint density of $(Z,W)$ by cross validation. The bandwidth of the marginal of $Z$ is $n^{1/5-7/24}$ times the cross-validation bandwidth. As we see from Table 2, $S_n^{1\textsl{e}}$ is slightly more powerful than the test of Blundell2007. If we choose a stronger sequence, however, then our test statistic $S_n^{2\textsl{e}}$ becomes considerably more powerful. \paragraph{Nonparametric Specification} Let us now study the finite sample of our test in the case of nonparametric specification. We generate the pair $(Z,W)$ as in the parametric case described above. For the generation of the dependent variable $Y$ we distinguish two cases. Besides the structural function $\sol_1(z)=\sum_{j=1}^\infty (-1)^{j+1}j^{-2} \sin(j\pi z)$ we also consider the function $\sol_2(z)=\sum_{j=1}^\infty ((-1)^{j+1}+1)/4\,j^{-2} \sin(j\pi z)$. Again, for computational reasons we truncate the infinite sum at $K=100$. The resulting functions are displayed in Figure (ref). Further, $Y$ is generated by (ref) either with $\sol_1$ and $c_U=0.2$ or $\sol_2$ and $c_U=0.8$. In both cases, we estimate the structural relationship using Lagrange polynomials with $\k=4$ for $n=500$ and $n=1000$.

If $H_{\textsl{np}}$ is false then $\Ex[U|W]\neq 0$ and we let $\Ex[U|W]=\Ex[\rho(Z)|W]$ where $\rho$ is defined below. Consequently, when $H_{\textsl{np}}$ is false we generate realizations of $Y$ from

equation*[equation* omitted — 45 chars of source]

for $l=1,2$ and $j\geq 1$ where $\rho_j(z)=c_j(\exp(2jz)\1_{\{z\leq 1/2\}}+\exp(2j(1-z))\1_{\{z>1/2\}}-1)$ and $c_j$ is a normalizing constant such that $\int_0^1 \rho_j(z)dz=0.5$. The functions $\rho_j$ are continuous but not differentiable at $0.5$. Roughly speaking, the degree of roughness of $\rho_j$ is larger for larger $j$.

figure[figure omitted — 142 chars of source]
table[table omitted — 1,150 chars of source]

In Table 3, we depict the empirical rejection probabilities when using $S_n^\textsl{np}$ with either no smoothing or additional smoothing ${\tau_j}=j^{-2}$, $j\geq 1$, which we denote by $S_n^{0\textsl{np}}$ or $S_n^{2\textsl{np}}$, respectively. When no additional smoothing is applied then the number of basis functions $f_j$ is given by $\m=11$ if $n=500$ and $\m=15$ if $n=1000$ and hence, the choice of $\m$ is slightly larger than $n^{1/3}$ as suggested by the theoretical results. The critical values of these statistics are estimated as described in Remark (ref) where in the case of $S_n^{2\textsl{np}}$ we choose $M=100$. We compare our results with the test of Horowitz2009. We observe that the statistic $S_n^{0\textsl{np}}$ is less powerful than $S_n^{2\textsl{np}}$ against the alternatives $\rho_1$ and $\rho_2$.

In the following, we illustrate that using additional weighting can be inappropriate. Table 4 illustrates the power of our tests when the structural function $\sol_2$ is considered and realizations $(Z,W)$ were generated by $ W \sim \cU[0,1]$, $Z = (0.8\, W +0.3\, \varepsilon)^2$ where $\varepsilon\sim\cN(0.5, 0.05)$. In this case, we generate $Y$ using (ref) where $c_U=0.8$. In this case, the estimates of the generalized coefficients of $T(\sol-\sol_0)$ are more fluctuating and using weights is not appropriate here. Indeed, as we can see from Table 4, the test statistic $S_n^{0\textsl{np}}$ with no smoothing is more powerful than $S_n^{2\textsl{np}}$ were weighting $\tau_j=j^{-2}$, $j\geq 1$, is used. In particular, $S_n^{0\textsl{np}}$ is much more powerful than the test of Horowitz2009.

table[table omitted — 1,167 chars of source]

Conclusion

Based on the methodology of series estimation, we have developed in this paper a family of goodness-of-fit statistics and derived their asymptotic properties. The implementation of these statistics is straightforward. We have seen that the asymptotic results depend crucially on the choice of the smoothing operator $L$. By choosing a stronger decaying sequence $\tau$, our test becomes more powerful with respect to local alternatives but might lose desirable consistency properties. We gave heuristic arguments how to choose the weights in practice. In addition, in a Monte Carlo investigation our tests perform well in finite samples.