Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,073 characters · 18 sections · 51 citation commands
Goodness-of-Fit Tests based on Series Estimators in Nonparametric Instrumental Regression
\vskip 1cm
While parametric instrumental variables estimators are widely used in econometrics, its nonparametric extension has not been introduced until the last decade. The study of nonparametric instrumental regression models was initiated by F03eswc and NP03econometrica. In these models, given a scalar dependent variable $Y$, a vector of regressors $Z$, and a vector of instrumental variables $W$, the structural function $\sol$ satisfies
for an error term $U$. Here, $Z$ contains potentially endogenous entries, that is, $\Ex[U|Z]$ may not be zero. Model (ref) does not involve the a priori assumption that the structural function is known up to finitely many parameters. By considering this nonparametric model, we minimize the likelihood of misspecification. On the other hand, implementing the nonparametric instrumental regression model can be challenging.
Nonparametric instrumental regression models have attracted increasing attention in the econometric literature. For example, AC03econometrica, BCK07econometrica, ChenReiss2008, NP03econometrica or Johannes2009 consider sieve minimum distance estimators of $\sol$, while DFR02, HallHorowitz2005, GS06 or FJVB07 study penalized least squares estimators. When the methods of analysis are widened to include nonparametric techniques, one must confront two mayor challenges. First, identification in model (ref) requires far stronger assumptions about the instrumental variables than for the parametric case (cf. NP03econometrica). Second, the accuracy of any estimator of $\sol$ can be low, even for large sample sizes. More precisely, ChenReiss2008 showed that for a large class of joint distributions of $(Z,W)$ only logarithmic rates of convergence can be obtained. The reason for this slow convergence is that model (ref) leads to an inverse problem which is ill posed in general, that is, the solution does not depend continuously on the data.
In light of the difficulties of estimating the nonparametric function $\sol$ in model (ref), the need for statistically justified model simplifications is paramount. We do not face an ill posed inverse problem if a parametric structure of $\sol$ or exogeneity of $Z$ can be justified. If these model simplifications are not supported by the data, one might still be interested in whether a smooth solution to model (ref) exists and if some regressors could be omitted from the structural function $\sol$. These model simplifications have important potential since they might increase the accuracy of estimators of $\sol$ or lower the required conditions imposed on the instrumental variables to ensure identification.
In this work we present a new family of goodness-of-fit statistics which allows for several restricted specification tests of the model (ref). Our method can be used for testing either a parametric or nonparametric specification. In addition, we perform a test of exogeneity and of dimension reduction of the vector of regressors $Z$, that is, whether certain regressors can be omitted from the structural function $\sol$. By a withdrawal of regressors which are independent of the instrument, identification in the restricted model might be possible although $\sol$ is not identified in the original model (ref).
There is a large literature concerning hypothesis testing of restricted specification of regression. In the context of conditional moment equation, Donald2003 and Tripathi2003 make use of empirical likelihood methods to test parametric restrictions of the structural function. In addition, Santos12 allows for different hypothesis tests, such as a test of homogeneity. Based on kernel techniques, Horowitz2006, Blundell2007, and Horowitz2011 propose test statistics in which an additional smoothing step (on the exogenous entries of $Z$) is carried out. Horowitz2006 considers a parametric specification test. Blundell2007 establish a consistent test of exogeneity of the vector of regressors $Z$, whereas Horowitz2011 tests whether the endogenous part of $Z$ can be omitted from $\sol$. GS07 and Horowitz2009 develop nonparametric specification tests in an instrumental regression model. We like to emphasize that their test cannot be applied to model (ref) where some entries of $Z$ might be exogenous.
Our testing procedure is entirely based on series estimation and hence is easy to implement. We use approximating functions to estimate the conditional moment restriction implied by the model (ref) where $\sol$ is replaced by an estimator under each conjectured hypothesis. It is worth noting that by our methodology we can omit some assumptions typically found in related literature, such as smoothness conditions on the joint distribution of $(Z,W)$. In addition, a Monte Carlo indicates that the finite sample power of our tests exceed that of existing tests.
The paper is organized as follows. In Section 2, we start with a simple hypothesis test, that is, whether $\sol$ coincides with a known function $\sol_0$. We obtain the test's asymptotic distribution under the null hypothesis and its consistency against any fixed alternative model. Moreover, we judge its power by considering linear local alternatives and establish uniform consistency over a class of functions. In Sections 3--5 we consider a parametric specification test, a test of exogeneity, and a nonparametric specification test. The goodness-of-fit statistics are obtained by replacing $\sol_0$ in the statistic of Section 2 by an appropriate estimator. In each case, the asymptotic distribution under correct specification and power statements against alternative models are derived. In Section 6, we investigate the finite sample properties of our tests by Monte Carlo simulations. All proofs can be found in the appendix.
In this section, we propose a goodness-of-fit statistic for testing the hypothesis $ H_0:\,\sol=\sol_0$, where $\sol_0$ is a known function, against the alternative $\sol\neq\sol_0$. We develop a test statistic based on $\cL^2$ distance. As we will see in the following chapters, it is sufficient to replace $\sol_0$ by an appropriate estimator to allow for tests of the general model against other specifications. We first give basic assumptions, then obtain the asymptotic distribution of the proposed statistic, and further discuss its power and consistency properties.
\paragraph{The model revisited} The nonparametric instrumental regression model (ref) leads to a linear operator equation. To be more precise, let us introduce the conditional expectation operator $T\phi:=\Ex[\phi(Z)|W]$ mapping $\cL^2_Z=\{\phi:\,\Ex|\phi(Z)|^2<\infty\}$ to $\cL^2_W=\{\psi:\, \Ex|\psi(W)|^2<\infty\}$. Consequently, model (ref) can be written as
where the function $g:=\Ex[Y | W]$ belongs to $\cL^2_W$. Throughout the paper we assume that an iid. $n$-sample of $(Y,Z,W)$ from the model (ref) is available. \paragraph{Assumptions.} Our test statistic based on a sequence of approximating functions $\{\basW_l\}_{l\geq 1}$ in $\cL^2_W$. Let $\cW$ denote the support of $W$ and the marginal density of $W$ by $p_W$. Let $\nu$ be a probability density function that is strictly positive on $\cW$. We assume throughout the paper that $\{\basW_l\}_{l\geq 1}$ forms an orthonormal basis in $\cL_\nu^2(\mathbb R^{d_w}):=\{\phi:\,\int \phi^2(s)\nu(s)ds<\infty\}$ where $d_w$ denotes the dimension of $W$. For instance, if $\cW\subset[a,b]$ then a natural choice of $\nu$ would be $\nu(w)=1/(b-a)$ for $w\in[a,b]$ and zero otherwise.
Assumption (ref) $(i)$ restricts the magnitude of the approximating functions $\{f_j\}_{j\geq1}$ which is necessary for our proof to determine the asymptotic behavior of our test statistic. This assumption holds for sufficiently large $\eta_f$ if the basis $\{f_l\}_{l\geq1}$ is uniformly bounded, such as trigonometric bases. Moreover, Assumption (ref) $(i)$ is satisfied by Hermite polynomials. Assumption (ref) $(ii)$ is satisfied if, for instance, $p_W/\nu$ is continuous and $\cW$ is compact.
The results derived below involve assumptions on the conditional moments of the random variables $U$ given $W$ gathered in the following assumption.
The conditional moment condition on the error term $U$ helps to establish the asymptotic distribution of our test statistics. The following assumption ensures identification of $\sol$ in the model (ref).
Under Assumption (ref), the hypothesis $H_0$ is equivalent to $g=T\sol_0$ which is used to construct our test statistic below. Note that the asymptotic results under null hypotheses considered in Sections 2--4 hold true even if $T$ is singular. If Assumption (ref) fails, however, our test has no power against alternative models whose structural function satisfies $\sol=\sol_0+\delta$ with $\delta$ belonging to the null space of $T$.
We will see below that the power of our test can be increased by carrying out an additional smoothing step. Therefore, we introduce a smoothing operator $L$ mapping $\cL_W^2$ to $\cL_W^2$. In contrast to the unknown conditional expectation operator $\op$, which has to be estimated, the operator $L$ can be chosen by the econometrician. Let $L$ have an eigenvalue decomposition given by $\{\tau_j^{1/2},f_j\}_{j\geq 1}$. We allow in this paper for a wide range of smoothing operators. In particular, $L$ may be the identity operator, that is, no smoothing step is carried out. We only require the following condition on the operator $L$ determined by the sequence of eigenvalues $\tau=(\tau_j)_{j\geq 1}$.
Assumption (ref) ensures that the operator $L$ is nonsingular.
\paragraph{Notation.} For a matrix $A$ we denote its transposed by $A^t$, its inverse by $A^{-1}$, and its generalized inverse by $A^{-}$. The euclidean norm is denoted by $\|\cdot\|$ which in case of a matrix denotes the spectral norm, that is $\|A\|=(\text{trace}(A^tA))^{1/2}$. The norms on $L_Z^2$ and $L_W^2$ are denoted by $\|\phi\|_Z^2:=\Ex|\phi(Z)|^2$ for $\phi\in L_Z^2$ and $\|\psi\|_W^2:=\Ex|\psi(W)|^2$ for $\psi\in L_W^2$. The $k\times k$ identity matrix is denoted by $I_k$. For a vector $V$ we write diag$(V)$ for the diagonal matrix with diagonal elements being the values of $V$. Moreover, $\basZ_\um(Z)$ and $\basW_\um(W)$ denote random vectors with entries $\basZ_j(Z)$ and $\basW_j(W)$, $1\leq j\leq m$, respectively. For any weighting sequence $w$ we introduce vectors $\basZ_\um^w(Z)$ and $\basW_\um^w(W)$ with entries $\basZ_j^w(Z)=\sqrt{w_j}e_j(Z)$ and $\basW_j^w(W)=\sqrt{w_j}f_j(W)$, $1\leq j\leq m$. We write $a_n\sim b_n$ when there exist constants $c,c' > 0$ such that $c b_n\leq a_n\leq c' b_n$ for all sufficiently large $n$.
Nonsingularity of the conditional expectation operator $T$ and the smoothing operator $L$ implies that the null hypothesis $H_0$ is equivalent to $L(g-T\sol_0)=0$. Note that $\|L(g-T\sol_0)\|_W=0$ if and only if $\int\big|L(g-T\sol_0)(w)p_W(w)/\nu(w)\big|^2\nu(w)dw=0$ since the Lebesgue measure $\nu$ is strictly positive on $\cW$. Moreover, since $\{f_j\}_{j\geq 1}$ is an orthonormal basis with respect to $\nu$ we obtain by Parseval's identity
Now we truncate the infinite sum at some integer $\m$ which grows with the sample size $n$. This ensures consistency of our testing procedure. Further, replacing the expectation by sample mean we obtain our test statistic
We reject the hypothesis $H_0$ if $n\,S_n$ becomes too large. When no additional smoothing is carried out, that is, $L$ is the identity operator, then $\tau_j=1$ for all $j\geq 1$. To achieve asymptotic normality we need to standardize our test statistic $S_n$ by appropriate mean and variance, which we introduce in the following definition.
Indeed the next result shows that $nS_n$ after standardization is asymptotically normally distributed if $\m$ increases appropriately as the sample size $n$ tends to infinity.
In the following result, we establish the asymptotic distribution of our test when the sequence of weights $\tau$ may have a stronger decay than in Theorem (ref), that is, we consider the case where $\tau$ satisfies $\sum_{j=1}^\m\tau_j=O(1)$. This holds, for instance, if the sequence $\tau$ satisfies $\tau_j\sim j^{-(1+\varepsilon)}$ for any $\varepsilon>0$. In this case, the asymptotic distribution changes and additional definitions have to be made. Let $\varSigma$ be the covariance matrix of the infinite dimensional centered vector $\big(Uf_j^\tau(W)\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma$ are denoted by $(\lambda_j)_{j\geq1}$. Below, we introduce a sequence $\{\chi_{1j}^2\}_{j\geq 1}$ of independent random variables that are distributed as chi-square with one degree of freedom.
Let us study the power of the test statistic $S_n$, that is, the probability to reject a false hypothesis, against a sequence of linear local alternatives that tends to zero as $n\to\infty$. It is shown that the power of our tests essentially relies on the choice of the weighting sequence $\tau$.
Let us start with the case $\varsigma_\m^{-1}=o(1)$. We consider the following sequence of linear local alternatives
for some function $\delta\in \cL_Z^4:=\{\phi:\Ex|\phi(Z)|^4<\infty\}$. The next result establishes asymptotic normality for the standardized test statistic $S_n$. Let us denote $\delta_j:=\sqrt{\tau_j}\,\Ex[\delta(Z)f_j(W)]$.
As we see below the test statistic $S_n$ has power advantages if $\sum_{j=1}^\m\tau_j=O(1)$. Let us consider the sequence of linear local alternatives
for some function $\delta\in \cL_Z^4$. For the next result, the sequence $\{\chi_{1j}^2(\delta_j/\lambda_j)\}_{j\geq 1}$ denotes independent random variables that are distributed as non-central chi-square with one degree of freedom and non-centrality parameters $\delta_j/\lambda_j$.
In this subsection, we establish consistency against a fixed alternative and uniform consistency of our test over appropriate function classes. Let us first consider the case of a fixed alternative. We assume that $H_0$ does not hold, that is, $\PP(\sol=\sol_0)<1$. The following proposition shows that our test has the ability to reject a false null hypothesis with probability $1$ as the sample size grows to infinity.
The consistency properties require the following additional assumption.
Assumption (ref) $(i)$ implies that $\|LT(\sol-\sol_0)\|_W>0$ for any structural function $\sol$ in the alternative. Further, Assumption (ref) implies that $\sum_{j=1}^\m\tau_j^2=O(\varsigma_\m^2)$.
In the following, we specify a class of functions over which our test $S_n$ is uniformly consistent. This essentially implies that there are no alternative functions in this class over which our test has low power. We show that our test is consistent uniformly over the class
where $C>0$ is a finite constant. Clearly, if $H_0$ is false then $\|LT(\sol-\sol_0)\|_W^2\geq \sr\,\varsigma_\m n^{-1}$ for all sufficiently large $n$ and some $\rho>0$. By Assumption (ref) the sequence $\tau$ is nonincreasing sequence with $\tau_1=1$ and hence, $\|LT(\sol-\sol_0)\|_W^2\leq \|T(\sol-\sol_0)\|_W^2\leq \|\sol-\sol_0\|_Z^2$ by Jensen's inequality. We conclude that $\cG_n^\sr$ contains all alternative functions whose $\cL_Z^2$-distance to the structural function $\sol_0$ is at least $n^{-1}\varsigma_\m$ within a constant. If the coefficients $\Ex[(\sol-\sol_0)(Z)f_j(W)]$ fluctuate for large $j$ then $\sol$ does not belong to $\cG_n^\sr$ if the decay of $\tau$ is too strong. On the other hand, if $\Ex[(\sol-\sol_0)(Z)f_j(W)]$ is sufficiently small for $j$ up to a finite constant then $\sol$ does not necessarily belong to $\cG_{n}^\sr$ with $\tau$ having a slow decay. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j\,\chi_{1j}^2$, respectively.
In this section, we present a test whether the structural function $\sol$ is known up to a finite dimensional parameter. Let $\varTheta$ be a compact subspace of $\mathbb R^k$ then we consider the null hypothesis $H_{\text{p}}:$ there exists some $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$ for a known function $\phi$. The alternative hypothesis is that there exists no $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$ holds true.
Under Assumptions (ref) and (ref), the null hypothesis $H_{\text{p}}$ is equivalent to $L(g-T\phi(\cdot,\vartheta))=0$ for some $\vartheta\in\varTheta$. Thereby, to verify $H_{\text{p}}$ we make use of the test statistic $S_n$ given in (ref) where $\sol_0$ is replaced by $\phi(\cdot,\hthet)$ with $\hthet$ being an estimator of $\vartheta$. Hence, our test statistic for a parametric specification is given by
If the test statistic $S_n^{\text{p}}$ becomes too large then $H_{\text{p}}$ has to be rejected. To obtain asymptotic results for the statistic $S_n^{\text{p}}$ we require smoothness conditions of the function $\phi$ with respect to its second argument. Below we denote the vector of partial derivatives of $\phi$ with respect to $\vartheta=(\vartheta_1,\dots,\vartheta_k)^t$ by $\phi_\vartheta=(\phi_{\vartheta_l})_{1\leq l\leq k}$ and the matrix of second-order partial derivatives by $\phi_{\vartheta\vartheta}=(\phi_{\vartheta_j\vartheta_l})_{1\leq j,l\leq k}$.
The following proposition establishes asymptotic normality of $S_n^{\text{p}}$ after standardization.
In the following theorem, we state the asymptotic distribution of $nS_n^{\textsl{p}}$ when $\sum_{j=1}^\m\tau_j=O(1)$. In this case, we assume that $\hthet$ satisfies under $H_{\textsl{p}}$
where $V_{i}:=(Y_i,Z_i,W_i,\thet)$ and $h_\uk(V_{i})=(h_1(V_i),\dots, h_k(V_i))^t$ where $h_j$, $1\leq j\leq k$, are real valued functions. It is well known that this representation holds if $\hthet$ is the generalized method of moments estimator. Let $\varSigma^\textsl{p}$ be the covariance matrix of the infinite dimensional centered vector $\big(Uf_j^\tau(W)-\Ex[f_j^\tau(W)\phi_\vartheta(Z,\thet)^t]h_\uk(V)\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{p}$ are denoted by $(\lambda_j^\textsl{p})_{j\geq1}$.
In the following, we study the power and consistency properties of the test statistic $S_n^{\text{p}}$. In the following, we consider a sequence of linear local alternatives (ref) or (ref) with $\sol_0=\phi(\thet,\cdot)$. Further, let $\delta_\perp$ denote the projection of $\delta$ onto the orthogonal complement of $\phi(\cdot,\thet)$; that is, $\Ex[\phi_\vartheta(Z,\thet)\delta_\perp(Z)]=0$. Let us denote $\delta_{j\perp}:=\sqrt{\tau_j}\,\Ex[\delta_\perp(Z)f_j(W)]$.
The next proposition establishes consistency of our test against a fixed alternative model. It is assumed that $H_\text p$ is false, that is, there exists no $\vartheta\in\varTheta$ such that $\sol(\cdot)=\phi(\cdot,\vartheta)$. In this situation, $\thet$ denotes the probability limit of the estimator $\hthet$.
In the following, we show that $S_n^{\textsl{p}}$ is consistent uniformly over the function class
for some constant $C>0$ and $\thet$ denotes the probability limit of $\hthet$. Similarly as in the previous section, it can be seen that $\cH_{n}^\sr$ only contains functions whose $\cL_Z^2$ distance to $\phi(\cdot,\vartheta_0)$ is at least $n^{-1}\varsigma_\m$ within a constant. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{p}}\,\chi_{1j}^2$, respectively.
Endogeneity of regressors is a common problem in econometric applications. Falsely assuming exogeneity of the regressors leads to inconsistent estimators. On the other hand, treating exogenous regressors as if they were endogenous can lower the accuracy of estimation dramatically. In this section, we propose a test whether the vector of regressors $Z$ is exogenous, that is, $\Ex[U|Z]=0$ or equivalently $\sol(Z)=\Ex[Y|Z]$. In this section, let $\sol_0(Z)=\Ex[Y|Z]$ then the hypothesis under consideration is given by $H_{\textsl{e}}:\, \sol=\sol_0$. The alternative hypothesis is that $\sol\neq \sol_0$.
To establish a test of exogeneity, let us first introduce an estimator of the conditional mean of $Y$ given $Z$. This estimator is based on a sequence of approximating functions $\{e_j\}_{j\geq 1}$ belonging to $\cL_Z^2$. Further, let $\bold Z_k$ denote a $n\times k$ matrix with entries $e_j(Z_i)$ for $1\leq i\leq n$ and $1\leq j\leq k$. Moreover, let $\bold Y_n=(Y_1,\dots,Y_n)^t$. Then we define the estimator
In contrast to the parametric case we need to allow for $k$ tending to infinity as $n\to\infty$ in order to ensure consistency of the estimator $\osol_k$. Under conditions given below $\bold Z_\k^t\bold Z_\k$ will be nonsingular with probability approaching one and hence its generalized inverse will be the standard inverse. Note that the asymptotic behavior of the estimator $\osol_k$ was studied, for example, by Newey1997.
Under Assumptions (ref) and (ref), the null hypothesis $H_{\textsl{e}}$ is equivalent to $L(g-T\sol_0)=0$. Consequently, our test of exogeneity of $Z$ is based on the goodness-of-fit statistic $S_n$ introduced in (ref) but where $\sol_0$ is replaced by the series estimator $\osol_\k$. The proposed test statistic for $H_{\textsl{e}}$ is now given by
where $\k$ and $\m$ tend to infinity as $n\to\infty$. The hypothesis of exogeneity of $Z$ has to be rejected if $S_n^{\textsl{e}}$ becomes too large.
For controlling the bias of the estimator $\osol_\k$ we specify in the following a rate of approximation (cf. Newey1997). Let $\sw=(\sw_j)_{j\geq 1}$ be a nondecreasing sequence with $\sw_1=1$. We assume that $\sol_0$ belongs to
Here, the sequence of weights $\sw$ measures the approximation error of $\sol_0$ with respect to the functions $\{e_j\}_{j\geq1}$.
Assumption (ref) $(i)$ determines the required asymptotic behavior of the rate $\sw$. For splines and power series this assumption is satisfied if the number of continuous derivatives of $\sol_0$ divided by the dimension of $Z$ equals two. Assumption (ref) $(ii)$ and $(iii)$ restrict the magnitude of the approximating functions $\{e_j\}_{j\geq1}$ and impose nonsingularity of their second moment matrix.
We are now in the position to proof the following asymptotic result for the standardized test statistic $S_n^\text e$. Here, a key requirement is that $\k=o(\varsigma_\m)$ implying that $\k=o(\sum_{j=1}^\m\tau_j)$ and, in particular, $\k=o(\m)$ if the smoothing operator $L$ is the identity.
The next result states an asymptotic distribution result for the statistic $S_n^{\textsl{e}}$ if $\sum_{j=1}^\m\tau_j=O(1)$. Let $\varSigma^\textsl{e}$ be the covariance matrix of the infinite dimensional centered vector $\big(U(f_j^\tau(W)-\sum_{l\geq 1}\Ex[f_j^\tau(W)e_l(Z)]e_l(Z))\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{e}$ are denoted by $(\lambda_j^\textsl{e})_{j\geq1}$.
Similar to the previous sections we study the power and consistency properties of our test. Let us study the power of our test of exogeneity under linear local alternatives (ref) or (ref). In these cases, it holds $\Ex[U|W]=0$ but $\Ex[U|Z]= -\varsigma_\m^{1/2}n^{-1/2} \delta(Z)$ under (ref) or $\Ex[U|Z]=-n^{-1/2}\delta(Z)$ under (ref).
Let us now establish consistency of our tests when $H_\text e$ does not hold, that is, $\PP\big(\sol=\sol_0\big)<1$.
In the following we show that our tests are consistent uniformly over the function class
form some constant $C>0$. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{e}}\,\chi_{1j}^2$, respectively.
A solution to the linear operator equation (ref) only exists if $g$ belongs to the range of $\op$. This might be violated if, for instance, the instrument is not valid, that is, $\Ex[U|W]\neq 0$. In many economic applications a priori smoothness restriction on the unknown function can be justified which we capture by a set of functions $\cF$. We consider the hypothesis $H_{\textsl{np}}$: there exists a solution $\sol_0\in\cF$ to (ref). The alternative hypothesis is that there exists a solution (ref) which does not belong to $\cF$. Under the alternative only unsmooth functions solve the conditional moment restriction which can be interpreted as a failure of validity of the instrument $W$. We see in this section that our results allow also for a test of dimension reduction of the vector of regressors $Z$, that is, whether some regressors can be omitted from the structural function $\sol_0$.
\paragraph{The nonparametric estimator.} In the following, we derive an estimator of $\sol_0$ under the null hypothesis $H_{\textsl{np}}$. For simplicity, assume that $\cZ=\cW$ and consider a sequence $\{e_j\}_{j\geq 1}$ of approximating functions which are orthonormal on $\cZ$ with respect to the Lebesque measure $\nu$. Under conditions given below, $\sol_0$ has the expansion $\sol_0(\cdot)=\sum_{l=1}^\infty \int\sol_0(z)e_l(z)\nu(z)dz\,e_l(\cdot)$. Thereby, the conditional moment restriction under $H_{\textsl{np}}$ leads to the following unconditional moment restrictions
for $j \geq 1$. This motivates the following orthogonal series type estimator. Let $\bold Z_k$ and $\bold Y_n$ be as in the previous section and let $\bold X_k$ denote a $n\times k$ matrix with entries $e_j(W_i)$ for $1\leq i\leq n$ and $1\leq j\leq k$. Then for each $k\geq 1$ we consider the estimator
Under conditions given below $\bold X_\k^t\bold Z_\k$ will be nonsingular with probability approaching one and hence its generalized inverse will be the standard inverse. The nonparametric estimator $\hsol_k$ given in (ref) was studied by Johannes2009, Horowitz2011, and Horowitz2009.
\paragraph{Additional assumptions.} In the following, we specify a priori smoothness assumptions captured by the set $\cF$. As noted by Horowitz2009, uniformly consistent testing of $H_{\textsl{np}}$ is only possible if the null is restricted that any solution to (ref) is smooth. Here, we assume that under the null hypothesis $\sol_0$ belongs to the ellipsoid $\cF:=\cF_{\sw}^\sr := \big\{\phi\in \cL_Z^2: \sum_{j= 1}^\infty\sw_j\Ex[\phi(Z)e_j(Z)]^2\leq \sr\big\}$. As in the previous section, $\sw=(\sw_j)_{j\geq 1}$ measures the approximation error of $\sol_0$ with respect to the basis $\{e_j\}_{j\geq 1}$.
Further, as usual in the context of nonparametric instrumental regression, we specify some mapping properties of the conditional expectation operator $T$. Denote by $\cT$ the set of all nonsingular operators mapping $\cL_Z^2$ to $\cL^2_W$. Given a sequence of weights $\Opw:=(\Opw_j)_{j\geqslant 1}$ and $\Opd\geqslant 1$ we define the subset $\Opwd$ of $\cT$ by
If $p_Z/\nu$ is bounded from above and $p_W/\nu$ is uniformly bounded away from zero then the conditional expectation operator $T$ belongs to $\Opwd$ with $\opw_j=1$, $j\geq 1$, due to Jensen's inequality. Notice that for all $T\in\Opwd$ it follows that $\normV{T\basZ_j}_W^2\leq d\,\eta_p \Opw_j$ and thereby, the condition $T\in\Opwd$ links the operator $T$ to the basis $\{e_j\}_{j\geq 1}$. In the following, we denote $[\op]_\uk=\Ex[e_\uk(W)e_\uk(Z)^t]$ which is assumed to be a nonsingular matrix. In what follows, we introduce a stronger condition on the basis $\{e_l\}_{l\geq1}$. We denote by $\cTdDw$ for some $\tD\geq \td$ the subset of $\cTdw $ given by
The class $\cTdDw$ only contains operators $T$ whose off-diagonal elements of $[T]_\uk^{-1}$ are sufficiently small for all $k\geq 1$. A similar diagonality restriction has been used by HallHorowitz2005 or BJ2011. Besides the mapping properties for the operator $T$ we need a stronger assumption for the basis under consideration. The following condition gathers conditions on the sequences $\sw$ and $\opw$.
Due to Assumption (ref) $(iv)$ the degree of additional smoothing for our testing procedure must not be stronger than the degree of ill-posedness implied by the conditional expectation operator $T$. Under similar assumptions as above, Johannes2009 show that mean integrated squared error loss of $\hsol_{\k}$ attains the optimal rate of convergence $\cR_n:=\max\big(\sw_{\k}^{-1}, \sum_{j=1}^{\k}(n\opw_j)^{-1}\big)$. Due to Assumption (ref) $(v)$ we do not require orthonormal bases with respect to the unknown distribution $(Z,W)$ (cf. Remark 3.2 of BJ2011).
As in the previous sections, our test is based on the observation that the null hypothesis $H_{\textsl{np}}$ is equivalent to $L(g-T\sol_0)=0$. Our goodness-of-fit statistic for testing nonparametric specifications is given by $S_n$ where $\sol_0$ is replaced by the nonparametric estimator $\hsol_{\k}$ given in (ref), that is,
If $S_n^{\textsl{np}}$ becomes too large then there exists no function in $\cF_\sw^\sr$ solving (ref). The next result establishes asymptotic normality of $S_n^{\textsl{np}}$ after standardization. Again, a key requirement to obtain this asymptotic distribution is that $\k=o(\varsigma_\m)$ implying that $\k=o(\m)$ if the smoothing operator $L$ is the identity. This corresponds to the test of overidentification in the parametric framework where more orthogonality restrictions than parameters are required.
The next result states an asymptotic distribution of our test if $\sum_{j=1}^\m\tau_j=O(1)$. Let $\varSigma^\textsl{np}$ be the covariance matrix of the infinite dimensional centered vector $\big(U(f_j^\tau(W)-e_j^\tau(W))\big)_{j\geq 1}$. The ordered eigenvalues of $\varSigma^\textsl{np}$ are denoted by $(\lambda_j^\textsl{np})_{j\geq1}$.
In the mildly ill-posed case, the estimation precision suffers from the curse of dimensionality. Hence, by the test of dimension reduction of $Z$ we can increase the accuracy of estimation of $\sol_0$. On the other hand, in the severely ill-posed case the rate of convergence is independent of the dimension of $Z$ (cf. ChenReiss2008). As the next example illustrates, a dimension reduction test can also weaken the required restrictions on the instrument to obtain identification of $\sol$ in the restricted model.
Similar to the previous sections we study the power and consistency properties of our test. To study the power against local alternatives of the statistic $S_n^{\textsl{np}}$ we consider alternative models with the function $\sol_\k(\cdot)=e_\k(\cdot)^t[T]_\k^{-1}\Ex[Yf_\ukn(W)]$. We consider alternative models
for some function $\delta\in L_Z^4$ and where $\Ex[U|W]=0$. Let $\sol$ be such that $\Ex[Y-\sol(Z)|W]=0$. Due to (ref) $\sol$ does not belong $\cF_\sw^\rho$ and hence $H_{\textsl{np}}$ fails. Indeed, if $\sol\in \cF_\sw^\rho$ then we show in the appendix that $\|T(\sol-\sol_\k)\|_W^2=O(\opw_\k\sw_\k^{-1})=o(\varsigma_\m n^{-1})$ due to condition (ref) (or (ref)), which is in contrast to (ref).
In the next proposition, we establish consistency of our test when $H_{\textsl{np}}$ does not hold, that is, the solution to (ref) does not belong to $\cF_\sw^\sr$ for any sequence $\sw$ satisfying Assumption (ref) and any sufficiently large constant $0<\sr<\infty$.
In the following we show that our tests are consistent uniformly over the function class
where $\sol_0\in\cF_\sw^\rho$ solves (ref) and $C>0$ is a finite constant. For the next result let $q_{1\alpha}$ and $q_{2\alpha}$ denote the $1-\alpha$ quantile of $\cN(0,1)$ and $\sum_{j=1}^\infty\lambda_j^{\textsl{np}}\,\chi_{1j}^2$, respectively.
In this section, we study the finite-sample performance of our test by presenting the results of Monte Carlo experiments. There are $1000 $ Monte Carlo replications in each experiment. Results are presented for the nominal level $0.05$. Realizations of Y were generated from
for some constant $c_U>0$ specified below. The structural function $\sol$ and the joint distribution of $(Z,W,U)$ varies in the experiments below. As basis $\{f_j\}_{j\geq 1}$ we choose cosine basis functions given by $f_{j}(t)=\sqrt{2}\cos(\pi j t)$ for $j=1,2,\dots$ throughout this simulation study. \paragraph{Parametric Specification} Let us investigate the finite sample performance of our tests in the case of parametric specifications. Realizations $(Z,W)$ were generated by $ W \sim \cU[0,1]$, $Z = (\xi\, W +(1-\xi)\, \varepsilon)^2$ where $\xi=0.8$ and $\varepsilon\sim\cN(0.5, 0.1)$. Moreover, let $U = \kappa\,\varepsilon + \sqrt{1-\kappa^2}\,\epsilon$ with $\kappa=0.3$ and $\epsilon\sim N(0, 1)$. Then realizations of $Y$ where generated by (ref) with $c_U=0.2$ by an either linear function
a polynomial of second degree
or a polynomial of third degree
Given (ref) is the correct model, then $\theta_3=1.5$ if (ref) is the null model and $\theta_3=3$ if (ref) is the null model. In Table 1 we depict the empirical rejection probabilities when using $S_n^\textsl{p}$ with additional smoothing where either $\tau_j= j^{-1}$ or $\tau_j= j^{-2}$, $j\geq 1$, which we denote by $S_n^{1\textsl{p}}$ or $S_n^{2\textsl{p}}$, respectively.
When $\tau_j=j^{-1}$ then the number of basis functions used is $m=200$ while in the case of $\tau_j=j^{-2}$ a choice of $m=100$ is sufficient. The critical values are estimated as described in Remark (ref) where $M=150$ if $\tau_j=j^{-1}$ and $M=100$ if $\tau_j=j^{-2}$. This choice of $M$ ensures that the estimated eigenvalues $\widehat\lambda_j$ are sufficiently close to zero for all $j\geq M$. We compare our test statistic with the test of Horowitz2006. We follow his implementation using biweight kernels. The bandwidth used to estimate the joint density of $(Z,W)$ was also selected by cross validation. As Table 1 illustrates, the results for $S_n^{1\textsl{p}}$ and $S_n^{2\textsl{p}}$ are quite similar. In both situations, our test is more powerful than the test of Horowitz2006 when testing (ref) against (ref). In this simulation study, we observed that the estimated coefficients of $T(\sol-\phi(\thet,\cdot))$ have a fast decay. Consequently, the test statistic $S_n$ with no weighting has less power, as we discussed in Subsection 2.4. In contrast, we will demonstrate by the end of this section that using weights can be inappropriate. \paragraph{Testing Exogeneity} We now turn to the test of exogeneity where the realizations $(Z,W)$ are generated by $W \sim\cU[0,1]$ and $Z = \xi\,W +\sqrt{1-\xi^2}\,\varepsilon$ with $\xi=0.7$, and $\varepsilon\sim \cU[0,1]$. Moreover, let $U = \kappa\,\varepsilon + \sqrt{1-\kappa^2}\,\epsilon$ with $\epsilon\sim \cU[0,1]$. Here, $\kappa$ measures the degree of endogeneity of $Z$ and is varied among the experiments. The null hypothesis $H_0$ holds true if $\kappa=0$ and is false otherwise. Now realizations of $Y$ where generated by (ref) with $c_U=1$ and the nonparametric structural function $\sol_1(z)=\sum_{j=1}^\infty (-1)^{j+1}\,j^{-1} \sin(j\pi z)$. For computational reasons we truncate the infinite sum at $K=100$. The resulting function is displayed in Figure (ref). We estimate the structural relationship using Lagrange polynomials. Indeed, only a few basis functions are necessary to accurately approximate the true function. If we choose $\k$ too small or too large then the estimator will be a poor approximate of the true structural function and hence, the test statistic will reject $H_{\textsl{np}}$. In this experiment we set $\k=4$ for $n=250$ and $n=500$.
In Table 2 we depict the empirical rejection probabilities when using $S_n^\textsl{e}$ with additional smoothing where either ${\tau_j} =j^{-1}$ or ${\tau_j}=j^{-2}$, $j\geq 1$, which we denote by $S_n^{1\textsl{e}}$ or $S_n^{2\textsl{e}}$, respectively. The critical values of these statistics are estimated as described in Remark (ref) with $M=50$ in case of $\tau_j=j^{-1}$ and $M=40$ in case of $\tau_j=j^{-2}$. We compare our results with the test of Blundell2007. We follow their approach by choosing the bandwidth of the joint density of $(Z,W)$ by cross validation. The bandwidth of the marginal of $Z$ is $n^{1/5-7/24}$ times the cross-validation bandwidth. As we see from Table 2, $S_n^{1\textsl{e}}$ is slightly more powerful than the test of Blundell2007. If we choose a stronger sequence, however, then our test statistic $S_n^{2\textsl{e}}$ becomes considerably more powerful. \paragraph{Nonparametric Specification} Let us now study the finite sample of our test in the case of nonparametric specification. We generate the pair $(Z,W)$ as in the parametric case described above. For the generation of the dependent variable $Y$ we distinguish two cases. Besides the structural function $\sol_1(z)=\sum_{j=1}^\infty (-1)^{j+1}j^{-2} \sin(j\pi z)$ we also consider the function $\sol_2(z)=\sum_{j=1}^\infty ((-1)^{j+1}+1)/4\,j^{-2} \sin(j\pi z)$. Again, for computational reasons we truncate the infinite sum at $K=100$. The resulting functions are displayed in Figure (ref). Further, $Y$ is generated by (ref) either with $\sol_1$ and $c_U=0.2$ or $\sol_2$ and $c_U=0.8$. In both cases, we estimate the structural relationship using Lagrange polynomials with $\k=4$ for $n=500$ and $n=1000$.
If $H_{\textsl{np}}$ is false then $\Ex[U|W]\neq 0$ and we let $\Ex[U|W]=\Ex[\rho(Z)|W]$ where $\rho$ is defined below. Consequently, when $H_{\textsl{np}}$ is false we generate realizations of $Y$ from
for $l=1,2$ and $j\geq 1$ where $\rho_j(z)=c_j(\exp(2jz)\1_{\{z\leq 1/2\}}+\exp(2j(1-z))\1_{\{z>1/2\}}-1)$ and $c_j$ is a normalizing constant such that $\int_0^1 \rho_j(z)dz=0.5$. The functions $\rho_j$ are continuous but not differentiable at $0.5$. Roughly speaking, the degree of roughness of $\rho_j$ is larger for larger $j$.
In Table 3, we depict the empirical rejection probabilities when using $S_n^\textsl{np}$ with either no smoothing or additional smoothing ${\tau_j}=j^{-2}$, $j\geq 1$, which we denote by $S_n^{0\textsl{np}}$ or $S_n^{2\textsl{np}}$, respectively. When no additional smoothing is applied then the number of basis functions $f_j$ is given by $\m=11$ if $n=500$ and $\m=15$ if $n=1000$ and hence, the choice of $\m$ is slightly larger than $n^{1/3}$ as suggested by the theoretical results. The critical values of these statistics are estimated as described in Remark (ref) where in the case of $S_n^{2\textsl{np}}$ we choose $M=100$. We compare our results with the test of Horowitz2009. We observe that the statistic $S_n^{0\textsl{np}}$ is less powerful than $S_n^{2\textsl{np}}$ against the alternatives $\rho_1$ and $\rho_2$.
In the following, we illustrate that using additional weighting can be inappropriate. Table 4 illustrates the power of our tests when the structural function $\sol_2$ is considered and realizations $(Z,W)$ were generated by $ W \sim \cU[0,1]$, $Z = (0.8\, W +0.3\, \varepsilon)^2$ where $\varepsilon\sim\cN(0.5, 0.05)$. In this case, we generate $Y$ using (ref) where $c_U=0.8$. In this case, the estimates of the generalized coefficients of $T(\sol-\sol_0)$ are more fluctuating and using weights is not appropriate here. Indeed, as we can see from Table 4, the test statistic $S_n^{0\textsl{np}}$ with no smoothing is more powerful than $S_n^{2\textsl{np}}$ were weighting $\tau_j=j^{-2}$, $j\geq 1$, is used. In particular, $S_n^{0\textsl{np}}$ is much more powerful than the test of Horowitz2009.
Based on the methodology of series estimation, we have developed in this paper a family of goodness-of-fit statistics and derived their asymptotic properties. The implementation of these statistics is straightforward. We have seen that the asymptotic results depend crucially on the choice of the smoothing operator $L$. By choosing a stronger decaying sequence $\tau$, our test becomes more powerful with respect to local alternatives but might lose desirable consistency properties. We gave heuristic arguments how to choose the weights in practice. In addition, in a Monte Carlo investigation our tests perform well in finite samples.