Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,383 characters · 7 sections · 76 citation commands
A Projection Approach to Nonparametric Significance and Conditional Independence Testing
\doublespacing
Testing the significance of a subset of explanatory variables in a regression model is a fundamental problem in statistics and econometrics and is essential for variable selection and model specification. Given the risks of inconsistent estimation and misleading inference associated with misspecified parametric models, the development of tests within a nonparametric framework has attracted considerable attention. Theories have been established for cross-sectional studies based on independent and identically distributed (i.i.d.) data, with early contributions by lewbel1995consistent, fan1996consistent, lavergne2000nonparametric, and delgado2001significance, and more recent work such as zhu2018dimension and lundborg2024projected. Further developments address significance tests for more complex data structures, such as time series settings, categorical regressors, high-dimensional models, spatial point patterns, and network data, with related work including chen1999consistent, li1999consistent, racine2006testing, zhu2018significance, kojaku2018generalised, and dvovrak2024nonparametric. In addition, modern machine learning methods have motivated new significance testing problems; see horel2020significance, wu2021generalized, and giesecke2025aico.
Generally, the literature on nonparametric significance testing can be categorized into two mainstreams: local smoothing-based and global nonsmoothing-based methods. The local smoothing-based approach, exemplified by fan1996consistent and lavergne2000nonparametric, typically constructs test statistics using kernel-based squared distance measures. While these tests are consistent against omnibus alternatives, they often suffer from the \textquotedblleft curse of dimensionality\textquotedblright, where the local power degrades rapidly as the dimension of the covariates increases. Although the test of zhu2018dimension addresses the problem via a dimension-reduction method, nonparametric estimation remains necessary, yielding a nonparametric convergence rate and making it difficult to fully resolve the issue. An alternative approach is the global nonsmoothing-based approach adopted by delgado2001significance, which is based on the integrated conditional moment (ICM) principle of bierens1982consistent and bierens1990consistent, originally developed for specification tests of parametric regressions. The ICM idea in the context of nonparametric significance testing is to transform the nonparametric-type conditional moment restrictions under the null into an infinite number of unconditional moment restrictions. The tests proposed by delgado2001significance achieve the fastest possible parametric rate and are less sensitive to bandwidth choices. However, these tests still face challenges in effectively leveraging the asymptotic null behavior of the $U$-process when nuisance functions of the null model are estimated.
More specifically, the nonparametric estimation of the conditional mean function inevitably entails a random denominator problem, a difficulty that has been emphasized in lundborg2024projected as an essential barrier to the broader use of more flexible testing procedures. A standard way to control this problem is to construct density–weighted residuals so that the estimated unconditional moment restrictions can be written as a nondegenerate $U$–process, at the cost of introducing an additional \textquotedblleft nonparametric estimation effect\textquotedblright. This additional effect creates a substantial technical hurdle when attempting to obtain critical values via computationally efficient bootstrap schemes. To address this difficulty, previous work, for example, delgado2001significance, modifies the multiplier bootstrap version of the $U$-process used in their procedure and adopts more elaborate setups. These approaches typically rely on technically demanding trimming schemes or impose stronger assumptions on the density of the covariates under the null, for instance, by requiring it to be bounded away from zero. This type of restriction excludes many commonly used distributions, such as the normal and the Student's $t$.
The primary contributions of this paper are to address the issues discussed above by introducing a nonparametric projection and to derive multiplier bootstrap critical values in a computationally efficient and interpretable manner. To be precise, we project the weighting function used in constructing the integrated conditional moments onto the orthocomplement of the space spanned by the density evaluated at all covariates under consideration, including both those already known to be significant and those whose significance is being tested. It is worth noting that, although, to our knowledge, the \textquotedblleft nonparametric estimation effect\textquotedblright has not been emphasized in the existing literature, an analogous phenomenon has been recognized earlier in testing parametric models; see durbin1973distribution. Moreover, the novel projection approach that we develop to address this effect is motivated by the parametric projection approaches proposed in sant2019specification, yang2024model, and song2025unified. Thanks to the projection, the critical values for our statistics can be simulated via the computationally fast multiplier bootstrap procedure rather than the computationally intensive wild bootstrap, as in much of the nonparametric significance testing literature, for example, gu2007bootstrap. Finally, given that testing the conditional independence assumption is closely related to testing significance in the mean, we also extend the tailor-made projection procedure to test this important assumption.
The rest of the paper is organized as follows. Section (ref) outlines the testing procedure, including the nonparametric projection-based methodology and the construction of our statistics. The asymptotic properties with some reasonable assumptions of the test under the null, local alternatives, and global alternatives are established in Section (ref). We further detail the implementation of the multiplier bootstrap procedure in Section (ref) and extend the above projection test and the multiplier bootstrap procedure to testing the conditional independence assumption in Section (ref). The finite-sample performance of the proposed test is investigated via a set of Monte Carlo simulations in Section (ref). Finally, Section (ref) concludes the paper. Additional simulation results and detailed proofs of the theoretical results are provided in the online supplement.
Let $(S,\mathcal{F},\mathbb{P})$ be the probability space of the random vector $\chi = (Y,W^\top)^\top$, where $Y$ is a scalar response variable and $W=(X^\top,Z^\top)^\top$ is the vector of explanatory variables, with $X$ being $\mathbb{R}^q$-valued and $Z$ being $\mathbb{R}^p$-valued. Henceforth, $A^\top$ denotes the matrix transpose of $A$. We consider testing whether the vector of covariates $Z$ has no additional explanatory power for the conditional mean of the dependent variable $Y$, given that $X$ has a statistically significant effect. The null hypothesis of interest is
and the alternative hypothesis, $H_1$, is the negation of $H_0$. Let $f_X(\cdot)$ denote the probability density function (PDF) of $X$. Using the fact that the PDF $f_X(X)>0$ $a.s.$, we can rewrite $H_0$ in (ref) as
where $\epsilon=Y-\mathbb{E}[Y\vert X]:=Y-m(X)$ is the nonparametric error, with $m(X)$ denoting the unknown conditional mean function under the null.
It is well established in the testing literature that the conditional moment restriction in (ref) is equivalent to a continuum of unconditional moment restrictions. More specifically, following the ICM principle introduced by bierens1982consistent, and further developed by stute1997nonparametric, stute1998bootstrap, delgado2001significance, stute2002model, and escanciano2006consistent, we employ the indicator function $1_w(W):=1(W\leq w)=1(X\leq x)1(Z\leq z)$ with $w=(x^\top,z^\top)^\top\in\mathbb R^{q+p}$ as the weighting scheme to transform the conditional moment restriction to the unconditional version. Consequently, the null hypothesis $H_0$ in (ref) can be equivalently characterized by the following infinite number of unconditional moment conditions indexed by $w$:
which converts the original nonparametric significance testing problem into a global testing problem. Therefore, testing (ref) achieves the fastest possible parametric rate.
Suppose that we have a random sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$ of size $n\geq 1$ consisting of independent and identically distributed (i.i.d.) random variables, the natural sample analog for (ref) is given by
where $\hat{\epsilon}_i=Y_i-\hat{m}(X_i)$ is the nonparametric residual, with the leave-one-out nonparametric estimators for $m(X_i)$ and $f_X(X_i)$ given by
and
respectively. Here, $a=a(n)\in\mathbb{R}^{+}$ is a bandwidth parameter shrinking to zero at a suitable rate as $n\to\infty$ and $K(u)=\Pi_{j=1}^qk(u^{(j)})$ is a product kernel for a $q$-dimensional vector $u$, where $k(\cdot)$ is the univariate kernel and $u^{(j)}$ denotes the $j$-th component of $u$. Note that the random denominator problem is effectively eliminated by incorporating the estimated density $\hat{f}_X(X_i)$ into the $U$-process $\hat T_n(\cdot)$, significantly facilitating the theoretical analysis of $\hat T_n(\cdot)$.
By imposing the standard regularity conditions considered in delgado2001significance (see their Assumptions A1--A5), under the null $H_0$, we have
where $F_{Z\vert X}(z\vert x)$ is the conditional distribution function of $Z$ given $X$. It is observed that in the asymptotically uniform expansion (ref), the term
is the infeasible process for testing the null hypothesis in (ref) if $\epsilon_i$ is known (equivalently if $m(X_i)$ is known), and the term
can be regarded as the “nonparametric estimation effect” due to using the nonparametric residual $\hat{\epsilon}_i$ to replace $\epsilon_i$ (via the estimation of $m(X_i)$).
Much less attention has been paid to studying the above term, unlike the “parametric estimation effect” also known as the “Durbin problem” mentioned in durbin1973distribution. Specifically, since the asymptotic null distributions of the test statistics (constructed as continuous functionals of $\hat T_n(\cdot)$) are typically case-dependent, the use of bootstrap methods to obtain critical values is unavoidable. However, the “nonparametric estimation effect” in the asymptotically uniform expansion of $\hat T_n(\cdot)$ poses a substantial challenge to the direct application of the computationally efficient multiplier bootstrap. While the wild bootstrap is a potential alternative, it incurs significantly higher computational complexity. To address this, with the help of (ref), delgado2001significance propose a specifically constructed multiplier bootstrap-based $U$-process as follows:
with $\{V_i\}_{i=1}^n$ being i.i.d. random variables (i.e., multipliers) that have mean zero, unit variance, and are independent of the original sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$. However, this approach necessitates estimating the conditional distribution function $F_{Z\vert X}(z\vert X_i)$ nonparametrically within the multiplier bootstrap procedure, which is
See Section $3$ of delgado2001significance for a detailed description of this implementation. Unfortunately, this inevitably reintroduces the random denominator issue in the multiplier bootstrap, thereby undermining the purpose of using $\hat T_n(w)$. Indeed, the solution adopted by delgado2001significance to ensure the validity of this bootstrap is to impose stronger assumptions, such as requiring the density $f_X(X)$ to be bounded away from zero (see their Assumptions A9), which was exactly intended to be avoided in the initial construction of $\hat T_n(w)$ through the density weighting $\hat f_X(X_i)$. Such a restrictive bounded support assumption constitutes a significant barrier to applying the test to data with unbounded support, including the commonly used normal distribution. Another potential solution is to use trimming; however, the prohibitive complexity of the required theoretical proofs and the practical choice of the trimming level have prevented its successful implementation.
In this paper, to eliminate the nonparametric estimation effect, we propose a tailor-made nonparametric significance test that depends on a novel nonparametric-type projection of the weighting function $1_z(Z_i)$ onto the orthocomplement of the space spanned by $f_W(W_i)$ (hereafter, the PDF of $W$), where orthogonality is understood conditionally on $X_i$. To this end, we first consider the following infeasible projected weighting function:
where $f_{Z\vert X}(z\vert x)$ is the conditional density function of $Z$ given $X$,
In fact, the projection constructed above admits an intuitive interpretation as the linear regression of $1_z(Z_i)$ on $f_W(W_i)$ conditional on $X_i$, where $\Delta^{-1}(X_i)G(z;X_i)$ represents the coefficients of the linear projection. In other words, the term $f_W(W_i)\Delta^{-1}(X_i)G(z;X_i)$ constitutes the best linear predictor of $1_z(Z_i)$ given $f_W(W_i)$. Consequently, our constructed weighting function $\mathcal{P}1_z(Z_i)$ corresponds precisely to the associated prediction error, which implies that $\mathcal{P}1_z(Z_i)$ is orthogonal to $f_W(W_i)$ conditional on $X_i$, i.e.,
Building on the sample version of the nonparametric projection $\mathcal{P}1_z(Z_i)$, our proposed projected $U$-process is constructed as follows:
where
is the natural estimator for $\mathcal{P}1_z(Z_i)$. Here,
are the sample versions of $\Delta(X_i)$ and $G(z;X_i)$, respectively, with
being the leave-one-out kernel density estimator for the PDF $f_W(X_i,z)$, where $L(\cdot)$ and $b=b(n)\in\mathbb R^+$ are the kernel and bandwidth, respectively.
Note that the proposed process $\hat R_n(w)$ based on the nonparametric-type orthogonal projection $\hat{\mathcal{P}}_n1_z(Z_i)$ can account for effectively the “nonparametric estimation effect” discussed before, although at the cost of nonparametrically estimating the density $f_W$. As such, it can be implemented directly using a convenient multiplier bootstrap without reintroducing the random denominator, in sharp contrast to the multiplier bootstrap used by delgado2001significance. Further details are provided in Section (ref). The test statistics are then constructed as continuous functionals of $\hat R_n(w)$. In this paper, we focus on the two most commonly employed choices, namely, the Cram\'{e}r--von Mises (CvM) and the Kolmogorov--Smirnov (KS) test statistics, which are explicitly given as
From a theoretical perspective, the integrating measure $F_{W_n}(\cdot)$ appearing in the $CvM_n$ is assumed to be a random measure that converges in probability to $F_W(\cdot)$, which is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{q+p}$. In practice, we adopt the approach suggested in escanciano2006consistent: for the computation of $CvM_n$ we take $F_{W_n}(\cdot)$ to be the empirical distribution function of $\{W_i\}_{i=1}^n$, and for the computation of $KS_n$ we replace the theoretical supremum by its maximum taken over the sample points. Under the null hypothesis $H_0$, the test statistics $CvM_n$ and $KS_n$ are expected to take sufficiently small values, whereas relatively large realizations provide evidence against $H_0$ in favor of the alternative hypothesis $H_1$. The critical values used to assess the magnitude of the test statistics are obtained via the multiplier bootstrap procedure detailed in Section (ref).
In this section, the large-sample properties of the proposed $CvM_n$ and $KS_n$ test statistics will be considered, with particular attention given to their asymptotic null distributions, local power, and consistency. Before presenting the theoretical results, it is necessary to adopt several definitions as introduced in delgado2001significance and to state a set of regularity conditions that hold uniformly across all cases. Let $\mathscr{K}_l$ with $l\geq1$ be the class of even functions of uniformly bounded variation $k:\mathbb{R}\to\mathbb{R}$, which satisfy
where $\delta_{ij}$ is the Kroneker's delta. Let $\mathscr{L}_\beta^\alpha$ with $\alpha>0$ and $\beta>0$ be the class of functions $g:\mathbb{R}^q\to\mathbb{R}$, which satisfy uniformly $(b-1)$-times continuously differentiability for $b-1\leq \beta\leq b$. In addition, we impose the assumption that there exist a positive constant $\rho$ and a function $d$ with finite $\alpha$-th moments such that
where $Q$ is a $(b-1)$-th degree homogeneous polynomial in $v-u$ with coefficients given by the partial derivatives of $g$ at $u$ of orders up to $b-1$, all of which are assumed to have finite $\alpha$-th moments; in particular, $Q=0$ when $b=1$.
Assumption (ref) concerns the moments of $\epsilon=Y-m(X)$ and $\Delta^{-1}(X)f_X^2(X)$. While fan1996consistent required the existence of the fourth moment, this condition was relaxed in delgado2001significance. We adopt the relaxed version to guarantee the finiteness of the variances of the statistics, which is crucial for establishing the theoretical results. Moreover, the assumption concerning $\Delta^{-1}(X)f_X^2(X)$ is essential to guarantee the convergence of the proposed projection structure. In fact, it can be reformulated as the moment condition $\mathbb{E}\vert\int f_{Z\vert X}^2(\bar z|X)\,d\bar z\vert^{-2-(4+2\delta_2)/\delta_1}<\infty$, which restricts the tail properties of the conditional density function $f_{Z|X}$ and the density function $f_X$ and is satisfied by most commonly used distributions. This assumption permits an unbounded covariate $X$ and is weaker than the stringent bounded–support assumption on $X$ as imposed in delgado2001significance (see their Assumptions A9). This theoretical refinement extends the applicability of the ICM–type tests for nonparametric moment restrictions to data generated from normal and other commonly used unbounded distributions that could not be handled under the conditions in delgado2001significance, as confirmed by the simulation results for the unbounded covariate case in Section (ref).
The function classes specified in Assumption (ref) impose the Lipschitz continuity of the functions $f_X(\cdot)$, $f_{Z\vert X}(z|\cdot)$, and $m(\cdot)$, which in turn ensures that the bias of the constructed empirical process is asymptotically negligible and the variance is finite.\footnote{We note that the condition $\sup_{(x,z)}\vert f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$ is sufficient but not necessary and permits an unbounded covariate $X$. This condition is assumed to control uniformly the bias term of $\hat{\Delta}_n(x)$, as shown in Lemma (ref). In fact, alternative conditions may also deliver the desired result. Let $d_{f_X}(\cdot)$ denote the Lipschitz coefficient of $f_X(\cdot)$ and $d_{f_{Z\vert X}}(z\vert \cdot)$ denote the Lipschitz coefficient of $f_{Z\vert X}(z\vert \cdot)$, in the sense of the definitions of $\mathscr{L}_{\lambda_1}^\infty$ and $\mathscr{L}_{\lambda_2}^\infty$, respectively. For example, if the conditions $\sup_{(x,z)}\vert d_{f_X}(x)f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$, $\sup_{(x,z)}\vert d_{f_{Z\vert X}}(z\vert x)\vert<\infty$, and $\sup_{(x,z)}\vert d_{f_X}(x)d_{f_{Z\vert X}}(z\vert x)/f_X(x)\vert<\infty$ are imposed in place of $\sup_{(x,z)}\vert f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$, the same conclusion continues to hold. } Note that standard densities, such as the normal distribution, and widely used linear models satisfy this assumption. In particular, the function class is restricted to be bounded when $\alpha=\infty$, for example, in the case of density functions. Moreover, the requirements on the kernel function $k(\cdot)$ strengthen the conventional high-order kernel assumptions by imposing a decay rate condition, as mentioned in robinson1988root. In addition, the order of the kernel function depends on the moments of the functions $f_X(\cdot)$, $f_{Z\vert X}(z|\cdot)$, and $m(\cdot)$.
Assumption (ref) determines the lower and upper bounds for the bandwidth sequences $a$ and $b$ and thus guarantees the convergence of the proposed $U$-process. On the one hand, the lower bounds on the bandwidths, which are consistent with those in robinson1988root and delgado2001significance, serve to ensure the asymptotic negligibility of the bias of the $U$-process. On the other hand, the upper bounds on the bandwidths in Assumption (ref) provide a sufficient, though not necessarily weakest possible, condition to establish the uniform convergence of the nonparametric estimator $\hat{\Delta}_n(\cdot)$ for $\Delta(\cdot)$ as shown in Lemma (ref). The slightly strengthened assumptions on the upper bounds are used to simplify the theoretical analysis. Specifically, we expand $\hat{\Delta}_n(\cdot)$ only up to the second order, although the upper bounds could be relaxed with higher order expansions of $\hat{\Delta}_n(\cdot)$. The upper bounds cannot, however, be relaxed beyond the rate $(na^{2q}b^{p})^{-1}\to 0$, which guarantees that the variance of the test statistics remains bounded.
In the following, let “$\Longrightarrow$” represent weak convergence on $(l^{\infty}(\Pi),\mathcal{B}_\infty)$ in the sense of Hoffmann–-J\orgensen, where $\mathcal{B}_\infty$ denotes the corresponding Borel $\sigma$-algebra, see, e.g., Definition $1.3.3$ in van1996weak. We establish formally the uniform decomposition of $\hat R_n(\cdot)$, which indicates that the nonparametric estimation effects arising from $\hat m$ and $\hat f_W$ are asymptotically negligible. Denote
To complement the theoretical results under the null hypothesis, and building upon Theorem (ref), the asymptotic distributions of the $CvM_n$ and $KS_n$ statistics introduced in Section (ref) are established by applying the continuous mapping theorem, as discussed in van1996weak.
Theorem (ref) establishes the theoretical advantage of the proposed test, namely, that despite the applications of kernel smoothing methods, the resulting empirical process still achieves the parametric convergence rate, with a centered Gaussian process as its limiting process. Subsequently, the Corollary (ref) ensures that the statistics $nCvM_{n}$ and $\sqrt{n}KS_{n}$ converge to the squared norm and the sup norm of the aforementioned Gaussian process, respectively. The uniform decomposition given in (ref) is both theoretically and empirically attractive for implementing the multiplier bootstrap detailed in Section (ref). To investigate the power of the proposed test, we introduce a sequence of local alternatives converging to the null hypothesis at the parametric rate,
under which the asymptotic theory of the constructed empirical process is subsequently established. We note that $\Psi:\mathbb{R}^{p+q}\to\mathbb{R}$ is a bounded nonzero function.
Owing to the similarity between the theoretical derivations under the sequence of local alternatives and those established under the null hypothesis as in Corollary (ref), the results $nCvM_{n}\stackrel{d}\longrightarrow \int\vert R_{\infty}(w)+\mu(w)\vert^2F_W(dw)$ and $\sqrt{n}KS_{n}\stackrel{d}\longrightarrow \sup\limits_{w}\vert R_{\infty}(w)+\mu(w)\vert$ are presented directly, where we omit the verification details of the convergence of the statistics. Consequently, the local powers of the proposed statistics are characterized by $\mu(\cdot)$. Specifically, the deterministic shift term determines the non-centrality of the limiting process in comparison with the centered Gaussian limiting process under the null hypothesis, and it leads to the fact that whenever the deterministic shift function is nonzero for at least some $w\in\mathbb R^{p+q}$ with a positive Lebesgue measure, the proposed $CvM_n$ and $KS_n$ statistics possess nontrivial power against the local alternatives. Indeed, since $\mu(\cdot)$ depends on the conditional $L^2$ inner product of $\Psi(X,Z)$ and $\mathcal{P}1_z(Z)$ given $X$, in the sense that $\mu(w) = \mathbb{E}\{f_X(X)1_x(X)\mathbb{E}[\Psi(W)\mathcal{P}1_z(Z)\vert X]\}$, the nontrivial local power is well-established as long as for any constant $\gamma$,
Finally, to derive the asymptotic global power properties of the proposed test statistics, we investigate the asymptotic behavior of the constructed empirical process under the alternative hypothesis.
The correspondence between local and global power is straightforward. The local power is driven by $\mu(\cdot)$, and the global power, in turn, depends on the conditional $L^2$ inner product of $\mathbb{E}(Y\vert W)-\mathbb{E}(Y\vert X)$ and $\mathcal{P}1_z(Z)$ given $X$. More specifically, once
is satisfied for any constant $\gamma$, the unconditional expectation $C(w)\neq 0$. This in turn implies that, as $n$ goes to infinity, the constructed empirical process converges uniformly in probability to a non-vanishing function $C(\cdot)$, which further leads the test statistics to diverge to positive infinity in probability. However, we do not regard the potential loss of power along certain directions of the alternative space as a primary empirical concern. As shown in bierens1997asymptotic, ICM–type tests are admissible under suitable regularity conditions, so that no uniformly more powerful procedure exists within a broad class of alternatives. These trade–offs are detailed in the simulation studies in Section (ref).
The case-dependent asymptotic limiting process established in Section (ref) necessitates the consideration of bootstrap procedures. In the traditional literature, two distinct bootstrap schemes are frequently considered. The residual-based bootstrap employed in koul1994bootstrapping generates bootstrap samples by resampling from the empirical distribution of the estimated residuals (possibly after a smoothing step), and has been widely employed in classical nonparametric testing procedures. In contrast, the wild bootstrap, as mentioned in mammen1993bootstrap, constructs bootstrap errors by multiplying the estimated residuals with independent random multipliers of zero mean and unit variance, thereby creating bootstrap samples without explicitly resampling from the residual distribution.
Although traditional bootstrap methods may be justified after careful theoretical verification, we consider more computationally efficient alternatives. Motivated by van1996weak, the uniform decomposition in (ref) inspires an intuitive, easy-to-implement multiplier bootstrap procedure. Towards this end, define the multiplier bootstrapped projected process as
where $\{V_i\}_{i=1}^n$ is a sequence of random variables (i.e., multipliers) that are mean zero, unit variance, and independent of the original sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$. It is worth noting that in contrast to the multiplier bootstrap proposed in delgado2001significance, which is motivated by the uniform decomposition in (ref), our multiplier bootstrap version $\hat R_n^\ast(w)$ does not involve a random denominator thanks to the novel projection, and thus it is not necessary to impose the stringent compact support assumption $\mathbb{P}(f_X(X)>\nu)=1$ for some constant $\nu>0$. This fact allows us to accommodate important distributions for X with unbounded support, like the t and normal.
It is worth noting that, under the null hypothesis, the bootstrap version coincides with the original sample-based empirical process in both convergence rate and asymptotic distribution. By constructing the multiplier bootstrap versions of the $CvM_n$ and $KS_n$ statistics in a manner analogous to Section (ref),
and applying the continuous mapping theorem in a way analogous to that employed in Corollary (ref), we establish that under the null hypothesis the bootstrap $CvM_n^\ast$ and $KS_n^\ast$ statistics consistently approximate the asymptotic null distributions of $CvM_n$ and $KS_n$, respectively, thereby ensuring that the empirical sizes of the two tests are theoretically close to the nominal level. Under local alternatives $H_{1n}$, however, the local power depends on $\mu(\cdot)$ introduced in Section (ref). This is because the bootstrap versions converge to the same limiting processes as under the null, whereas the original statistics converge to limiting processes with additional drift components. Therefore, taking the $CvM_n$ statistic as an example, the different limiting distributions of $CvM_n$ and its bootstrap counterpart $CvM_n^\ast$ imply that the proportion of rejecting $H_0$ will exceed the nominal level, thereby confirming the nontrivial local power of the proposed test under $H_{1n}$. As a consequence of the above analysis, the asymptotic critical value (still taking the $CvM_n$ statistic as an example) at the significance level $\alpha$ is $c^\ast_{\alpha}=\inf\{c_\alpha\in[0,\infty):\lim_{n\to\infty}\mathbb{P}_n^\ast(nCvM_n^\ast>c_\alpha)=\alpha\}$, where $\mathbb{P}_n^\ast$ is the bootstrap probability under the bootstrap law. In practice, $c^\ast_{\alpha}$ can be approximated as $c^\ast_{n,\alpha} = \{nCvM_n^\ast\}_{B(1-\alpha)}$, the $B(1-\alpha)$-th order statistic for $B$ replicates $\{nCvM_{n,b}^\ast\}_{b=1}^B$ and we reject $H_0$ if $nCvM_n>c^\ast_{n,\alpha}$. Finally, under the alternative hypothesis $H_1$, when the condition (ref) is satisfied, the consistency of our bootstrap statistics against $H_1$ follows from the fact that the original statistics diverge to infinity in probability, whereas the bootstrap versions remain stochastically bounded.
The projection techniques developed in this article to handle the nonparametric estimation effect can readily be extended to testing other restrictions on regression curves, thereby allowing ICM-type tests, which are particularly advantageous in certain settings, to be applied to these important problems. A particularly important example is the problem of testing conditional independence, which has attracted considerable attention in recent years and has been emphasized as a key issue in, among others, su2014testing, huang2016flexible, wang2018characteristic, li2020nonparametric, shah2020hardness, neykov2021minimax and cai2022distribution. Formally, we are interested in testing whether $Y$ is independent of $Z$ given $X$, denoted by $H_0^{CI}:Y\perp Z\vert X$. Using the previous notations, we consider testing
which is further equivalent to
where $\epsilon(y) = 1_y(Y)-\mathbb{E}[1_y(Y)\vert X] = 1_y(Y)-F_{Y\vert X}(y\vert X)$, with $F_{Y\vert X}(y\vert x)$ the conditional distribution function of $Y$ given $X$. Similar to $T_n(w)$ for the case of significance testing in the conditional mean, delgado2001significance have proposed to use the following unprojected $U$-process to test $H_0^{CI}$:
which is shown to satisfy
Motivated by the uniform decomposition in (ref) and the expected case-dependent asymptotic limiting process, we would need to obtain critical values for the tests based on $\hat L_n(\cdot)$ via a multiplier bootstrap scheme similar to that described in Section (ref). This, however, requires the nonparametric estimation of $F_{Z\vert X}(z|X_i)$. The random denominator problem leads either to overly stringent assumptions on the distribution of $X$ or to substantial additional technical difficulties. Similar to the discussion in Section (ref), when ICM-type tests are used to test conditional independence, \textquotedblleft nonparametric estimation effect\textquotedblright arises. In particular, these difficulties stem from the presence of the term
Thanks to the projection-based principle proposed in previous sections, the following projected empirical process can be defined to solve the problems discussed above:
where $\hat{\epsilon}_i(y)=1_y(Y_i)-\hat{F}_{Y\vert X}(y\vert X_i)$ is the estimator for $\epsilon_i(y)$, with the leave-one-out nonparametric estimator
All other quantities are defined as before. The analysis of $\hat I_n(y,w)$ is identical to $\hat R_n(w)$ but with $\hat{\epsilon}_i$ substituted by $\hat{\epsilon}_i(y)$. Thus, reasoning as in the Section (ref), we have
This leads to the fact that our constructed projected $U$-process $\sqrt n\hat I_n(\cdot)$ weakly converges to a centered Gaussian process $R_\infty^{CI}(\cdot)$ with covariance structure
at the parametric rate under the null. Correspondingly, under the alternative hypothesis, the constructed empirical process converges uniformly in probability to
As a result, the test is consistent as long as for any constant $\gamma$, the following requirement
holds. In addition, this condition also characterizes the nontrivial asymptotic local power of the test for conditional independence. The widely used CvM and KS test statistics can be constructed as
so that large values provide evidence against the null in favor of the alternative hypothesis, where $F_{Y_n,W_n}(\cdot)$ is a random measure that converges in probability to $F_{Y,W}(\cdot)$ that is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{1+q+p}$. To implement the proposed test for conditional independence, still motivated by (ref), we construct the multiplier bootstrap version of $\hat I_n(y,w) $ as follows:
and use it to obtain the critical values. The desired asymptotic properties follow from the following result:
The corresponding multiplier bootstrap $CvM^{CI,\ast}_n$ and $KS^{CI,\ast}_n$ test statistics can be defined in a way analogous to that described in Section (ref), and can be shown, under the null and under local alternatives, to provide good approximations to the sup norm and the squared norm of $R_\infty^{CI}(\cdot)$, respectively. Under fixed alternatives, the bootstrap versions of these statistics remain stochastically bounded, whereas the original statistics diverge to infinity whenever condition (ref) is satisfied, thereby establishing the validity of the bootstrap procedure. The implementation of the resulting test can therefore follow the same steps as those outlined in Section (ref). More Simulation results are presented in Section (ref).
In this section, a Monte Carlo study is conducted to evaluate the finite-sample performance of the proposed nonparametric projection method. The main objectives are as follows. First, the proposed CvM test based on projection (denoted by PJ) is compared with the non-projected CvM test of delgado2001significance (denoted by DM) in terms of empirical size and power.\footnote{For brevity, we have focused on comparing the CvM-type statistics. The comparison using the KS-type statistics is similar and available upon request.} Second, the robustness of the test performance with respect to bandwidth choice is examined under various bandwidth settings. Third, the behavior of the projection-based test is investigated across different covariate distributions and alternative frequencies.
In the simulation design presented in the main text, the covariates $X$ and $Z$ are both one-dimensional and follow a specified dependence structure $X=Z+U$, where $Z$ and $U$ are independent $N(0,1)$ variables. It is worth emphasizing that our procedures are not restricted to this case. The tests are applicable when $X$ and $Z$ are multivariate, and to illustrate this property, we also consider designs with $p=2$ and $q=2$. The proposed tests also apply when $X$ and $Z$ follow the uniform distribution $U(0,1)$. Moreover, when some parts of Assumptions (ref)--(ref) are violated, for instance when both covariates are $N(0,1)$, the size accuracy of the test is affected but remains within an acceptable range and compares favorably with the DM test, which suggests that our procedure has a broader practical scope. For the sake of brevity, and because the main conclusions are similar, these additional results are reported in Appendix (ref), whereas the main text focuses on the case $X=Z+U$ with $Z$ and $U$ being independent $N(0,1)$. To compare the performance of the tests under alternatives with different frequencies, the response variable $Y$ is generated under the following data-generating process (DGP),
where $\epsilon\sim N(0,1)$ and is independent of $(X,Z)$. Under the null hypothesis, the parameter $\gamma$ is fixed at zero, whereas under the alternative hypothesis, the frequency is determined by the value of $\gamma$. Moreover, to assess, in the spirit of bierens1997asymptotic, the admissibility properties of ICM-type tests, we also report in Appendix (ref) simulation results for designs with interaction effects between covariates. In particular, we consider alternatives of the form $X(Z^2-1)$ and examine the power of the proposed procedures in this case. In addition, the adoption of the kernel function and method of bandwidth selection follows those in delgado2001significance, where an $l$-th order Epanechnikov kernel is employed, and the bandwidth is determined by the rule-of-thumb with coefficient $c\in\{0.5,1,2\}$ to assess the sensitivity to bandwidth. The order of the kernel and the explicit bandwidth formula are consistent with the requirements of Assumption (ref). Specifically, $b=cn^{-1/3}$ is used for $l=2$, and $b=cn^{-1/6}$ is adopted for $l=4$. The results were examined under sample sizes $n\in\{200,400\}$ with the nominal significance level set at $\alpha\in\{0.01,0.05,0.1\}$, and all outcomes were obtained from $1000$ Monte Carlo replications and $199$ multiplier bootstrap samples.
Tables (ref)--(ref) report results for the case $p=q=1$. Table (ref) corresponds to the design $X=Z+U$, where $Z$ and $U$ are independent $N(0,1)$ variables, which satisfies Assumptions (ref)--(ref) but not the bounded-support assumption on the distribution of $X$ imposed in delgado2001significance. Table (ref) considers the case where $X$ and $Z$ are independent $N(0,1)$ variables, and Table (ref) the case where $X$ and $Z$ are independent $U(0,1)$ variables. As is typically observed, both size and power improve substantially as sample size increases, demonstrating that the proposed procedure possesses desirable large-sample properties.
For comparison with the method introduced in delgado2001significance (named DM), the performance of each method in each case is reported in the tables. It is observed that DM tests exhibit lower accuracy, whereas the projection-based method achieves significant improvements. Regarding the power properties, Table (ref) shows that the proposed projection-based tests (PJ) perform markedly better than the DM tests, which illustrates the suitability of our procedure for a wider class of unbounded distributions that satisfy Assumptions (ref)--(ref). In Table (ref), we investigate how the two tests behave when commonly used distributions that violate these assumptions are employed. We observe that the DM tests suffer from substantial size distortion, with empirical rejection frequencies well above the nominal level (oversized), whereas the PJ tests exhibit much better size accuracy. In Table (ref), both tests perform well, which is mainly due to the fact that the data are generated from the standard uniform distribution, a commonly used design that is fully compatible with the theoretical assumptions.
Further notable advantages of the proposed tests include reasonable robustness to bandwidth selection, as shown by results across different values of $c$, and relatively high power, as evidenced by comparisons across different values of $\gamma$. These features are characteristic of ICM-type tests in general and are naturally also reflected in the performance of the DM tests.
Finally, as further support for the theoretical results established in Section (ref), we use the above DGPs to test the conditional independence assumption and the results are reported in Tables (ref)--(ref). The only difference in the design is that
where the conditional independence under test pertains only to the variance structure. Note that the significance tests discussed in Sections (ref)--(ref) can also be interpreted as tests of conditional independence for the conditional mean function, and the conclusions drawn from Tables (ref)--(ref) are similar to those obtained above.
This paper introduces a novel nonparametric significance test that uses a tailored projected weighting function to address the nonparametric estimation effect. The approach relaxes the restrictive compact support assumption, provides a computationally efficient multiplier bootstrap, and extends straightforwardly to testing conditional independence. Numerical evidence confirms the strong performance of our proposal in finite samples.