The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
70,370 characters
A Projection Approach to Nonparametric Significance and Conditional Independence Testing
\title{A Projection Approach to Nonparametric Significance and Conditional Independence Testing\footnote{The authors contributed equally to this work and are listed in alphabetical order.}}
\author{Xiaojun Song\thanks{Corresponding author: Department of Business Statistics and Econometrics, Guanghua School of Management, Peking University. Email: \texttt{[email removed]}. This work was supported by the National Natural Science Foundation of China [Grant Numbers 72373007 and 72333001]. The author also gratefully acknowledges the research support from the Center for Statistical Science of Peking University, China, and the Key Laboratory of Mathematical Economics and Quantitative Finance (Peking University) of the Ministry of Education, China.}\\
\and Jichao Yuan\thanks{Department of Business Statistics and Econometrics, Guanghua School of Management, Peking University. Email: \texttt{[email removed]}.}
}
\maketitle
\begin{abstract}
This paper develops a novel nonparametric significance test based on a tailored nonparametric-type projected weighting function that exhibits appealing theoretical and numerical properties. We derive the asymptotic properties of the proposed test and show that it can detect local alternatives at the parametric rate. Using the nonparametric orthogonal projection, we construct a computationally convenient multiplier bootstrap to obtain critical values from the case-dependent asymptotic null distribution. Compared with the existing literature, our approach overcomes the need for a stronger compact support assumption on the density of covariates arising from random denominators. We also extend the tailor-made projection procedure to test the conditional independence assumption. The simulation experiments further illustrate the advantages of our proposed method in testing significance and conditional independence in finite samples.
\end{abstract}
\newpage
\doublespacing
\section{Introduction}
Testing the significance of a subset of explanatory variables in a regression model is a fundamental problem in statistics and econometrics and is essential for variable selection and model specification. Given the risks of inconsistent estimation and misleading inference associated with misspecified parametric models, the development of tests within a nonparametric framework has attracted considerable attention. Theories have been established for cross-sectional studies based on independent and identically distributed (i.i.d.) data, with early contributions by \cite{lewbel1995consistent}, \cite{fan1996consistent}, \cite{lavergne2000nonparametric}, and \cite{delgado2001significance}, and more recent work such as \cite{zhu2018dimension} and \cite{lundborg2024projected}. Further developments address significance tests for more complex data structures, such as time series settings, categorical regressors, high-dimensional models, spatial point patterns, and network data, with related work including \cite{chen1999consistent}, \cite{li1999consistent}, \cite{racine2006testing}, \cite{zhu2018significance}, \cite{kojaku2018generalised}, and \cite{dvovrak2024nonparametric}. In addition, modern machine learning methods have motivated new significance testing problems; see \cite{horel2020significance}, \cite{wu2021generalized}, and \cite{giesecke2025aico}.
Generally, the literature on nonparametric significance testing can be categorized into two mainstreams: local smoothing-based and global nonsmoothing-based methods. The local smoothing-based approach, exemplified by \cite{fan1996consistent} and \cite{lavergne2000nonparametric}, typically constructs test statistics using kernel-based squared distance measures. While these tests are consistent against omnibus alternatives, they often suffer from the \textquotedblleft curse of dimensionality\textquotedblright{}, where the local power degrades rapidly as the dimension of the covariates increases. Although the test of \cite{zhu2018dimension} addresses the problem via a dimension-reduction method, nonparametric estimation remains necessary, yielding a nonparametric convergence rate and making it difficult to fully resolve the issue.
An alternative approach is the global nonsmoothing-based approach adopted by \cite{delgado2001significance}, which is based on the integrated conditional moment (ICM) principle of \cite{bierens1982consistent} and \cite{bierens1990consistent}, originally developed for specification tests of parametric regressions. The ICM idea in the context of nonparametric significance testing is to transform the nonparametric-type conditional moment restrictions under the null into an infinite number of unconditional moment restrictions. The tests proposed by \cite{delgado2001significance} achieve the fastest possible parametric rate and are less sensitive to bandwidth choices. However, these tests still face challenges in effectively leveraging the asymptotic null behavior of the $U$-process when nuisance functions of the null model are estimated.
More specifically, the nonparametric estimation of the conditional mean function inevitably entails a random denominator problem, a difficulty that has been emphasized in \cite{lundborg2024projected} as an essential barrier to the broader use of more flexible testing procedures. A standard way to control this problem is to construct density–weighted residuals so that the estimated unconditional moment restrictions can be written as a nondegenerate $U$–process, at the cost of introducing an additional \textquotedblleft nonparametric estimation effect\textquotedblright{}. This additional effect creates a substantial technical hurdle when attempting to obtain critical values via computationally efficient bootstrap schemes. To address this difficulty, previous work, for example, \cite{delgado2001significance}, modifies the multiplier bootstrap version of the $U$-process used in their procedure and adopts more elaborate setups. These approaches typically rely on technically demanding trimming schemes or impose stronger assumptions on the density of the covariates under the null,
for instance, by requiring it to be bounded away from zero. This type of restriction excludes many commonly used distributions, such as the normal and the Student's $t$.
The primary contributions of this paper are to address the issues discussed above by introducing a nonparametric projection and to derive multiplier bootstrap critical values in a computationally efficient and interpretable manner. To be precise, we project the weighting function used in constructing the integrated conditional moments onto the orthocomplement of the space spanned by the density evaluated at all covariates under consideration, including both those already known to be significant and those whose significance is being tested. It is worth noting that, although, to our knowledge, the \textquotedblleft nonparametric estimation effect\textquotedblright{} has not been emphasized in the existing literature, an analogous phenomenon has been recognized earlier in testing parametric models; see \cite{durbin1973distribution}. Moreover, the novel projection approach that we develop to address this effect is motivated by the parametric projection approaches proposed in \cite{sant2019specification}, \cite{yang2024model}, and \cite{song2025unified}. Thanks to the projection, the critical values for our statistics can be simulated via the computationally fast multiplier bootstrap procedure rather than the computationally intensive wild bootstrap, as in much of the nonparametric significance testing literature, for example, \cite{gu2007bootstrap}.
Finally, given that testing the conditional independence assumption is closely related to testing significance in the mean, we also extend the tailor-made projection procedure to test this important assumption.
The rest of the paper is organized as follows. Section \ref{sec.Test} outlines the testing procedure, including the nonparametric projection-based methodology and the construction of our statistics. The asymptotic properties with some reasonable assumptions of the test under the null, local alternatives, and global alternatives are established in Section \ref{sec.Asy}. We further detail the implementation of the multiplier bootstrap procedure in Section \ref{sec.boot} and extend the above projection test and the multiplier bootstrap procedure to testing the conditional independence assumption in Section \ref{sec.CI}. The finite-sample performance of the proposed test is investigated via a set of Monte Carlo simulations in Section \ref{sec.Simulation}. Finally, Section \ref{sec.Conclusion} concludes the paper. Additional simulation results and detailed proofs of the theoretical results are provided in the online supplement.
\section{Testing procedure}\label{sec.Test}
Let $(S,\mathcal{F},\mathbb{P})$ be the probability space of the random vector $\chi = (Y,W^\top)^\top$, where $Y$ is a scalar response variable and $W=(X^\top,Z^\top)^\top$ is the vector of explanatory variables, with $X$ being $\mathbb{R}^q$-valued and $Z$ being $\mathbb{R}^p$-valued. Henceforth, $A^\top$ denotes the matrix transpose of $A$. We consider testing whether the vector of covariates $Z$ has no additional explanatory power for the conditional mean of the dependent variable $Y$, given that $X$ has a statistically significant effect. The null hypothesis of interest is
\begin{align}\label{hyp.null}
H_0:\,\mathbb{E}\left[Y\vert W\right] = \mathbb{E}\left[Y\vert X\right]\quad a.s.,
\end{align}
and the alternative hypothesis, $H_1$, is the negation of $H_0$. Let $f_X(\cdot)$ denote the probability density function (PDF) of $X$. Using the fact that the PDF $f_X(X)>0$ $a.s.$, we can rewrite $H_0$ in \eqref{hyp.null} as
\begin{align}\label{hyp.null f}
H_0:\, f_X(X)\mathbb{E}\left[\epsilon\vert W\right]=0\quad a.s.,
\end{align}
where $\epsilon=Y-\mathbb{E}[Y\vert X]:=Y-m(X)$ is the nonparametric error, with $m(X)$ denoting the unknown conditional mean function under the null.
It is well established in the testing literature that the conditional moment restriction in \eqref{hyp.null f} is equivalent to a continuum of unconditional moment restrictions. More specifically, following the ICM principle introduced by \cite{bierens1982consistent}, and further developed by \cite{stute1997nonparametric}, \cite{stute1998bootstrap}, \cite{delgado2001significance}, \cite{stute2002model}, and \cite{escanciano2006consistent}, we employ the indicator function $1_w(W):=1(W\leq w)=1(X\leq x)1(Z\leq z)$ with $w=(x^\top,z^\top)^\top\in\mathbb R^{q+p}$ as the weighting scheme to transform the conditional moment restriction to the unconditional version. Consequently, the null hypothesis $H_0$ in \eqref{hyp.null f} can be equivalently characterized by the following infinite number of unconditional moment conditions indexed by $w$:
\begin{align}\label{hyp.null eq}
H_0:\,\mathbb{E}\left[\epsilon f_X(X)1_w(W)\right]=0\quad \text{for all }w\in\mathbb{R}^{q+p},
\end{align}
which converts the original nonparametric significance testing problem into a global testing problem. Therefore, testing \eqref{hyp.null eq} achieves the fastest possible parametric rate.
Suppose that we have a random sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$ of size $n\geq 1$ consisting of independent and identically distributed (i.i.d.) random variables, the natural sample analog for \eqref{hyp.null eq} is given by
\begin{align*}
\hat T_n(w) =& \frac{1}{n}\sum_{i=1}^n\hat{\epsilon}_i\hat{f}_X(X_i)1_w(W_i)\\
=&\frac{1}{n(n-1)}\sum_{i=1}^n\sum_{j=1,j\neq i}^n\frac{1}{a^q}K\left(\frac{X_i-X_j}{a}\right)(Y_i-Y_j)1_w(W_i),
\end{align*}
where $\hat{\epsilon}_i=Y_i-\hat{m}(X_i)$ is the nonparametric residual, with the leave-one-out nonparametric estimators for $m(X_i)$ and $f_X(X_i)$ given by
\begin{align*}
\hat{m}(X_i) = \frac{1}{\hat{f}_X(X_i)}\frac{1}{(n-1)a^q}\sum_{j=1,j\neq i}^nK\left(\frac{X_i-X_j}{a}\right)Y_j
\end{align*}
and
\begin{align*}
\hat{f}_{X}(X_i) = \frac{1}{(n-1)a^q}\sum_{j=1,j\neq i}^nK\left(\frac{X_i-X_j}{a}\right),
\end{align*}
respectively. Here, $a=a(n)\in\mathbb{R}^{+}$ is a bandwidth parameter shrinking to zero at a suitable rate as $n\to\infty$ and $K(u)=\Pi_{j=1}^qk(u^{(j)})$ is a product kernel for a $q$-dimensional vector $u$, where $k(\cdot)$ is the univariate kernel and $u^{(j)}$ denotes the $j$-th component of $u$. Note that the random denominator problem is effectively eliminated by incorporating the estimated density $\hat{f}_X(X_i)$ into the $U$-process $\hat T_n(\cdot)$, significantly facilitating the theoretical analysis of $\hat T_n(\cdot)$.
By imposing the standard regularity conditions considered in \cite{delgado2001significance} (see their Assumptions A1--A5), under the null $H_0$, we have
\begin{align}\label{stat.delgado}
\sup_{w}\left\vert \hat T_n(w)-\frac{1}{n}\sum_{i=1}^n\epsilon_if_X(X_i)1_x(X_i)\left[1_z(Z_i)-F_{Z\vert X}(z\vert X_i)\right]\right\vert = o_p\left(n^{-1/2}\right),
\end{align}
where $F_{Z\vert X}(z\vert x)$ is the conditional distribution function of $Z$ given $X$. It is observed that in the asymptotically uniform expansion \eqref{stat.delgado}, the term
\begin{align*}
\frac{1}{n}\sum_{i=1}^n\epsilon_if_X(X_i)1_x(X_i)1_z(Z_i)
\end{align*}
is the infeasible process for testing the null hypothesis in \eqref{hyp.null eq} if $\epsilon_i$ is known (equivalently if $m(X_i)$ is known), and the term
\begin{align*}
\frac{1}{n}\sum_{i=1}^n\epsilon_if_X(X_i)1_x(X_i)F_{Z\vert X}(z\vert X_i)
\end{align*}
can be regarded as the ``nonparametric estimation effect'' due to using the nonparametric residual $\hat{\epsilon}_i$ to replace $\epsilon_i$ (via the estimation of $m(X_i)$).
Much less attention has been paid to studying the above term, unlike the ``parametric estimation effect'' also known as the ``Durbin problem'' mentioned in \cite{durbin1973distribution}. Specifically, since the asymptotic null distributions of the test statistics (constructed as continuous functionals of $\hat T_n(\cdot)$) are typically case-dependent, the use of bootstrap methods to obtain critical values is unavoidable. However, the ``nonparametric estimation effect'' in the asymptotically uniform expansion of $\hat T_n(\cdot)$ poses a substantial challenge to the direct application of the computationally efficient multiplier bootstrap. While the wild bootstrap is a potential alternative, it incurs significantly higher computational complexity. To address this, with the help of \eqref{stat.delgado}, \cite{delgado2001significance} propose a specifically constructed multiplier bootstrap-based $U$-process as follows:
\begin{align*}
\frac{1}{n}\sum_{i=1}^nV_i\hat{\epsilon}_i\hat{f}_X(X_i)1_x(X_i)\left[1_z(Z_i)-\hat{F}_{Z\vert X}(z\vert X_i)\right],
\end{align*}
with $\{V_i\}_{i=1}^n$ being i.i.d. random variables (i.e., multipliers) that have mean zero, unit variance, and are independent of the original sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$. However, this approach necessitates estimating the conditional distribution function $F_{Z\vert X}(z\vert X_i)$ nonparametrically within the multiplier bootstrap procedure, which is
\begin{align*}
\frac{1}{\hat{f}_X(X_i)}\frac{1}{(n-1)a^q}\sum_{j=1,j\neq i}^nK\left(\frac{X_i-X_j}{a}\right)1_z(Z_j).
\end{align*}
See Section $3$ of \citet{delgado2001significance} for a detailed description of this implementation. Unfortunately, this inevitably reintroduces the random denominator issue in the multiplier bootstrap, thereby undermining the purpose of using $\hat T_n(w)$.
Indeed, the solution adopted by \citet{delgado2001significance} to ensure the validity of this bootstrap is to impose stronger assumptions, such as requiring the density $f_X(X)$ to be bounded away from zero (see their Assumptions A9), which was exactly intended to be avoided in the initial construction of $\hat T_n(w)$ through the density weighting $\hat f_X(X_i)$. Such a restrictive bounded support assumption constitutes a significant barrier to applying the test to data with unbounded support, including the commonly used normal distribution. Another potential solution is to use trimming; however, the prohibitive complexity of the required theoretical proofs and the practical choice of the trimming level have prevented its successful implementation.
In this paper, to eliminate the nonparametric estimation effect, we propose a tailor-made nonparametric significance test that depends on a novel nonparametric-type projection of the weighting function $1_z(Z_i)$ onto the orthocomplement of the space spanned by $f_W(W_i)$ (hereafter, the PDF of $W$), where orthogonality is understood conditionally on $X_i$. To this end, we first consider the following infeasible projected weighting function:
\begin{align*}
\mathcal{P}1_z(Z_i) :=& 1_z(Z_i)-f_{Z\vert X}(Z_i\vert X_i)\left[\int_{-\infty}^{\infty}f_{Z\vert X}^2(\bar z\vert X_i)\,d\bar z\right]^{-1}\int_{-\infty}^zf_{Z\vert X}(\bar z\vert X_i)\,d\bar z\notag\\
=&1_z(Z_i)-f_W(W_i)\Delta^{-1}(X_i)G(z;X_i),
\end{align*}
where $f_{Z\vert X}(z\vert x)$ is the conditional density function of $Z$ given $X$,
\begin{align*}
\Delta(X_i) = \int_{-\infty}^{\infty}f_{W}^2(X_i,\bar z)\,d\bar z, \quad \text{and}\quad G(z;X_i) = \int_{-\infty}^zf_{W}(X_i,\bar z)\,d\bar z.
\end{align*}
In fact, the projection constructed above admits an intuitive interpretation as the linear regression of $1_z(Z_i)$ on $f_W(W_i)$ conditional on $X_i$, where $\Delta^{-1}(X_i)G(z;X_i)$ represents the coefficients of the linear projection. In other words, the term $f_W(W_i)\Delta^{-1}(X_i)G(z;X_i)$ constitutes the best linear predictor of $1_z(Z_i)$ given $f_W(W_i)$. Consequently, our constructed weighting function $\mathcal{P}1_z(Z_i)$ corresponds precisely to the associated prediction error, which implies that $\mathcal{P}1_z(Z_i)$ is orthogonal to $f_W(W_i)$ conditional on $X_i$, i.e.,
\begin{align*}
\mathbb{E}\left[f_W(W_i)\mathcal{P}1_z(Z_i)\vert X_i\right] = f_X^{-1}(X_i)\left[G(z;X_i) - \Delta(X_i)\Delta^{-1}(X_i)G(z;X_i)\right]\equiv 0 \,\,\,a.s.
\end{align*}
\begin{remark}
While this paper introduces a nonparametric orthogonal projection approach, analogous parametric orthogonal projection-based techniques are well established in the testing literature, primarily for assessing the correct specification of parametric models. Examples include \cite{escanciano2014specification}, \cite{sant2019specification}, \cite{yang2024model}, and \cite{song2025unified}, which utilize orthogonal projected weighting functions onto the score functions to handle the ``parametric estimation effect'' in testing the (partially) linear quantile regression processes, propensity score models, and partially linear spatial autoregressive models.
\end{remark}
Building on the sample version of the nonparametric projection $\mathcal{P}1_z(Z_i)$, our proposed projected $U$-process
is constructed as follows:
\begin{align}\label{stat.stat}
\hat R_n(w) = \frac{1}{n}\sum_{i=1}^n\hat{\epsilon}_i\hat{f}_X(X_i)1_x(X_i)\hat{\mathcal{P}}_n1_z(Z_i).
\end{align}
where
\begin{align*}
\hat{\mathcal{P}}_n1_z(Z_i) = 1_z(Z_i)-\hat{f}_{W}(W_i)\hat{\Delta}_n^{-1}(X_i)\hat{G}_n(z;X_i)
\end{align*}
is the natural estimator for $\mathcal{P}1_z(Z_i)$. Here,
\begin{align*}
\hat{\Delta}_n(X_i) = \int_{-\infty}^{\infty}\hat{f}^2_W(X_i,\bar z)\,d\bar z \quad \text{and}\quad \hat{G}_n(z;X_i) = \int_{-\infty}^z\hat{f}_W(X_i,\bar z)\,d\bar z
\end{align*}
are the sample versions of $\Delta(X_i)$ and $G(z;X_i)$, respectively, with
\begin{align*}
\hat{f}_W(X_i,z) = \frac{1}{(n-1)a^qb^{p}}\sum_{j=1,j\neq i}^nK\left(\frac{X_i-X_j}{a}\right)L\left(\frac{z-Z_j}{b}\right)
\end{align*}
being the leave-one-out kernel density estimator for the PDF $f_W(X_i,z)$, where $L(\cdot)$ and $b=b(n)\in\mathbb R^+$ are the kernel and bandwidth, respectively.
\begin{remark}
We construct the kernel $L(\cdot)$ for the $p$-dimensional vector $Z$ as a product of the univariate kernel $l(\cdot)$ and introduce a second bandwidth $b$ that shrinks to zero at a suitable rate as $n\to\infty$. Such a setting is standard in the literature on smoothing-based methods, where the additional bandwidth $b$ introduces flexibility in characterizing bias and is typically used to control the relative rate with respect to $a$; see, for example, \cite{fan1996consistent} and \cite{lavergne2000nonparametric}. In our setting, however, the restrictions on the bandwidths $a$ and $b$ are less stringent, requiring only that they satisfy Assumption \ref{ass.bandwidth} in Section \ref{sec.Asy}. One may even set $a=b$ and take $K(\cdot)$ and $L(\cdot)$ to be the same kernel without affecting the theoretical results and the implementation of the proposed tests, as illustrated in the simulations reported in Section \ref{sec.Simulation}.
\end{remark}
Note that the proposed process $\hat R_n(w)$ based on the nonparametric-type orthogonal projection $\hat{\mathcal{P}}_n1_z(Z_i)$ can account for effectively the ``nonparametric estimation effect'' discussed before, although at the cost of nonparametrically estimating the density $f_W$. As such, it can be implemented directly using a convenient multiplier bootstrap without reintroducing the random denominator, in sharp contrast to the multiplier bootstrap used by \citet{delgado2001significance}. Further details are provided in Section \ref{sec.boot}. The test statistics are then constructed as continuous functionals of $\hat R_n(w)$. In this paper, we focus on the two most commonly employed choices, namely, the Cram\'{e}r--von Mises (CvM) and the Kolmogorov--Smirnov (KS) test statistics, which are explicitly given as
\begin{align*}
CvM_n = \int\left\vert \hat R_n(w)\right\vert^2F_{W_n}(dw) \quad \text{and}\quad KS_n = \sup_{w}\left\vert \hat R_n(w)\right\vert.
\end{align*}
From a theoretical perspective, the integrating measure $F_{W_n}(\cdot)$ appearing in the $CvM_n$ is assumed to be a random measure that converges in probability to $F_W(\cdot)$, which is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{q+p}$. In practice, we adopt the approach suggested in \cite{escanciano2006consistent}: for the computation of $CvM_n$ we take $F_{W_n}(\cdot)$ to be the empirical distribution function of $\{W_i\}_{i=1}^n$, and for the computation of $KS_n$ we replace the theoretical supremum by its maximum taken over the sample points. Under the null hypothesis $H_0$, the test statistics $CvM_n$ and $KS_n$ are expected to take sufficiently small values, whereas relatively large realizations provide evidence against $H_0$ in favor of the alternative hypothesis $H_1$. The critical values used to assess the magnitude of the test statistics are obtained via the multiplier bootstrap procedure detailed in Section \ref{sec.boot}.
\section{Theoretical results}\label{sec.Asy}
In this section, the large-sample properties of the proposed $CvM_n$ and $KS_n$ test statistics will be considered, with particular attention given to their asymptotic null distributions, local power, and consistency. Before presenting the theoretical results, it is necessary to adopt several definitions as introduced in \cite{delgado2001significance} and to state a set of regularity conditions that hold uniformly across all cases. Let $\mathscr{K}_l$ with $l\geq1$ be the class of even functions of uniformly bounded variation $k:\mathbb{R}\to\mathbb{R}$, which satisfy
\begin{align*}
k(u) = O\left((1+\left\vert u\right\vert^{l+1+\eta})^{-1}\right) \text{ for some }\eta>0 \text{ and }\int_\mathbb{R}u^ik(u)\,du=\delta_{i0}\text{ for }i=0,\cdots,l-1,
\end{align*}
where $\delta_{ij}$ is the Kroneker's delta. Let $\mathscr{L}_\beta^\alpha$ with $\alpha>0$ and $\beta>0$ be the class of functions $g:\mathbb{R}^q\to\mathbb{R}$, which satisfy uniformly $(b-1)$-times continuously differentiability for $b-1\leq \beta\leq b$. In addition, we impose the assumption that there exist a positive constant $\rho$ and a function $d$ with finite $\alpha$-th moments such that
\begin{align*}
\sup_{\left\vert v-u\right\vert<\rho}\left\vert g(v)-g(u)-Q(v,u)\right\vert/\left\vert v-u\right\vert^\beta\leq d(u)\text{ for all } u,
\end{align*}
where $Q$ is a $(b-1)$-th degree homogeneous polynomial in $v-u$ with coefficients given by the partial derivatives of $g$ at $u$ of orders up to $b-1$, all of which are assumed to have finite $\alpha$-th moments; in particular, $Q=0$ when $b=1$.
\begin{Assumption}\label{ass.sample}
$\{(Y_i,W_i^\top)^\top\}_{i=1}^n$ are i.i.d. observations drawn from $(Y,W^\top)^\top$, which satisfies $\mathbb{E}\vert Y-m(X)\vert^{2+\delta_1+\delta_2}<\infty$ and $\mathbb{E}\vert \Delta^{-1}(X)f_X^2(X)\vert^{2+(4+2\delta_2)/\delta_1}<\infty$ for some positive constants $\delta_1$ and $\delta_2$.
\end{Assumption}
\begin{Assumption}\label{ass.structual}
$f_X(\cdot)\in\mathscr{L}_{\lambda_1}^{\infty}$, $f_{Z\vert X}(z|\cdot)\in\mathscr{L}_{\lambda_2}^{\infty}$, and $m(\cdot)\in\mathscr{L}_\tau^2$ for some positive constants $\lambda_1$, $\lambda_2$, and $\tau$. $k(\cdot)\in\mathscr{K}_{\max\{l_1,l_2\}+t-1}$, where $l_1-1<\lambda_1\leq l_1$, $l_2-1<\lambda_2\leq l_2$, and $t-1<\tau\leq t$. In addition, we assume $\sup_{(x,z)}\vert f_{Z\vert X}(z|x)/f_X(x)\vert<\infty$.
\end{Assumption}
\begin{Assumption}\label{ass.bandwidth}
$(na^{2q}b^{2p})^{-1}(\log n)^{2}+na^{2\min(\tau,\lambda_1)}+nb^{2\min(\tau,\lambda_2)}\to 0$ as $n\to \infty$.
\end{Assumption}
Assumption \ref{ass.sample} concerns the moments of $\epsilon=Y-m(X)$ and $\Delta^{-1}(X)f_X^2(X)$.
While \cite{fan1996consistent} required the existence of the fourth moment, this condition was relaxed in \cite{delgado2001significance}. We adopt the relaxed version to guarantee the finiteness of the variances of the statistics, which is crucial for establishing the theoretical results. Moreover, the assumption concerning $\Delta^{-1}(X)f_X^2(X)$ is essential to guarantee the convergence of the proposed projection structure. In fact, it can be reformulated as the moment condition $\mathbb{E}\vert\int f_{Z\vert X}^2(\bar z|X)\,d\bar z\vert^{-2-(4+2\delta_2)/\delta_1}<\infty$, which restricts the tail properties of the conditional density function $f_{Z|X}$ and the density function $f_X$ and is satisfied by most commonly used distributions. This assumption permits an unbounded covariate $X$ and is weaker than the stringent bounded–support assumption on $X$ as imposed in \cite{delgado2001significance} (see their Assumptions A9).
This theoretical refinement extends the applicability of the ICM–type tests for nonparametric moment restrictions to data generated from normal and other commonly used unbounded distributions that could not be handled under the conditions in \cite{delgado2001significance}, as confirmed by the simulation results for the unbounded covariate case in Section \ref{sec.Simulation}.
The function classes specified in Assumption \ref{ass.structual} impose the Lipschitz continuity of the
functions $f_X(\cdot)$, $f_{Z\vert X}(z|\cdot)$, and $m(\cdot)$, which in turn ensures that the bias of the constructed empirical process is asymptotically negligible and the variance is finite.\footnote{We note that the condition $\sup_{(x,z)}\vert f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$ is sufficient but not necessary and permits an unbounded covariate $X$. This condition is assumed to control uniformly the bias term of $\hat{\Delta}_n(x)$, as shown in Lemma \ref{lemma.delta}. In fact, alternative conditions may also deliver the desired result. Let $d_{f_X}(\cdot)$ denote the Lipschitz coefficient of $f_X(\cdot)$ and $d_{f_{Z\vert X}}(z\vert \cdot)$ denote the Lipschitz coefficient of $f_{Z\vert X}(z\vert \cdot)$, in the sense of the definitions of $\mathscr{L}_{\lambda_1}^\infty$ and $\mathscr{L}_{\lambda_2}^\infty$, respectively. For example, if the conditions $\sup_{(x,z)}\vert d_{f_X}(x)f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$, $\sup_{(x,z)}\vert d_{f_{Z\vert X}}(z\vert x)\vert<\infty$, and $\sup_{(x,z)}\vert d_{f_X}(x)d_{f_{Z\vert X}}(z\vert x)/f_X(x)\vert<\infty$ are imposed in place of $\sup_{(x,z)}\vert f_{Z\vert X}(z\vert x)/f_X(x)\vert<\infty$, the same conclusion continues to hold.
} Note that standard densities, such as the normal distribution, and widely used linear models satisfy this assumption. In particular, the function class is restricted to be bounded when $\alpha=\infty$, for example, in the case of density functions. Moreover, the requirements on the kernel function $k(\cdot)$ strengthen the conventional high-order kernel assumptions by imposing a decay rate condition, as mentioned in \cite{robinson1988root}. In addition, the order of the kernel function depends on the moments of the
functions $f_X(\cdot)$, $f_{Z\vert X}(z|\cdot)$, and $m(\cdot)$.
Assumption \ref{ass.bandwidth} determines the lower and upper bounds for the bandwidth sequences $a$ and $b$ and thus guarantees the convergence of the proposed $U$-process. On the one hand, the lower bounds on the bandwidths, which are consistent with those in \cite{robinson1988root} and \cite{delgado2001significance}, serve to ensure the asymptotic negligibility of the bias of the $U$-process. On the other hand, the upper bounds on the bandwidths in Assumption \ref{ass.bandwidth} provide a sufficient, though not necessarily weakest possible, condition to establish the uniform convergence of the nonparametric estimator $\hat{\Delta}_n(\cdot)$ for $\Delta(\cdot)$ as shown in Lemma \ref{lemma.delta}. The slightly strengthened assumptions on the upper bounds are used to simplify the theoretical analysis. Specifically, we expand $\hat{\Delta}_n(\cdot)$ only up to the second order, although the upper bounds could be relaxed with higher order expansions of $\hat{\Delta}_n(\cdot)$. The upper bounds cannot, however, be relaxed beyond the rate $(na^{2q}b^{p})^{-1}\to 0$, which guarantees that the variance of the test statistics remains bounded.
In the following, let ``$\Longrightarrow$'' represent weak convergence on $(l^{\infty}(\Pi),\mathcal{B}_\infty)$ in the sense of Hoffmann–-J\textup{\o{}}rgensen, where $\mathcal{B}_\infty$ denotes the corresponding Borel $\sigma$-algebra, see, e.g., Definition $1.3.3$ in \cite{van1996weak}. We establish formally the uniform decomposition of $\hat R_n(\cdot)$, which indicates that the nonparametric estimation effects arising from $\hat m$ and $\hat f_W$ are asymptotically negligible. Denote
\begin{align*}
\zeta(Y,X,Z;x,z)=\left[Y-m(X)\right]f_X(X)1_x(X)\mathcal{P}1_z(Z).
\end{align*}
\begin{theorem}\label{thm.null}
Suppose that Assumptions \ref{ass.sample}--\ref{ass.bandwidth} hold. Under the null hypothesis $H_0$,
\begin{align}\label{thm.eq null}
\sup_{w}\left\vert \hat R_n(w)-\frac{1}{n}\sum_{i=1}^n\zeta(Y_i,X_i,Z_i;x,z)\right\vert=o_p\left(n^{-1/2}\right).
\end{align}
Furthermore,
\begin{equation*}
\sqrt{n}\hat R_{n}(\cdot) \Longrightarrow R_{\infty}(\cdot),
\end{equation*}
where $R_{\infty}(\cdot)$ is a centered Gaussian process with the covariance structure given by
\begin{align*}
&Cov\left[R_{\infty}(w), R_{\infty}(w^\prime)\right]= \mathbb{E}\left[\zeta(Y,X,Z;x,z)\zeta(Y,X,Z;x^\prime,z^\prime)\right].
\end{align*}
\end{theorem}
To complement the theoretical results under the null hypothesis, and building upon Theorem \ref{thm.null}, the asymptotic distributions of the $CvM_n$ and $KS_n$ statistics introduced in Section \ref{sec.Test} are established by applying the continuous mapping theorem, as discussed in \cite{van1996weak}.
\begin{corollary}\label{Cor.null}
Suppose that Assumptions \ref{ass.sample}--\ref{ass.bandwidth} hold. Under the null hypothesis $H_0$,
\begin{align*}
&nCvM_{n}\stackrel{d}\longrightarrow \int\left\vert R_{\infty}(w)\right\vert^2F_W(dw) \text{ and } \sqrt{n}KS_{n}\stackrel{d}\longrightarrow \sup\limits_{w}\left\vert R_{\infty}(w)\right\vert.
\end{align*}
\end{corollary}
Theorem \ref{thm.null} establishes the theoretical advantage of the proposed test, namely, that despite the applications of kernel smoothing methods, the resulting empirical process still achieves the parametric convergence rate, with a centered Gaussian process as its limiting process. Subsequently, the Corollary \ref{Cor.null} ensures that the statistics $nCvM_{n}$ and $\sqrt{n}KS_{n}$ converge to the squared norm and the sup norm of the aforementioned Gaussian process, respectively. The uniform decomposition given in \eqref{thm.eq null} is both theoretically and empirically attractive for implementing the multiplier bootstrap detailed in Section \ref{sec.boot}. To investigate the power of the proposed test, we introduce a sequence of local alternatives converging to the null hypothesis at the parametric rate,
\begin{align}\label{hyp.local}
H_{1n}:\mathbb{E}(Y \vert W)=\mathbb{E}(Y \vert X)+\frac{\Psi(W)}{\sqrt{n}} \quad a.s.,
\end{align}
under which the asymptotic theory of the constructed empirical process is subsequently established. We note that $\Psi:\mathbb{R}^{p+q}\to\mathbb{R}$ is a bounded nonzero function.
\begin{theorem}\label{thm.local}
Suppose that Assumptions \ref{ass.sample}--\ref{ass.bandwidth} hold. Under the sequence of local alternatives $H_{1n}$,
\begin{align*}
\sqrt{n}\hat R_{n}(\cdot) \Longrightarrow R_{\infty}(\cdot)+\mu(\cdot),
\end{align*}
where $R_{\infty}(\cdot)$ is the centered Gaussian process as described in Theorem \ref{thm.null}, and the deterministic shift function $\mu(w)$ is given by
\begin{align*}
\mu(w) = \mathbb{E}\left\{\Psi(W)f_X(X)1_x(X)\left[1_z(Z)-f_W(W)\Delta^{-1}(X)G(z;X)\right]\right\}.
\end{align*}
\end{theorem}
Owing to the similarity between the theoretical derivations under the sequence of local alternatives and those established under the null hypothesis as in Corollary \ref{Cor.null}, the results $nCvM_{n}\stackrel{d}\longrightarrow \int\vert R_{\infty}(w)+\mu(w)\vert^2F_W(dw)$ and $\sqrt{n}KS_{n}\stackrel{d}\longrightarrow \sup\limits_{w}\vert R_{\infty}(w)+\mu(w)\vert$ are presented directly, where we omit the verification details of the convergence of the statistics. Consequently, the local powers of the proposed statistics are characterized by $\mu(\cdot)$. Specifically, the deterministic shift term determines the non-centrality of the limiting process in comparison with the centered Gaussian limiting process under the null hypothesis, and it leads to the fact that whenever the deterministic shift function is nonzero for at least some $w\in\mathbb R^{p+q}$ with a positive Lebesgue measure, the proposed $CvM_n$ and $KS_n$ statistics possess nontrivial power against the local alternatives. Indeed, since $\mu(\cdot)$ depends on the conditional $L^2$ inner product of $\Psi(X,Z)$ and $\mathcal{P}1_z(Z)$ given $X$, in the sense that $\mu(w) = \mathbb{E}\{f_X(X)1_x(X)\mathbb{E}[\Psi(W)\mathcal{P}1_z(Z)\vert X]\}$, the nontrivial local power is well-established as long as for any constant $\gamma$,
\begin{align*}
\mathbb{P}\left[\Psi(W)=\gamma f_W(W)\right]<1.
\end{align*}
Finally, to derive the asymptotic global power properties of the proposed test statistics, we investigate the asymptotic behavior of the constructed empirical process under the alternative hypothesis.
\begin{theorem}\label{thm.alt}
Suppose that Assumptions \ref{ass.sample}--\ref{ass.bandwidth} hold. Under the alternative hypothesis $H_1$,
\begin{equation*}
\sup_{w\in\Pi}\left\vert \hat R_{n}(w) - C(w) \right\vert = o_p(1),
\end{equation*}
where
\begin{align*}
&C(w) = \mathbb{E}\left\{\left[\mathbb{E}\left(Y\vert W\right)-\mathbb{E}\left(Y\vert X\right)\right]f_X(X)1_x(X)\left[1_z(Z)-f_W(W)\Delta^{-1}(X)G(z;X)\right]\right\}.
\end{align*}
\end{theorem}
The correspondence between local and global power is straightforward. The local power is driven by $\mu(\cdot)$, and the global power, in turn, depends on the conditional $L^2$ inner product of $\mathbb{E}(Y\vert W)-\mathbb{E}(Y\vert X)$ and $\mathcal{P}1_z(Z)$ given $X$. More specifically, once
\begin{align}\label{thm.alt eq}
\mathbb{P}\left\{\left[\mathbb{E}\left(Y\vert W\right)-\mathbb{E}\left(Y\vert X\right)\right]=\gamma f_W(W)\right\}<1
\end{align}
is satisfied for any constant $\gamma$, the unconditional expectation $C(w)\neq 0$. This in turn implies that, as $n$ goes to infinity, the constructed empirical process converges uniformly in probability to a non-vanishing function $C(\cdot)$, which further leads the test statistics to diverge to positive infinity in probability. However, we do not regard the potential loss of power along certain directions of the alternative space as a primary empirical concern. As shown in \cite{bierens1997asymptotic}, ICM–type tests are admissible under suitable regularity conditions, so that no uniformly more powerful procedure exists within a broad class of alternatives. These trade–offs are detailed in the simulation studies in Section \ref{sec.Simulation}.
\section{Multiplier bootstrap}\label{sec.boot}
The case-dependent asymptotic limiting process established in Section \ref{sec.Asy} necessitates the consideration of bootstrap procedures. In the traditional literature, two distinct bootstrap schemes are frequently considered. The residual-based bootstrap employed in \cite{koul1994bootstrapping} generates bootstrap samples by resampling from the empirical distribution of the estimated residuals (possibly after a smoothing step), and has been widely employed in classical nonparametric testing procedures. In contrast, the wild bootstrap, as mentioned in \cite{mammen1993bootstrap}, constructs bootstrap errors by multiplying the estimated residuals with independent random multipliers of zero mean and unit variance, thereby creating bootstrap samples without explicitly resampling from the residual distribution.
Although traditional bootstrap methods may be justified after careful theoretical verification, we consider more computationally efficient alternatives. Motivated by \cite{van1996weak}, the uniform decomposition in \eqref{thm.eq null} inspires an intuitive, easy-to-implement multiplier bootstrap procedure. Towards this end, define the multiplier bootstrapped projected process as
\begin{align}\label{stat.boot}
\hat R_n^\ast(w)=\frac{1}{n}\sum_{i=1}^n V_i\hat{\epsilon}_i\hat{f}_X(X_i)1_x(X_i)\hat{\mathcal{P}}_n1(Z_i\leq z),
\end{align}
where $\{V_i\}_{i=1}^n$ is a sequence of random variables (i.e., multipliers) that are mean zero, unit variance, and independent of the original sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$. It is worth noting that in contrast to the multiplier bootstrap proposed in \cite{delgado2001significance}, which is motivated by the uniform decomposition in \eqref{stat.delgado}, our multiplier bootstrap version $\hat R_n^\ast(w)$ does not involve a random denominator thanks to the novel projection, and thus it is not necessary to impose the stringent compact support assumption $\mathbb{P}(f_X(X)>\nu)=1$ for some constant $\nu>0$. This fact allows us to
accommodate important distributions for X with unbounded support, like the t and normal.
\begin{theorem}\label{thm.boot}
Suppose that Assumptions \ref{ass.sample}--\ref{ass.bandwidth} hold. Under the null hypothesis $H_0$ or the sequence of local alternatives $H_{1n}$,
\begin{align*}
\sqrt{n}\hat R^{\ast}_{n}(\cdot) \underset{\ast}{\Longrightarrow } R_{\infty}(\cdot)\quad \text{in probability},
\end{align*}
where \textquotedblleft$\underset{\ast}{\Longrightarrow }$\textquotedblright denotes weak convergence under the bootstrap law, i.e., conditional on the original sample $\{(Y_i,W_i^\top)^\top\}_{i=1}^n$, and $R_{\infty}(\cdot)$ is the centered Gaussian process as described in Theorem \ref{thm.null}. Under the alternative hypothesis $H_1$,
\begin{align*}
\sqrt{n}\hat R^{\ast}_{n}(\cdot) \underset{\ast}{\Longrightarrow } R^1_{\infty}(\cdot)\quad \text{in probability},
\end{align*}
where $R^1_{\infty}(\cdot)$ is a centered Gaussian process that is different from $R_{\infty}(\cdot)$.
\end{theorem}
It is worth noting that, under the null hypothesis, the bootstrap version coincides with the original sample-based empirical process in both convergence rate and asymptotic distribution. By constructing the multiplier bootstrap versions of the $CvM_n$ and $KS_n$ statistics in a manner analogous to Section \ref{sec.Test},
\begin{align*}
CvM_n^\ast = \int\left\vert \hat R^\ast_n(w)\right\vert^2F_{W_n}(dw) \quad \text{and}\quad KS_n^\ast = \sup_{w}\left\vert \hat R^\ast_n(w)\right\vert,
\end{align*}
and applying the continuous mapping theorem in a way analogous to that employed in Corollary \ref{Cor.null}, we establish that under the null hypothesis the bootstrap $CvM_n^\ast$ and $KS_n^\ast$ statistics consistently approximate the asymptotic null distributions of $CvM_n$ and $KS_n$, respectively, thereby ensuring that the empirical sizes of the two tests are theoretically close to the nominal level. Under local alternatives $H_{1n}$, however, the local power depends on $\mu(\cdot)$ introduced in Section \ref{sec.Asy}. This is because the bootstrap versions converge to the same limiting processes as under the null, whereas the original statistics converge to limiting processes with additional drift components. Therefore, taking the $CvM_n$ statistic as an example, the different limiting distributions of $CvM_n$ and its bootstrap counterpart $CvM_n^\ast$ imply that the proportion of rejecting $H_0$ will exceed the nominal level, thereby confirming the nontrivial local power of the proposed test under $H_{1n}$. As a consequence of the above analysis, the asymptotic critical value (still taking the $CvM_n$ statistic as an example) at the significance level $\alpha$ is $c^\ast_{\alpha}=\inf\{c_\alpha\in[0,\infty):\lim_{n\to\infty}\mathbb{P}_n^\ast(nCvM_n^\ast>c_\alpha)=\alpha\}$, where $\mathbb{P}_n^\ast$ is the bootstrap probability under the bootstrap law. In practice, $c^\ast_{\alpha}$ can be approximated as $c^\ast_{n,\alpha} = \{nCvM_n^\ast\}_{B(1-\alpha)}$, the $B(1-\alpha)$-th order statistic for $B$ replicates $\{nCvM_{n,b}^\ast\}_{b=1}^B$
and we reject $H_0$ if $nCvM_n>c^\ast_{n,\alpha}$. Finally, under the alternative hypothesis $H_1$, when the condition \eqref{thm.alt eq} is satisfied, the consistency of our bootstrap statistics against $H_1$ follows from the fact that the original statistics diverge to infinity in probability, whereas the bootstrap versions remain stochastically bounded.
\section{Testing conditional independence}\label{sec.CI}
The projection techniques developed in this article to handle the nonparametric estimation effect can readily be extended to testing other restrictions on regression curves, thereby allowing ICM-type tests, which are particularly advantageous in certain settings, to be applied to these important problems. A particularly important example is the problem of testing conditional independence, which has attracted considerable attention in recent years and has been emphasized as a key issue in, among others, \cite{su2014testing}, \cite{huang2016flexible}, \cite{wang2018characteristic}, \cite{li2020nonparametric}, \cite{shah2020hardness}, \cite{neykov2021minimax} and \cite{cai2022distribution}. Formally, we are interested in testing whether $Y$ is independent of $Z$ given $X$, denoted by $H_0^{CI}:Y\perp Z\vert X$. Using the previous notations, we consider testing
\begin{align*}
H_0^{CI}:\mathbb{E}\left[1_y(Y)\vert W\right] = \mathbb{E}\left[1_y(Y)\vert X\right]\quad a.s. \text{ for all } y\in \mathbb{R},
\end{align*}
which is further equivalent to
\begin{align*}
H_0^{CI}:\mathbb{E}\left[\epsilon(y)f_X(X)1_w(W)\right] = 0\text{ for all } (y,w)\in \mathbb{R}^{1+q+p},
\end{align*}
where $\epsilon(y) = 1_y(Y)-\mathbb{E}[1_y(Y)\vert X] = 1_y(Y)-F_{Y\vert X}(y\vert X)$, with $F_{Y\vert X}(y\vert x)$ the conditional distribution function of $Y$ given $X$. Similar to $T_n(w)$ for the case of significance testing in the conditional mean, \cite{delgado2001significance} have proposed to use the following unprojected $U$-process to test $H_0^{CI}$:
\begin{align*}
\hat L_n(y,w) = &\frac{1}{n}\sum_{i=1}^n\hat{\epsilon_i}(y)\hat{f}_X(X_i)1_w(W_i)\\
=&\frac{1}{n(n-1)}\sum_{i=1}^n\sum_{j=1,j\neq i}^n\frac{1}{a^q}K\left(\frac{X_i-X_j}{a}\right)\left[1_y(Y_i)-1_y(Y_j)\right]1_w(W_i),
\end{align*}
which is shown to satisfy
\begin{align}\label{stat.delgado ci}
\sup_w\left\vert \hat L_n(y,w)-\frac{1}{n}\sum_{i=1}^n\epsilon_i(y)f_X(X_i)1_x(X_i)\left[1_z(Z_i)-F_{Z\vert X}(z\vert X_i)\right]\right\vert = o_p\left(n^{-1/2}\right).
\end{align}
Motivated by the uniform decomposition in \eqref{stat.delgado ci} and the expected case-dependent asymptotic limiting process, we would need to obtain critical values for the tests based on $\hat L_n(\cdot)$ via a multiplier bootstrap scheme similar to that described in Section \ref{sec.boot}. This, however, requires the nonparametric estimation of $F_{Z\vert X}(z|X_i)$. The random denominator problem leads either to overly stringent assumptions on the distribution of $X$ or to substantial additional technical difficulties. Similar to the discussion in Section \ref{sec.Test}, when ICM-type tests are used to test conditional independence, \textquotedblleft nonparametric estimation effect\textquotedblright{} arises. In particular, these difficulties stem from the presence of the term
\begin{align*}
\frac{1}{n}\sum_{i=1}^n\epsilon_i(y)f_X(X_i)1_x(X_i)F_{Z\vert X}(z\vert X_i).
\end{align*}
Thanks to the projection-based principle proposed in previous sections, the following projected empirical process can be defined to solve the problems discussed above:
\begin{align*}
\hat I_n(y,w) = \frac{1}{n}\sum_{i=1}^n\hat{\epsilon}_i(y)\hat{f}_X(X_i)1_x(X_i)\hat{\mathcal{P}}_n1_z(Z_i),
\end{align*}
where $\hat{\epsilon}_i(y)=1_y(Y_i)-\hat{F}_{Y\vert X}(y\vert X_i)$ is the estimator for $\epsilon_i(y)$, with the leave-one-out nonparametric estimator
\begin{align*}
\hat{F}_{Y\vert X}(y\vert X_i) = \frac{1}{\hat{f}_X(X_i)}\frac{1}{(n-1)a^q}\sum_{j=1,j\neq i}^n K\left(\frac{X_i-X_j}{a}\right)1_y(Y_j).
\end{align*}
All other quantities are defined as before. The analysis of $\hat I_n(y,w)$ is identical to $\hat R_n(w)$ but with $\hat{\epsilon}_i$ substituted by $\hat{\epsilon}_i(y)$. Thus, reasoning as in the Section \ref{sec.Asy}, we have
\begin{align}\label{stat.CI}
\sup_{w}\left\vert \hat I_n(y,w)-\frac{1}{n}\sum_{i=1}^n\epsilon_i(y)f_X(X_i)1_x(X_i)\mathcal{P}1_z(Z_i)\right\vert = o_p\left(n^{-1/2}\right).
\end{align}
This leads to the fact that our constructed projected $U$-process $\sqrt n\hat I_n(\cdot)$ weakly converges to a centered Gaussian process $R_\infty^{CI}(\cdot)$ with covariance structure
\begin{align*}
\mathbb{E}\left[\epsilon(y)\epsilon(y^\prime)f^2_X(X)1_x(X)1_{x^\prime}(X)\mathcal{P}1_z(Z)\mathcal{P}1_{z^\prime}(Z)\right]
\end{align*}
at the parametric rate under the null. Correspondingly, under the alternative hypothesis, the constructed empirical process converges uniformly in probability to
\begin{align*}
C^{CI}(y,w) = \mathbb{E}\left\{\left[\mathbb{E}\left(1_y(Y)\vert W\right)-\mathbb{E}\left(1_y(Y)\vert X\right)\right]f_X(X)1_x(X)\mathcal{P}1_z(Z)\right\}.
\end{align*}
As a result, the test is consistent as long as for any constant $\gamma$, the following requirement
\begin{align}\label{stat.alt CI}
\mathbb{P}\left\{\left[\mathbb{E}\left(1_y(Y)\vert W\right)-\mathbb{E}\left(1_y(Y)\vert X\right)\right]=\gamma f_W(W)\right\}<1
\end{align}
holds. In addition, this condition also characterizes the nontrivial asymptotic local power of the test for conditional independence. The widely used CvM and KS test statistics can be constructed as
\begin{align*}
CvM^{CI}_n = \int\left\vert \hat I_n(y,w)\right\vert^2F_{Y_n,W_n}(dy,dw) \quad \text{and}\quad KS^{CI}_n = \sup_{(y,w)}\left\vert \hat I_n(y,w)\right\vert,
\end{align*}
so that large values provide evidence against the null in favor of the alternative hypothesis, where $F_{Y_n,W_n}(\cdot)$ is a random measure that converges in probability to $F_{Y,W}(\cdot)$ that is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{1+q+p}$. To implement the proposed test for conditional independence, still motivated by \eqref{stat.CI}, we construct the multiplier bootstrap version of $\hat I_n(y,w) $ as follows:
\begin{align*}
\hat I_n^\ast(y,w) = \frac{1}{n}\sum_{i=1}^nV_i\hat{\epsilon}_i(y)\hat{f}_X(X_i)1_x(X_i)\hat{\mathcal{P}}_n1_z(Z_i)
\end{align*}
and use it to obtain the critical values. The desired asymptotic properties follow from the following result:
\begin{align*}
\sup_{w}\left\vert \hat I^\ast_n(y,w)-\frac{1}{n}\sum_{i=1}^nV_i\epsilon_i(y)f_X(X_i)1_x(X_i)\mathcal{P}1_z(Z_i)\right\vert=o_p\left(n^{-1/2}\right).
\end{align*}
The corresponding multiplier bootstrap $CvM^{CI,\ast}_n$ and $KS^{CI,\ast}_n$ test statistics can be defined in a way analogous to that described in Section \ref{sec.boot}, and can be shown, under the null and under local alternatives, to provide good approximations to the sup norm and the squared norm of $R_\infty^{CI}(\cdot)$, respectively. Under fixed alternatives, the bootstrap versions of these statistics remain stochastically bounded, whereas the original statistics diverge to infinity whenever condition \eqref{stat.alt CI} is satisfied, thereby establishing the validity of the bootstrap procedure. The implementation of the resulting test can therefore follow the same steps as those outlined in Section \ref{sec.boot}. More Simulation results are presented in Section \ref{sec.AppendixA}.
\section{Simulation study}\label{sec.Simulation}
In this section, a Monte Carlo study is conducted to evaluate the finite-sample performance of the proposed nonparametric projection method. The main objectives are as follows. First, the proposed CvM test based on projection (denoted by PJ) is compared with the non-projected CvM test of \cite{delgado2001significance} (denoted by DM) in terms of empirical size and power.\footnote{For brevity, we have focused on comparing the CvM-type statistics. The comparison using the KS-type statistics is similar and available upon request.} Second, the robustness of the test performance with respect to bandwidth choice is examined under various bandwidth settings. Third, the behavior of the projection-based test is investigated across different covariate distributions and alternative frequencies.
In the simulation design presented in the main text, the covariates $X$ and $Z$ are both one-dimensional and follow a specified dependence structure $X=Z+U$, where $Z$ and $U$ are independent $N(0,1)$ variables. It is worth emphasizing that our procedures are not restricted to this case. The tests are applicable when $X$ and $Z$ are multivariate, and to illustrate this property, we also consider designs with $p=2$ and $q=2$. The proposed tests also apply when $X$ and $Z$ follow the uniform distribution $U(0,1)$. Moreover, when some parts of Assumptions \ref{ass.sample}--\ref{ass.bandwidth} are violated, for instance when both covariates are $N(0,1)$, the size accuracy of the test is affected but remains within an acceptable range and compares favorably with the DM test, which suggests that our procedure has a broader practical scope. For the sake of brevity, and because the main conclusions are similar, these additional results are reported in Appendix \ref{sec.AppendixA}, whereas the main text focuses on the case $X=Z+U$ with $Z$ and $U$ being independent $N(0,1)$.
To compare the performance of the tests under alternatives with different frequencies, the response variable $Y$ is generated under the following data-generating process (DGP),
\begin{align*}
Y = 1+X+\sin\left(\gamma Z\right)+\epsilon,
\end{align*}
where $\epsilon\sim N(0,1)$ and is independent of $(X,Z)$. Under the null hypothesis, the parameter $\gamma$ is fixed at zero, whereas under the alternative hypothesis, the frequency is determined by the value of $\gamma$. Moreover, to assess, in the spirit of \cite{bierens1997asymptotic}, the admissibility properties of ICM-type tests, we also report in Appendix \ref{sec.AppendixA} simulation results for designs with interaction effects between covariates. In particular, we consider alternatives of the form $X(Z^2-1)$ and examine the power of the proposed procedures in this case. In addition, the adoption of the kernel function and method of bandwidth selection follows those in \cite{delgado2001significance}, where an $l$-th order Epanechnikov kernel is employed, and the bandwidth is determined by the rule-of-thumb with coefficient $c\in\{0.5,1,2\}$ to assess the sensitivity to bandwidth. The order of the kernel and the explicit bandwidth formula are consistent with the requirements of Assumption \ref{ass.structual}. Specifically, $b=cn^{-1/3}$ is used for $l=2$, and $b=cn^{-1/6}$ is adopted for $l=4$. The results were examined under sample sizes $n\in\{200,400\}$ with the nominal significance level set at $\alpha\in\{0.01,0.05,0.1\}$, and all outcomes were obtained from $1000$ Monte Carlo replications and $199$ multiplier bootstrap samples.
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Significance test when $X=Z+U$, $Z\sim N(0,1)$ and $U\sim N(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.144 & 0.413 & 0.200 & 0.137 & 0.836 & 0.225\\
& 0.05 & 0.072 & 0.212 & 0.091 & 0.070 & 0.583 & 0.131\\
& 0.01 & 0.017 & 0.051 & 0.026 & 0.022 & 0.175 & 0.038\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.113 & 0.456 & 0.211 & 0.110 & 0.758 & 0.342\\
& 0.05 & 0.055 & 0.283 & 0.126 & 0.051 & 0.588 & 0.221\\
& 0.01 & 0.011 & 0.106 & 0.042 & 0.018 & 0.312 & 0.081\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.133 & 0.419 & 0.172 & 0.131 & 0.856 & 0.215\\
& 0.05 & 0.065 & 0.210 & 0.095 & 0.068 & 0.590 & 0.126\\
& 0.01 & 0.018 & 0.043 & 0.018 & 0.018 & 0.171 & 0.030\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.096 & 0.453 & 0.219 & 0.095 & 0.775 & 0.355\\
& 0.05 & 0.043 & 0.307 & 0.119 & 0.045 & 0.614 & 0.227\\
& 0.01 & 0.007 & 0.129 & 0.040 & 0.009 & 0.342 & 0.096\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.133 & 0.442 & 0.186 & 0.144 & 0.861 & 0.234\\
& 0.05 & 0.062 & 0.213 & 0.089 & 0.069 & 0.599 & 0.127\\
& 0.01 & 0.019 & 0.040 & 0.021 & 0.022 & 0.186 & 0.035\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.086 & 0.484 & 0.201 & 0.096 & 0.839 & 0.329\\
& 0.05 & 0.040 & 0.340 & 0.112 & 0.048 & 0.674 & 0.189\\
& 0.01 & 0.012 & 0.126 & 0.032 & 0.011 & 0.379 & 0.081\\
\bottomrule
\end{tabular}
\label{tab:SNsind1p1q1}
\end{table}
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Significance test when $X\sim N(0,1)$ and $Z\sim N(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.164 & 0.428 & 0.196 & 0.147 & 0.729 & 0.257\\
& 0.05 & 0.084 & 0.249 & 0.098 & 0.079 & 0.542 & 0.132\\
& 0.01 & 0.020 & 0.070 & 0.030 & 0.024 & 0.230 & 0.037\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.109 & 0.346 & 0.256 & 0.115 & 0.648 & 0.411\\
& 0.05 & 0.052 & 0.167 & 0.157 & 0.061 & 0.394 & 0.264\\
& 0.01 & 0.012 & 0.056 & 0.040 & 0.014 & 0.169 & 0.114\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.143 & 0.414 & 0.180 & 0.135 & 0.743 & 0.249\\
& 0.05 & 0.078 & 0.241 & 0.092 & 0.066 & 0.546 & 0.128\\
& 0.01 & 0.019 & 0.068 & 0.022 & 0.020 & 0.227 & 0.041\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.101 & 0.299 & 0.245 & 0.107 & 0.642 & 0.406\\
& 0.05 & 0.051 & 0.175 & 0.151 & 0.058 & 0.460 & 0.280\\
& 0.01 & 0.013 & 0.046 & 0.059 & 0.013 & 0.207 & 0.126\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.146 & 0.401 & 0.165 & 0.131 & 0.744 & 0.254\\
& 0.05 & 0.066 & 0.226 & 0.088 & 0.065 & 0.539 & 0.117\\
& 0.01 & 0.018 & 0.062 & 0.023 & 0.019 & 0.237 & 0.041\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.100 & 0.388 & 0.234 & 0.105 & 0.786 & 0.390\\
& 0.05 & 0.052 & 0.238 & 0.132 & 0.051 & 0.593 & 0.273\\
& 0.01 & 0.012 & 0.067 & 0.046 & 0.017 & 0.282 & 0.110\\
\bottomrule
\end{tabular}
\label{tab:SNsind2p1q1}
\end{table}
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Significance test when $X\sim U(0,1)$ and $Z\sim U(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.126 & 0.999 & 0.437 & 0.125 & 1.000 & 0.678\\
& 0.05 & 0.058 & 0.983 & 0.303 & 0.054 & 1.000 & 0.569\\
& 0.01 & 0.015 & 0.772 & 0.147 & 0.015 & 0.999 & 0.366\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.105 & 0.996 & 0.437 & 0.105 & 1.000 & 0.739\\
& 0.05 & 0.054 & 0.980 & 0.335 & 0.055 & 1.000 & 0.644\\
& 0.01 & 0.010 & 0.840 & 0.159 & 0.019 & 0.999 & 0.433\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.120 & 0.999 & 0.461 & 0.117 & 1.000 & 0.700\\
& 0.05 & 0.058 & 0.988 & 0.342 & 0.055 & 1.000 & 0.592\\
& 0.01 & 0.011 & 0.796 & 0.158 & 0.011 & 1.000 & 0.376\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.095 & 0.997 & 0.463 & 0.100 & 1.000 & 0.737\\
& 0.05 & 0.052 & 0.988 & 0.332 & 0.061 & 1.000 & 0.645\\
& 0.01 & 0.015 & 0.872 & 0.157 & 0.018 & 1.000 & 0.449\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.125 & 0.998 & 0.509 & 0.123 & 1.000 & 0.742\\
& 0.05 & 0.059 & 0.992 & 0.383 & 0.059 & 1.000 & 0.636\\
& 0.01 & 0.018 & 0.830 & 0.184 & 0.016 & 1.000 & 0.430\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.107 & 0.998 & 0.424 & 0.109 & 1.000 & 0.700\\
& 0.05 & 0.063 & 0.976 & 0.286 & 0.058 & 1.000 & 0.602\\
& 0.01 & 0.016 & 0.861 & 0.129 & 0.017 & 0.999 & 0.397\\
\bottomrule
\end{tabular}
\label{tab:SNsind3p1q1}
\end{table}
Tables \ref{tab:SNsind1p1q1}--\ref{tab:SNsind3p1q1} report results for the case $p=q=1$. Table \ref{tab:SNsind1p1q1} corresponds to the design $X=Z+U$, where $Z$ and $U$ are independent $N(0,1)$ variables, which satisfies Assumptions \ref{ass.sample}--\ref{ass.bandwidth} but not the bounded-support assumption on the distribution of $X$ imposed in \cite{delgado2001significance}. Table \ref{tab:SNsind2p1q1} considers the case where $X$ and $Z$ are independent $N(0,1)$ variables, and Table \ref{tab:SNsind3p1q1} the case where $X$ and $Z$ are independent $U(0,1)$ variables. As is typically observed, both size and power improve substantially as sample size increases, demonstrating that the proposed procedure possesses desirable large-sample properties.
For comparison with the method introduced in \cite{delgado2001significance} (named DM), the performance of each method in each case is reported in the tables. It is observed that DM tests exhibit lower accuracy, whereas the projection-based method achieves significant improvements. Regarding the power properties, Table \ref{tab:SNsind1p1q1} shows that the proposed projection-based tests (PJ) perform markedly better than the DM tests, which illustrates the suitability of our procedure for a wider class of unbounded distributions that satisfy Assumptions \ref{ass.sample}--\ref{ass.bandwidth}. In Table \ref{tab:SNsind2p1q1}, we investigate how the two tests behave when commonly used distributions that violate these assumptions are employed. We observe that the DM tests suffer from substantial size distortion, with empirical rejection frequencies well above the nominal level (oversized), whereas the PJ tests exhibit much better size accuracy. In Table \ref{tab:SNsind3p1q1}, both tests perform well, which is mainly due to the fact that the data are generated from the standard uniform distribution, a commonly used design that is fully compatible with the theoretical assumptions.
Further notable advantages of the proposed tests include reasonable robustness to bandwidth selection, as shown by results across different values of $c$, and relatively high power, as evidenced by comparisons across different values of $\gamma$. These features are characteristic of ICM-type tests in general and are naturally also reflected in the performance of the DM tests.
Finally, as further support for the theoretical results established in Section \ref{sec.CI}, we use the above DGPs to test the conditional independence assumption and the results are reported in Tables \ref{tab:CIsind1p1q1}--\ref{tab:CIsind3p1q1}. The only difference in the design is that
\begin{align*}
Y = 1+\left[X+\sin\left(\gamma Z\right)\right]\epsilon,
\end{align*}
where the conditional independence under test pertains only to the variance structure. Note that the significance tests discussed in Sections \ref{sec.Test}--\ref{sec.boot} can also be interpreted as tests of conditional independence for the conditional mean function, and the conclusions drawn from Tables \ref{tab:CIsind1p1q1}--\ref{tab:CIsind3p1q1} are similar to those obtained above.
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Testing conditional independence when $\Psi(X,Z)=\sin(\gamma Z)$, $X=Z+U$, $Z\sim N(0,1)$ and $U\sim N(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.200 & 0.137 & 0.152 & 0.169 & 0.162 & 0.133\\
& 0.05 & 0.105 & 0.078 & 0.096 & 0.079 & 0.088 & 0.083\\
& 0.01 & 0.032 & 0.023 & 0.033 & 0.022 & 0.025 & 0.025\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.122 & 0.119 & 0.129 & 0.100 & 0.140 & 0.107\\
& 0.05 & 0.071 & 0.066 & 0.076 & 0.048 & 0.069 & 0.061\\
& 0.01 & 0.023 & 0.018 & 0.022 & 0.012 & 0.021 & 0.019\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.161 & 0.128 & 0.126 & 0.170 & 0.158 & 0.119\\
& 0.05 & 0.084 & 0.062 & 0.073 & 0.076 & 0.087 & 0.070\\
& 0.01 & 0.024 & 0.016 & 0.024 & 0.018 & 0.025 & 0.020\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.102 & 0.098 & 0.103 & 0.086 & 0.117 & 0.093\\
& 0.05 & 0.049 & 0.045 & 0.060 & 0.048 & 0.059 & 0.047\\
& 0.01 & 0.013 & 0.012 & 0.020 & 0.010 & 0.025 & 0.015\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.197 & 0.121 & 0.122 & 0.197 & 0.138 & 0.123\\
& 0.05 & 0.100 & 0.055 & 0.067 & 0.096 & 0.082 & 0.062\\
& 0.01 & 0.026 & 0.013 & 0.020 & 0.020 & 0.025 & 0.019\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.092 & 0.096 & 0.099 & 0.084 & 0.102 & 0.093\\
& 0.05 & 0.044 & 0.045 & 0.050 & 0.052 & 0.057 & 0.051\\
& 0.01 & 0.008 & 0.006 & 0.010 & 0.012 & 0.018 & 0.018\\
\bottomrule
\end{tabular}
\label{tab:CIsind1p1q1}
\end{table}
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Testing conditional independence when $\Psi(X,Z)=\sin(\gamma Z)$, $X\sim N(0,1)$ and $Z\sim N(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.188 & 0.150 & 0.149 & 0.159 & 0.152 & 0.139\\
& 0.05 & 0.096 & 0.091 & 0.090 & 0.084 & 0.080 & 0.076\\
& 0.01 & 0.030 & 0.023 & 0.027 & 0.022 & 0.023 & 0.026\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.137 & 0.143 & 0.123 & 0.110 & 0.144 & 0.109\\
& 0.05 & 0.064 & 0.077 & 0.072 & 0.062 & 0.070 & 0.059\\
& 0.01 & 0.014 & 0.025 & 0.023 & 0.017 & 0.018 & 0.020\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.157 & 0.134 & 0.135 & 0.147 & 0.143 & 0.115\\
& 0.05 & 0.088 & 0.073 & 0.069 & 0.081 & 0.066 & 0.062\\
& 0.01 & 0.023 & 0.016 & 0.022 & 0.021 & 0.020 & 0.023\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.101 & 0.111 & 0.101 & 0.101 & 0.118 & 0.106\\
& 0.05 & 0.049 & 0.057 & 0.054 & 0.051 & 0.060 & 0.055\\
& 0.01 & 0.012 & 0.018 & 0.012 & 0.012 & 0.020 & 0.013\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.178 & 0.134 & 0.120 & 0.158 & 0.133 & 0.108\\
& 0.05 & 0.090 & 0.072 & 0.058 & 0.088 & 0.065 & 0.058\\
& 0.01 & 0.018 & 0.014 & 0.016 & 0.019 & 0.019 & 0.021\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.099 & 0.103 & 0.080 & 0.107 & 0.115 & 0.109\\
& 0.05 & 0.047 & 0.053 & 0.046 & 0.049 & 0.066 & 0.053\\
& 0.01 & 0.011 & 0.018 & 0.013 & 0.011 & 0.015 & 0.016\\
\bottomrule
\end{tabular}
\label{tab:CIsind2p1q1}
\end{table}
\begin{table}[htbp]
\centering
\setlength{\tabcolsep}{5pt}
\caption{Testing conditional independence when $\Psi(X,Z)=\sin(\gamma Z)$, $X\sim U(0,1)$ and $Z\sim U(0,1)$}
\begin{tabular}{llcccccccc}
\toprule
$c=0.5$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.184 & 0.954 & 0.275 & 0.180 & 1.000 & 0.506\\
& 0.05 & 0.095 & 0.882 & 0.147 & 0.095 & 0.999 & 0.269\\
& 0.01 & 0.034 & 0.563 & 0.042 & 0.028 & 0.957 & 0.073\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.102 & 0.863 & 0.281 & 0.104 & 0.997 & 0.592\\
& 0.05 & 0.050 & 0.764 & 0.156 & 0.056 & 0.992 & 0.400\\
& 0.01 & 0.010 & 0.521 & 0.057 & 0.012 & 0.955 & 0.173\\
\midrule
$c=1$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.189 & 0.959 & 0.278 & 0.199 & 1.000 & 0.547\\
& 0.05 & 0.107 & 0.902 & 0.148 & 0.110 & 0.999 & 0.282\\
& 0.01 & 0.030 & 0.623 & 0.044 & 0.030 & 0.969 & 0.078\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.086 & 0.910 & 0.285 & 0.107 & 0.999 & 0.637\\
& 0.05 & 0.052 & 0.828 & 0.173 & 0.058 & 0.996 & 0.440\\
& 0.01 & 0.013 & 0.575 & 0.063 & 0.015 & 0.968 & 0.200\\
\midrule
$c=2$ & & \multicolumn{3}{c}{\textbf{$n=200$}} & \multicolumn{3}{c}{\textbf{$n=400$}} \\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
& \textbf{$\alpha/\gamma$} & 0 & 5 & 10 & 0 & 5 & 10\\
\midrule
\multirow{3}{*}{DM}
& 0.10 & 0.363 & 0.964 & 0.305 & 0.374 & 1.000 & 0.594\\
& 0.05 & 0.196 & 0.916 & 0.167 & 0.211 & 1.000 & 0.305\\
& 0.01 & 0.059 & 0.697 & 0.045 & 0.066 & 0.983 & 0.081\\
\midrule
\multirow{3}{*}{PJ}
& 0.10 & 0.097 & 0.933 & 0.344 & 0.101 & 1.000 & 0.692\\
& 0.05 & 0.050 & 0.826 & 0.203 & 0.051 & 1.000 & 0.539\\
& 0.01 & 0.013 & 0.508 & 0.082 & 0.012 & 0.961 & 0.244\\
\bottomrule
\end{tabular}
\label{tab:CIsind3p1q1}
\end{table}
\section{Conclusion}\label{sec.Conclusion}
This paper introduces a novel nonparametric significance test that uses a tailored projected weighting function to address the nonparametric estimation effect.
The approach relaxes the restrictive compact support assumption, provides a computationally efficient multiplier bootstrap, and extends straightforwardly to testing conditional independence. Numerical evidence confirms the strong performance of our proposal in finite samples.
\bibliographystyle{apalike_revised}
\bibliography{Ref}
\newpage