The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
84,221 characters
Specification tests for regression models with measurement errors
\title{Specification tests for regression models with measurement errors\footnote{The authors contributed equally to this work and are listed in alphabetical order.}}
\author{Xiaojun Song\thanks{Guanghua School of Management, Peking University. Email: \texttt{[email removed]}. This work was supported by the National Natural Science Foundation of China [Grant Numbers 72373007 and 72333001]. The author also gratefully acknowledges the research support from the Center for Statistical Science of Peking University.}\\
\and Jichao Yuan\thanks{Guanghua School of Management, Peking University. Email: \texttt{[email removed]}.}
}
\maketitle
\begin{abstract}
In this paper, we propose new specification tests for regression models with measurement errors in the explanatory variables. Inspired by the integrated conditional moment (ICM) approach, we use a deconvoluted residual-marked empirical process and construct ICM-type test statistics based on it. The issue of measurement errors is addressed by applying a deconvolution kernel estimator in constructing the residuals. We demonstrate that employing an orthogonal projection onto the tangent space of nuisance parameters not only eliminates the parameter estimation effect but also facilitates the simulation of critical values via a computationally simple
multiplier bootstrap procedure.
It is the first time a multiplier bootstrap has been proposed in the literature of specification testing with measurement errors. We also develop specification tests and the multiplier bootstrap procedure when the measurement error distribution is unknown. The finite-sample performance of the proposed tests for both known and unknown measurement error distributions is evaluated through Monte Carlo simulations, which demonstrate their efficacy.
\end{abstract}
\newpage
\doublespacing
\section{Introduction}
There already exists a substantial body of literature on specification testing for regression models, see, e.g., \cite{bierens1982consistent}, \cite{hardle1993comparing}, \cite{JOHNXUZHENG1996263}, \cite{stute1997nonparametric}, \cite{stute1998bootstrap}, \cite{stinchcombe1998consistent}, \cite{stute2002model}, \cite{zhu2003testing}, \cite{zhu2005testing}, \cite{Escanciano_2006}, \cite{hall2007testing}, \cite{song2008model}, \cite{xu2015nonparametric}, \cite{sant2019specification}, \cite{otsu2021specification}, and \cite{TAN2025106113}. The list is undoubtedly not exhaustive.
However, among these, relatively few studies have adequately developed specification tests in the presence of measurement errors. Measurement error is a common issue in data across many disciplines, including economics, finance, and medicine, see, e.g., \cite{fuller2009measurement}, \cite{alexander2009deconvolution}, and \cite{hu2017measurement}. Unfortunately, specification tests often have incorrect size and low power in the presence of measurement errors.
In an earlier influential work, \cite{otsu2021specification} proposed a specification test for regression models with measurement errors in the explanatory variables based on nonparametric local smoothing estimators. Their test belongs to the class of local smoothing tests, and so does the minimum distance test proposed by \cite{song2008model}. These tests are typically constructed from the distance between the fitted nonparametric regression and the parametric fit under the null hypothesis, and have been shown to be more powerful against high-frequency alternatives. Complementary to such approaches, global smoothing tests, developed in \cite{hall2007testing}, exhibit higher power against low-frequency alternatives. As emphasized in \cite{otsu2021specification}, local smoothing and global smoothing tests serve as complements rather than substitutes. A similar complementarity in power properties between the two types of tests, in the absence of measurement errors, was also established in \cite{fan2000consistent}.
Within the class of global smoothing tests, the integrated conditional moment (ICM)-type tests proposed by \cite{bierens1982consistent} and \cite{bierens1997asymptotic} are particularly important and are the main focus of this paper. These tests typically compare the integrated regression function with the integrated parametric regression function under the null hypothesis. Specifically, a residual-based empirical process is employed to construct the test statistics, and, in general, the asymptotic null distributions depend on the data-generating process, thereby motivating the use of bootstrap techniques to obtain the critical values. The ICM-type tests have been widely applied to specification testing when measurement errors are absent, as shown in \cite{su1991lack}, \cite{stute1997nonparametric}, and \cite{Escanciano_2006}, among others. However, in the presence of measurement errors, the ICM-type tests face challenges arising from the unobservability of the true residuals, the stringent requirements for the estimators, and the computational complexity of obtaining critical values, which motivate the present study to address these issues. In this paper, we propose ICM-type tests based on a deconvoluted residual-marked empirical process (i.e., the residuals are constructed using the deconvolution kernel estimator). We employ an orthogonal projection onto the tangent space of nuisance parameters to eliminate the parameter estimation effect. As a desirable byproduct, the projection facilitates the simulation of critical values via a computationally straightforward multiplier bootstrap.
It is important to highlight our contribution regarding the simple multiplier bootstrap procedure for obtaining critical values. Conventional bootstrap methods face serious challenges in the presence of measurement errors because the true regressors are unobservable; consequently, we cannot resample them. Infeasibility of obtaining the critical values through the multiplier bootstrap procedure is due to the parameter estimation effect,
initially discussed in \cite{durbin1973distribution}. We introduce a novel projection to address the parameter estimation effect by imposing an orthogonality condition on the weight function of the residual-marked empirical process. The introduction of the projection also facilitates a computationally attractive multiplier bootstrap procedure to implement our tests, whose asymptotic validity can be easily justified. Notably, compared with other bootstrap methods, the multiplier bootstrap is simpler to justify theoretically, easier to implement, and computationally more efficient. To our knowledge, this is the first time a multiplier bootstrap has been proposed in the literature of specification testing in the presence of measurement errors. Furthermore, we establish a comprehensive theoretical framework and a multiplier bootstrap-based procedure for obtaining critical values when the measurement error distribution is unknown. In addition, our tests are more robust to bandwidth choices than the local smoothing approach.
The rest of this paper is organized as follows. In Section \ref{sec.Test}, we describe the testing framework and introduce our projection method. Section \ref{sec.Asy} discusses the asymptotic properties of the test statistics. Section \ref{sec.unknown ME} addresses the case of unknown measurement error distribution. In Section \ref{sec.boot}, we provide a multiplier bootstrap procedure and its implementation. Monte Carlo simulations are conducted in Section \ref{sec.Simulation}. We present our conclusions in Section \ref{sec.Conclusion}. Additional simulation results and the mathematical proofs are included in the online supplementary appendix.
\section{The testing framework}\label{sec.Test}
\subsection{The testing problem}
Let $(Y,X)^\top$ be a random vector in a two-dimensional Euclidean space, where $Y$ is an observed real-valued response variable and $X$ is an unobservable error-free explanatory variable. Consider the following regression model:
\begin{align}\label{mod.regression model}
Y=m(X)+U, \quad \text{with} \quad\mathbb{E}[U|X]=0\,\,\,\text{almost surely }(a.s.),
\end{align}
where $m(X)=\mathbb{E}[Y|X]$ is the regression function of $Y$ given $X$ and $U$ is the unpredictable part of $Y$. While the true explanatory variable $X$ is not directly observed, we instead observe a noisy variable $W$ through
\begin{align}\label{ME}
W=X+\epsilon, \quad \text{with} \quad \mathbb E[\epsilon]=0,
\end{align}
where the measurement error $\epsilon\in\mathbb{R}$ is assumed to be independent of $X$ and $Y$. In addition, we assume that the density function of $\epsilon$ is known and denoted by $f_\epsilon$. The case of unknown $f_\epsilon$ but with repeated measurements available is discussed in Section \ref{sec.unknown ME}. Suppose researchers consider using the following parametric regression model:
\begin{align}\label{mod.parametric regression model}
Y=g(X;\theta)+e(\theta),
\end{align}
where $g(X;\theta)$ is the parametric specification for the regression function $m(X)$ and $e(\theta)$ is the parametric disturbance of the model. In both econometrics and statistics, it is crucial to assess the adequacy of the putative parametric model $g(X;\theta)$. That is, our null hypothesis of interest is
\begin{align}\label{hyp.null1}
H_0: \mathbb{P}[m(X)=g(X;\theta_0)]=1 \,\,\,\text{for some}\,\,\,\theta_0\in\Theta\subset\mathbb{R}^d.
\end{align}
The alternative hypothesis is the negation of $H_0$, i.e.,
\begin{align}\label{hyp.alternative1}
H_1: \mathbb{P}[m(X)=g(X;\theta)]<1 \,\,\,\text{for any}\,\,\,\theta\in\Theta\subset\mathbb{R}^d.
\end{align}
Clearly, testing the null hypothesis in \eqref{hyp.null1} is equivalent to testing
\begin{align}\label{hyp.null2}
H_0: \mathbb{E}[e(\theta_0)|X]=0\,\,\,a.s.\,\,\text{for some}\,\,\,\theta_0\in\Theta\subset\mathbb{R}^d.
\end{align}
Note that \eqref{hyp.null2} is a standard conditional moment restriction. It is well known that \eqref{hyp.null2} can be equivalently expressed as a continuum number of unconditional moment restrictions using the exponential weight function, see, for example, \cite{bierens1982consistent}, \cite{bierens1990consistent}, and \cite{bierens1997asymptotic}. Specifically, we can rewrite \eqref{hyp.null2} as follows:
\begin{align}\label{null eq1}
S(\xi,\theta_0)=\mathbb{E}\left[e(\theta_0){\rm e}^{{\rm i} X\xi}\right]=0\,\,\,\forall\,\xi\in\Pi\subset\mathbb{R}\,\,\,\text{for some}\,\,\,\theta_0\in\Theta\subset\mathbb{R}^d,
\end{align}
where $e(\theta)=Y-g(X;\theta)$ is the parametric error, $\Pi$ is a properly chosen compact set with nonempty interior, and ${\rm i}=\sqrt{-1}$ denotes the imaginary unit.
To provide evidence of model misspecification, we can then compare a suitable estimator for $S(\xi,\theta_0)$ in \eqref{null eq1} with the zero function. Unfortunately, the classical residual-marked empirical process is clearly infeasible, as the true regressor $X_i$ is unobservable due to the measurement error $\epsilon$ in \eqref{ME}.
Motivated by \cite{dong2022nonparametric}, however, $S(\xi,\theta_0)$ can still be consistently estimated by noting that
\begin{align*}
S(\xi,\theta_0)=\iint\left(y-g(x;\theta_0)\right)f_{Y,X}(y,x){\rm e}^{{\rm i} x\xi}\,dy\,dx,
\end{align*}
where $f_{Y,X}(y,x)$ denotes the unknown joint density function of $\left(Y,X\right)^\top$. Given a random sample $\{(Y_i,W_i)^\top\}_{i=1}^n$ of size $n\geq 1$, and motivated by the deconvolution methods developed for density estimation [see, e.g., \cite{carroll1988optimal} and \cite{stefanski1990deconvolving}], we replace $f_{Y,X}(y,x)$ by a nonparametric estimator and employ a consistent estimator $\hat{\theta}_n$ for $\theta_0$. Then, $S(\xi,\theta_0)$ is estimated by
\begin{align*}
S_{n}(\xi,\hat{\theta}_n)=\iint\left(y-g(x;\hat{\theta}_n)\right)\hat{f}_{Y,X}(y,x){\rm e}^{{\rm i} x\xi}\,dy\,dx,
\end{align*}
where $\hat f_{Y,X}(y,x)$ is the deconvolution kernel density estimator for $f_{Y,X}(y,x)$, i.e.,
\begin{align*}
\hat{f}_{Y,X}(y,x)=\frac{1}{n}\sum_{i=1}^nK_b\left(\frac{y-Y_i}{b}\right)\mathcal{K}_b\left(\frac{x-W_i}{b}\right),
\end{align*}
with $K_b(a)=K(a)/b$, where $K(\cdot)$ is a symmetric kernel function and $b=b_n\in\mathbb{R}^+$ is a sequence of bandwidth parameters shrinking to zero at an appropriate rate specified later. In addition, $\mathcal{K}_b(a)$ is the univariate deconvolution kernel given by
\begin{align}\label{conv}
\mathcal{K}_b(a)=\frac{1}{2\pi b}\int {\rm e}^{-{\rm i} t a}\frac{K^{\text{ft}}(t)}{f_\epsilon^{\text{ft}}(t/b)}\,dt,
\end{align}
in which $K^{\text{ft}}(\cdot)$ and $f_\epsilon^{\text{ft}}(\cdot)$ are the Fourier transforms of the kernel function $K(\cdot)$ and the measurement error density $f_\epsilon(\cdot)$, respectively. The above deconvolution kernel density estimator has been widely used to construct density and nonparametric regression estimators in the presence of measurement errors; see, e.g., \cite{alexander2009deconvolution}. Such an approach is particularly important when kernel smoothers are employed, as shown in \cite{otsu2021specification}. We note that, following this approach, $S_{n}(\xi,\hat{\theta}_n)$ can be further rewritten as
\begin{align}
S_{n}(\xi,\hat{\theta}_n)=&\frac{1}{n}\sum_{i=1}^n\int\left[\int\left(y-g(x;\hat{\theta}_n)\right)K_b\left(\frac{y-Y_i}{b}\right)\,dy\right]\mathcal{K}_b\left(\frac{x-W_i}{b}\right){\rm e}^{{\rm i} x\xi}\,dx\notag\\
=&\frac{1}{n}\sum_{i=1}^n\int\left(Y_i-g(x;\hat{\theta}_n)\right)\mathcal{K}_b\left(\frac{x-W_i}{b}\right){\rm e}^{{\rm i} x\xi}\,dx. \label{interm}
\end{align}
It can be shown that under the null hypothesis and mild regularity conditions on the smoothness of $g(x;\theta)$ with respect to $\theta$, the smoothness of $f_\epsilon$ and kernel function $K(\cdot)$, the $\sqrt n$-consistency assumption on the estimator $\hat\theta_n$ as well as the bandwidth $b$,
\begin{align}
\sqrt{n}S_{n}(\xi,\hat{\theta}_n)=&\frac{1}{\sqrt{n}}\sum_{i=1}^n\int\left(Y_i-g(x;\theta_0)\right)\mathcal{K}_b\left(\frac{x-W_i}{b}\right){\rm e}^{{\rm i} x\xi}\,dx\notag\\
&-\sqrt{n}\left(\hat{\theta}-\theta_0\right)^\top G(\xi,\theta_0)+o_p(1), \label{no_proj_decom}
\end{align}
uniformly in $\xi\in\Pi$, where,
\begin{align}\label{classical_parametric effect}
G(\xi,\theta_0)=\mathbb{E}\left[\dot{g}(X;\theta_0){\rm e}^{{\rm i} X\xi}\right]:=\int\dot{g}(x;\theta_0){\rm e}^{{\rm i} x\xi}f_X(x)\,dx,
\end{align}
with $f_X(x)$ denoting the density function of $X$. Here and below, a dot denotes differentiation with respect to the variable $\theta$, i.e., $\dot{g}(x;\theta)=\partial g(x;\theta)/\partial\theta$.
However, the uniform asymptotic decomposition obtained in \eqref{no_proj_decom} requires the $\sqrt n$-consistency property of $\hat\theta_n$, which may not be satisfied in the measurement error context. It is known that $\hat\theta_n$ may have a lower convergence rate, for example, $\hat\theta_n-\theta_0=O_p(n^{\delta-1/2})$ for some $0\leq\delta<1/4$ in parametric models with measurement errors, see \cite{taupin1998estimation}. Even $\hat\theta_n$ is $\sqrt n$-consistent, due to the presence of $\sqrt{n}(\hat{\theta}_n-\theta_0)^\top G(\xi,\theta_0)$ (commonly known as the parameter estimation effect) in the decomposition of $\sqrt{n}S_n(\xi,\hat{\theta}_n)$, the asymptotic null distribution of $\sqrt{n}S_n(\xi,\hat{\theta}_n)$ depends on $\hat\theta_n$ and the computationally attractive multiplier bootstrap is also infeasible. A typical way to deal with the parameter estimation effect is to impose the asymptotically linear representation assumption of $\sqrt n(\hat\theta_n-\theta_0)$ as follows:
\begin{align}\label{param expansion}
\sqrt{n}\left(\hat{\theta}_n-\theta_0\right)=\frac{1}{\sqrt n}\sum_{i=1}^nl(Y_i,W_i;\theta_0)+o_p(1),
\end{align}
for some function $l(Y_i,W_i;\theta_0)$ with zero mean and finite variance. However, such an asymptotically linear representation is
even more difficult to obtain in the context of measurement errors.
\subsection{A projection approach}
In this section, we propose a class of orthogonal projection-based test statistics with attractive theoretical and empirical properties. This approach has been used in the literature to address parameter estimation effects in various testing problems; for example, specification analysis of linear quantile models in \cite{escanciano2014specification}, specification tests for the propensity score in \cite{sant2019specification}, model checking in partially linear spatial autoregressive models in \cite{yang2024model}, and testing for the nonparametric component in
partially linear quantile regression models in \cite{song2025unified}. It is worth emphasizing that the projection approach proposed here is particularly suitable for regression models with measurement errors, as it does not require assuming an asymptotically linear representation of $\sqrt n(\hat\theta_n-\theta_0)$, or even the $\sqrt n$-consistency of $\hat\theta_n$ to $\theta_0$. Specifically, we introduce a projection-based weight to eliminate the parameter estimation effect in \eqref{no_proj_decom}, thereby relaxing the requirement on the estimator $\hat\theta_n$ and preserving the parametric convergence advantage of our ICM-type tests relative to local-smoothing-based methods. To be precise, $n^{\delta}(\hat\theta_n-\theta_0)$ for some $\delta>1/4$ would suffice. Consequently, a computationally attractive multiplier bootstrap procedure is facilitated, which is particularly advantageous in the presence of measurement errors.
To motivate our approach, we consider a projection-based transformation of the weight function ${\rm e}^{{\rm i} x\xi}$ given by
\begin{align}\label{projection distribution}
\mathcal{P}(x;\xi,\theta)={\rm e}^{{\rm i} x\xi}-\dot g^\top(x;\theta)\Delta^{-1}(\theta)G(\xi,\theta),
\end{align}
with
\begin{align*}
G(\xi,\theta )=\mathbb{E}[\dot g(X;\theta ){\rm e}^{{\rm i} X\xi}]\quad\text{and}\quad\Delta(\theta) =\mathbb{E}[\dot g(X;\theta)\dot g^\top(X;\theta)].
\end{align*}
A similar modification of the weight function can also be found in \cite{escanciano2014specification} and \cite{sant2019specification}, among others. The intuition behind \eqref{projection distribution} is that $\Delta^{-1} \left( \theta \right)G\left(\xi,\theta \right) $ represents the vector of linear projection coefficients of regressing ${\rm e}^{{\rm i} X\xi} $ on $\dot g(X;\theta)$. Thus, it follows that $\dot g(X;\theta )^\top\Delta^{-1} \left( \theta \right) G\left(\xi,\theta
\right) $ is the best linear predictor of ${\rm e}^{{\rm i} X\xi} $ given $\dot g(X;\theta )$ and
\begin{align*}
\mathbb{E}\left[\dot g(X;\theta )\mathcal{P}(X;\xi,\theta)\right] & =\mathbb{E}\left[\dot g(X;\theta )\left({\rm e}^{{\rm i} X\xi}-\dot g^\top (X,\theta
)\Delta^{-1} \left( \theta \right) G\left(\xi,\theta \right)
\right)\right] \\
& =G(\xi,\theta )-\Delta \left( \theta \right) \Delta^{-1} \left( \theta
\right) G\left(\xi,\theta \right) =0.
\end{align*}
Based on the properties mentioned above, our tests are then based on continuous functionals of the following feasible projection-based deconvoluted residual-marked empirical process,
\begin{align}\label{Test stat}
S_{n}^{pro}(\xi,\hat\theta_n)=\frac{1}{n}\sum_{i=1}^n\int\left(Y_i-g(x;\hat{\theta}_n)\right)\mathcal{K}_b\left(\frac{x-W_i}{b}\right)\mathcal{P}_n(x;\xi,\hat{\theta}_n)\,dx,
\end{align}
where $\hat{\theta}_n$ is a suitably consistent estimator for $\theta_{0}$ under the null hypothesis with required convergence rates specified later (not necessarily $\sqrt n$-consistent), and $\mathcal{P}_n(x;\xi,\theta)$ is the sample analog of projection $\mathcal{P}(x;\xi,\theta)$ in (\ref{projection distribution}); namely,
\begin{align*}
\mathcal{P}_n(x;\xi,\hat{\theta}_n)={\rm e}^{{\rm i} x\xi}
-\dot g^\top(x;\hat\theta_n)\Delta_n^{-1}(\hat\theta_n)G_{n}(\xi,\hat\theta_n),
\end{align*}
where
\begin{equation*}
G_{n}(\xi,\theta)=\int\dot g(x;\theta)\hat f_X(x){\rm e}^{{\rm i} x\xi}\,dx \quad\text{and} \quad \Delta_n(\theta)=\int\dot g(x;\theta)\dot g^\top(x;\theta )\hat f_X(x)\,dx,
\end{equation*}
with
\begin{equation}
\hat f_X(x)=\frac{1}{n}\sum_{i=1}^n\mathcal{K}_b\left(\frac{x-W_i}{b}\right)\label{decon density}
\end{equation}
the deconvolution kernel density estimator for $f_X(x)$, where $\mathcal{K}_b(\cdot)$ is defined in \eqref{conv}.
With the assistance of \eqref{Test stat}, we can establish that under the null hypothesis of correct specification, the empirical process $S_{n}^{pro}(\cdot,\hat{\theta}_n)$ is expected to be close to the zero function. Under the alternative hypothesis of misspecification, $S_{n}^{pro}(\cdot,\hat{\theta}_n)$ tends to deviate from the zero function, and consequently $\sqrt nS_{n}^{pro}(\cdot,\hat{\theta}_n)$ diverges. It is therefore natural to construct the test statistic by measuring an appropriate distance between $S^{pro}_{n}(\cdot,\hat{\theta}_n)$ and the zero function, denoted by $\Gamma(S^{pro}_{n})$ for some continuous functional $\Gamma(\cdot)$. For example, we could consider the following two test statistics based on the popular Kolmogorov--Smirnov (KS) and Cram\'{e}r--von Mises (CvM) functionals, respectively,
\begin{align*}
KS_{n}=\sup_{\xi\in\Pi}\left\vert S^{pro}_{n}(\xi,\hat{\theta}_n)\right\vert
\quad\text{and}\quad CvM_{n}=\int_\Pi\left\vert S^{pro}_{n}(\xi,\hat{\theta}_n)\right\vert^2\,d\xi.
\end{align*}
In constructing the $CvM_{n}$ statistic, we adopt the uniform integrating measure on $\Pi$, which is also recommended in \cite{dong2022nonparametric}. Previous studies have discussed the use of integrating measures that are absolutely continuous with respect to the Lebesgue measure on $\Pi$. However, due to the difficulty of selecting an appropriate measure and the estimation challenges that arise when the true regressor $X$ is unobserved, such measures have rarely been employed in models with measurement errors. The test statistics $KS_{n}$ and $CvM_{n}$ should be small if the null hypothesis \eqref{hyp.null1} is true, while \textquotedblleft
large\textquotedblright{}\ values of $KS_{n}$ and $CvM_{n}$ imply the rejection of $H_{0}$. These \textquotedblleft
large\textquotedblright{}\ values will be determined by a convenient multiplier bootstrap procedure described in Section \ref{sec.boot}, thanks to the projection $\mathcal{P}_n(x;\xi,\hat{\theta}_n)$ that eliminates the parameter estimation effect.
\section{Asymptotic theory}\label{sec.Asy}
In this section, we establish the asymptotic distributions of the test statistics $KS_n$ and $CvM_n$ under the null hypothesis $H_0$. Subsequently, we investigate the asymptotic power behavior of these statistics under a sequence of local alternative hypotheses $H_{1n}$ that converge to $H_0$ at the parametric rate $n^{-1/2}$, as well as under the fixed alternative. Denote $f^{ft}(t)=\int {\rm e}^{{\rm i} tx}f(x)dx$ for a generic function $f$ and let $g^{(p)}(x;\theta)$ denote the $p$-times derivative of function $g$ with respect to variable $x$. Throughout this paper, $|c|$ is used to denote the Euclidean norm of a vector $c$ and $|A|$ is used to denote the Frobenius norm of a matrix $A$.
\subsection{Asymptotic null distributions}
We first impose some regularity conditions to derive the asymptotic null distributions of the test statistics $KS_n$ and $CvM_n$.
\begin{assumption}\label{ass.D}
\quad
\begin{enumerate}[label=(\roman*)]
\item $\{(Y_i,W_i)^\top\}_{i=1}^n$ is an i.i.d sample of $(Y,W)^\top$ which satisfies \eqref{mod.regression model}, \eqref{ME}, and $\mathbb E|Y|^2<\infty$.
\item The estimator $\hat\theta_n$ satisfies $\hat\theta_n-\theta_0=O_p(n^{\delta-1/2})$ for some $0\leq \delta<1/4$ in parametric regression models with measurement errors.
\item The measurement error $\epsilon$ is independent of $(Y,X)^\top$.
\end{enumerate}
\end{assumption}
Assumption \ref{ass.D}({\romannumeral1}) requires random sampling and the existence of the second moment of $Y$, both of which are frequently mentioned in the literature. We emphasize that Assumption \ref{ass.D}({\romannumeral2}) relaxes the requirement on the convergence rate of the estimator $\hat\theta_n$ used in our tests, making it more suitable for cases involving measurement errors. Most estimators cannot typically be expanded as \eqref{param expansion} detailed in Section \ref{sec.Test} in the presence of measurement errors. Even worse, they cannot achieve $\sqrt{n}$-rate convergence without imposing numerous restrictions [see \cite{taupin1998estimation}], making the multiplier bootstrap infeasible. Assumption \ref{ass.D}({\romannumeral3}) is common and essential, as mentioned in the classical literature on measurement error [see \cite{otsu2021specification}].
As shown in the classical measurement error literature, we categorize two separate cases by the decay rate of the tail of the characteristic function of the measurement errors: the ordinary smooth case and the supersmooth case. We first focus on the ordinary smooth case and impose the following assumptions.
\begin{assumption}\label{ass.O}
\quad
\begin{enumerate}[label=(\roman*)]
\item The functions $f_X(x)$, $g(x;\theta_0)$, and $h(x;\theta_0)$ are $p$-times continuously differentiable with respect to $x$ with bounded and integrable derivatives, where $h(x;\theta)$ denotes any partial derivative of $g(x;\theta)$ with respect to $\theta$ of order up to three and $p$ is a positive integer satisfying $p>\alpha$. Furthermore, we also impose additional assumption about Lipschitz continuous properties of $f_X^{(p)}(x)$, $\left[g(x;\theta_0)f_X(x)\right]^{(p)}$, $h^{(p)}(x;\theta_0)$, and $\left[g(x;\theta_0)h(x;\theta_0)\right]^{(p)}$ for almost every $x$ as follows:
\begin{align*}
& \left\vert f_X^{(p)}(x+y) - f_X^{(p)}(x)\right\vert \leq L_{f_X^{(p)}}(x)\vert y\vert,\\
& \left\vert \left[g(x+y;\theta_0)f_X(x+y)\right]^{(p)} - \left[g(x;\theta_0)f_X(x)\right]^{(p)}\right\vert \leq L_{[g f_X]^{(p)}}(x)\vert y\vert\\
& \left\vert h^{(p)}(x+y;\tilde{\theta}) - h^{(p)}(x;\tilde{\theta})\right\vert \leq L_{h^{(p)}}(x)\vert y\vert,\\
& \left\vert \left[g h\right]^{(p)}(x+y;\tilde{\theta}) - \left[g h\right]^{(p)}(x;\tilde{\theta})\right\vert \leq L_{[g h]^{(p)}}(x)\vert y\vert,
\end{align*}
for some bounded and integrable functions $L_{f_X^{(p)}}(x)$, $L_{[gf_X]^{(p)}}(x)$, $L_{h^{(p)}}(x)$, and $L_{[gh]^{(p)}}(x)$, where $\tilde{\theta}$ takes values between $\theta_0$ and $\hat{\theta}_n$.
\item The characteristic function of measurement error $\epsilon$ is of the following form for all $t\in\mathbb{R}$, where $c_0^{os},c_1^{os},\cdots,c_{\alpha}^{os}$ are finite constants with $c_0^{os} = 1$ and $\alpha>0$,
\begin{align*}
f_\epsilon^{ft}(t) = \frac{1}{c_0^{os}+c_{1}^{os}t+\cdots+c_{\alpha}^{os}t^{\alpha}}.
\end{align*}
\item The kernel function $K$ is differentiable to order $p+1$ and satisfies the following equations for $l=1,2,\cdots,p-1$:
\begin{align*}
&\int K(u)du = 1, \qquad \int u^{p}K(u)du\neq 0, \qquad \int u^lK(u)du = 0.
\end{align*}
In addition, $K^{ft}$ is compactly supported on a compact set that is symmetric around zero.
\item $nb^{2p}\to 0$ as $n\to\infty$.
\item For $c_l^{os}(\xi) = (-{\rm i})^l\sum\limits_{j=l}^{\alpha}c_j^{os}\binom{j}{l}\xi^{j-l}$, we have
\begin{align*}
&\mathbb{E}\left[\sup\limits_{\xi\in\Pi}\left\vert r^{os}_{h, \infty}(Y,W;\xi,\theta_0)\right\vert^2\right]<\infty,
\end{align*}
where
\begin{align*}
&r^{os}_{h, \infty}(Y,W;\xi,\theta_0) = \sum_{l=0}^\alpha c_l^{os}(\xi) \left[Yh^{(l)}(W;\theta_0)\right] - \sum_{l=0}^\alpha c_l^{os}(\xi) \left[(gh)^{(l)}(W;\theta_0)\right].
\end{align*}
\end{enumerate}
\end{assumption}
Assumption \ref{ass.O}({\romannumeral1}) imposes both requirements of the Lipschitz continuity and restrictions on the smoothness of the structural functions, mainly adopted from \cite{dong2022nonparametric}. The Lipschitz continuity assumptions are necessary to analyze the moment properties of structural functions under ordinary smooth conditions, as detailed in the proofs of the lemmas. Our assumptions are slightly stronger, requiring the Lipschitz continuity of $h^{(p)}$ and $[gh]^{(p)}$, enhancements that are essential for ensuring convergence. These assumptions hold in commonly used scenarios, such as polynomial $g$, uniformly distributed $X$, normally distributed $U$, and Laplace distributed $\epsilon$. Assumption \ref{ass.O}({\romannumeral2}) is commonly known as the ordinary smooth assumption, which specifies the polynomial decay rate of $f_\epsilon^{ft}$, the characteristic function of measurement error. As shown in \cite{fan1995average} and \cite{dong2022nonparametric}, it can be generalized to $f_\epsilon^{ft}(t)=\exp(it\zeta)/(c_0^{os}+c_{1}^{os}t+\cdots+c_{\alpha}^{os}t^{\alpha})$, including typical distributions such as the Laplace distribution. The characteristic function of the measurement error is estimated through repeated sampling when it is unknown, as outlined in Section \ref{sec.unknown ME}. Assumption \ref{ass.O}({\romannumeral3}) concerns the order of the kernel function, which is essential for constructing the deconvolution kernel, ensuring its integration properties, and eliminating the asymptotic bias introduced by nonparametric estimators. Moreover, the construction of higher-order kernel functions is feasible as discussed in \cite{alexander2009deconvolution}. Assumption \ref{ass.O}({\romannumeral4}) contains the undersmoothing condition frequently used in the literature. It also guarantees the asymptotic negligibility of the bias from nonparametric estimators and establishes the integration properties of the deconvolution kernel. It is worth noting that, in the context of the specification testing problem considered in this paper, we do not impose a lower bound on the bandwidth to ensure the existence of the variance, as is commonly done in studies using kernel-based estimators. In particular, we verify in Appendix \ref{sec.AppendixC} that, under the Lipschitz continuity of the underlying functions and an undersmoothing bandwidth condition, the use of kernel smoothing does not introduce any additional non-negligible variance. Consequently, it suffices to assume the boundedness of the asymptotic variance to guarantee the existence of the limiting process, as stated in Assumption \ref{ass.O}({\romannumeral5}). Similar assumptions are made in \cite{fan1995average} and \cite{dong2022nonparametric}.
Using the assumptions above, we can characterize the limiting behavior of $S_{n}^{pro}(\cdot,\hat\theta_n)$ for the ordinary smooth case. Let \textquotedblleft $\Longrightarrow$ \textquotedblright{} denote weak convergence on $(l^{\infty}(\Pi),\mathcal{B}_{\infty})$ in the sense of Hoffmann--J{\o}rgensen, where $\mathcal{B}_{\infty}$ denotes the corresponding Borel $\sigma$-algebra, see, e.g., Definition 1.3.3 in \cite{van1996weak}. Subsequently, we derive the asymptotic null distributions of the statistics $KS_n$ and $CvM_n$, as stated in the following theorem and corollary.
\begin{theorem}\label{theorem.ordinary smooth under H_0}
Suppose that Assumptions \ref{ass.D} and \ref{ass.O} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align}\label{result of param effect ass O H0}
&\sup_{\xi\in\Pi}\left\vert S_{n}^{pro}(\xi,\hat\theta_n)- S_{n}^{pro}(\xi,\theta_0)\right\vert=o_p\left(n^{-\frac{1}{2}}\right).
\end{align}
Furthermore,
\begin{align}\label{result of main term ass O H0}
\sqrt n S_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow S_{\infty}^{os}(\cdot,\theta_0),
\end{align}
where $S_{\infty}^{os}(\cdot,\theta_0)$ is a Gaussian process with mean zero and covariance structure
\begin{align*}
&Cov\left[S_{\infty}^{os}(\xi,\theta_0),S_{\infty}^{os}(\xi^\prime,\theta_0)\right] = \mathbb{E}\left[r^{os}_{\infty}(Y,W;\xi,\theta_0)\overline{r^{os}_{\infty}}(Y,W;\xi^\prime,\theta_0)\right],
\end{align*}
with
\begin{align}\label{theorem.ordinary smooth under H_0 ros}
&r^{os}_{\infty}(Y,W;\xi,\theta_0) = c_0^{os}(\xi)Y{\rm e}^{{\rm i} W\xi} - \sum\limits_{l=0}^{\alpha} c_l^{os}(\xi) [{\rm e}^{{\rm i} W\xi} g^{(l)}(W;\theta_0)] \notag\\
& - \left\{ \sum\limits_{l=0}^\alpha c_l^{os}(0) \left[Y\dot{g}^{(l)}(W;\theta_0)\right] - \sum\limits_{l=0}^\alpha c_l^{os}(0) \left[(g\dot{g})^{(l)}(W;\theta_0)\right] \right\}^\top\Delta^{-1}(\theta_0)G(\xi,\theta_0).
\end{align}
\end{theorem}
Based on Theorem \ref{theorem.ordinary smooth under H_0} and the continuous mapping theorem, see, e.g., \cite{van1996weak}, we can further derive the asymptotic null distributions of $KS_n$ and $CvM_n$ for the ordinary smooth case.
\begin{corollary}\label{Corollary.known ordinary}
Suppose that Assumptions \ref{ass.D} and \ref{ass.O} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align*}
\sqrt{n}KS_{n} \stackrel{d}\longrightarrow \sup\limits_{\xi\in\Pi}\left\vert S_{\infty}^{os}(\xi,\theta_0)\right\vert \quad\text{ and }
\quad nCvM_{n}\stackrel{d}\longrightarrow \int_{\Pi}\left\vert S_{\infty}^{os}(\xi,\theta_0)\right\vert^2\,d\xi.
\end{align*}
\end{corollary}
Theorem \ref{theorem.ordinary smooth under H_0} shows that $S_{n}^{pro}(\cdot,\hat\theta_n)$ converges weakly to a centered Gaussian process with its covariance structure depending on the data-generating process (DGP) and the parametric model $g(x;\theta_0)$ under the null. Subsequently, Corollary \ref{Corollary.known ordinary} shows that the $KS_n$ and $CvM_n$ statistics converge to the sup norm and the squared norm of the aforementioned Gaussian process, respectively. We note that $S_{n}^{pro}(\cdot,\hat\theta_n)$ exhibits $\sqrt{n}$-rate convergence for the ordinary smooth case and is asymptotically independent of the chosen bandwidth $b$, improving the convergence results of local smoothing tests.
For the supersmooth case where the characteristic function of measurement error decays at an exponential rate, e.g., the distribution of $\epsilon$ is normal, regularity conditions underlying the derivation of the asymptotic distribution of our proposed test statistics under the null are given as follows.
\begin{assumption}\label{ass.S}
\quad
\begin{enumerate}[label=(\roman*)]
\item The functions $f_X(x)$, $g(x;\theta_0)$, and $h(x;\theta_0)$ are infinitely differentiable with respect to $x$, where the function $h(x;\theta)$ is mentioned in Assumption \ref{ass.O}(i).
\item The measurement error $\epsilon$ follows a Gaussian distribution with the characteristic function of the following form for all $t\in\mathbb{R}$ and some positive constant $\mu$,
\begin{align*}
f^{ft}_\epsilon(t) = {\rm e}^{-\mu t^2}.
\end{align*}
\item The kernel function $K$ is infinitely differentiable and satisfies the following equations for all $l\in\mathbb{N}$,
\begin{align*}
\int K(u)du = 1, \qquad \int u^lK(u)du = 0.
\end{align*}
In addition, $K^{ft}$ has a compact support set that is symmetric around zero, and is bounded.
\item $b\to0$ as $n\to\infty$.
\item For $c_l^{ss}(\xi) = (-{\rm i})^l\sum\limits_{j\geq\frac{l}{2}}^{\infty}\frac{\mu^j}{j!}\binom{2j}{l}\xi^{2j-l}$, we have
\begin{align*}
&\mathbb{E}\left[\sup\limits_{\xi\in\Pi}\left\vert r^{ss}_{h, \infty}(Y,W;\xi,\theta_0)\right\vert^2\right]<\infty,
\end{align*}
where
\begin{align*}
&r^{ss}_{h, \infty}(Y,W;\xi,\theta_0) = \left\{ \sum_{l=0}^\infty c_l^{ss}(\xi) \left[Yh^{(l)}(W;\theta_0)\right] - \sum_{l=0}^\infty c_l^{ss}(\xi) \left[(gh)^{(l)}(W;\theta_0)\right] \right\}e^{{\rm i}W\xi}.
\end{align*}
\end{enumerate}
\end{assumption}
Assumption \ref{ass.S}({\romannumeral1}) requires structural functions and the distribution of latent variables to be infinitely smooth. Although assumptions about conditional mean functions are restrictive, commonly used functions, such as polynomials, circular functions, exponentials, and sums or products of such functions,
satisfy the requirement. Additionally, the construction of the class of infinitely differentiable density functions is also mentioned in \cite{alexander2009deconvolution}. Assumption \ref{ass.S}({\romannumeral2}) is a supersmooth condition. Assuming the function $f^{ft}_\epsilon$ has an exponential decay rate further enhances Assumption \ref{ass.O}({\romannumeral2}), with the Gaussian distribution being a typical example. The supersmoothness of measurement error and structural functions necessitate the use of an infinite-order kernel in Assumption \ref{ass.S}({\romannumeral3}) and integral properties to derive the asymptotic behavior of the statistic. Assumption \ref{ass.S}({\romannumeral4}) is simply a trivial bandwidth requirement, because the infinite-order smoothness condition removes the bias of the proposed empirical process under the null hypothesis, so that the undersmoothing condition required in Assumption \ref{ass.O}({\romannumeral4}) is no longer needed. Assumption \ref{ass.S}({\romannumeral5}) pertains to the boundedness of the asymptotic variance of our statistics, analogous to Assumption \ref{ass.O}({\romannumeral5}).
Using these assumptions, we derive the asymptotic behavior of the empirical process $S_{n}^{pro}(\cdot,\hat\theta_n)$ for the supersmooth case in the following theorem.
\begin{theorem}\label{theorem.supersmooth under H_0}
Suppose that Assumptions \ref{ass.D} and \ref{ass.S} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have \eqref{result of param effect ass O H0} and
\begin{align}\label{result of main term ass S H0}
\sqrt n S_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow S_{\infty}^{ss}(\cdot,\theta_0),
\end{align}
where $S_{\infty}^{ss}(\cdot,\theta_0)$ is a Gaussian process with mean zero and covariance structure
\begin{align*}
&Cov\left[S_{\infty}^{ss}(\xi,\theta_0),S_{\infty}^{ss}(\xi^\prime,\theta_0)\right] = \mathbb{E}\left[r^{ss}_{\infty}(Y,W;\xi,\theta_0)\overline{r^{ss}_{\infty}}(Y,W;\xi^\prime,\theta_0)\right],
\end{align*}
with
\begin{align}\label{theorem.supersmooth under H_0 rss}
&r^{ss}_{\infty}(Y,W;\xi,\theta_0) = c_0^{ss}(\xi)Y{\rm e}^{{\rm i} W\xi} - \sum_{l=0}^{\infty}c_l^{ss}(\xi)g^{(l)}(W;\theta_0){\rm e}^{{\rm i} W\xi}\notag\\
- & \left\{Y\sum_{l=0}^{\infty}c_l^{ss}(0)g^{(l)}(W;\theta_0){\rm e}^{{\rm i} W\xi} - \sum_{l=0}^{\infty}c_l^{ss}(0)\left[g\dot{g}\right]^{(l)}(W;\theta_0){\rm e}^{{\rm i} W\xi}\right\}^\top\Delta^{-1}(\theta_0)G(\xi,\theta_0).
\end{align}
\end{theorem}
Based on the theorem above, we use the continuous mapping theorem to derive the asymptotic null distributions of $KS_n$ and $CvM_n$ for the supersmooth case.
\begin{corollary}\label{Corollary.known super}
Suppose that Assumptions \ref{ass.D} and \ref{ass.S} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align*}
\sqrt{n}KS_{n} \stackrel{d}\longrightarrow \sup\limits_{\xi\in\Pi}\left\vert S_{\infty}^{ss}(\xi,\theta_0)\right\vert \quad \text{ and } \quad nCvM_{n} \stackrel{d}\longrightarrow \int_{\Pi}\left\vert S_{\infty}^{ss}(\xi,\theta_0)\right\vert^2\,d\xi.
\end{align*}
\end{corollary}
The projection plays a crucial role in Corollaries \ref{Corollary.known ordinary} and \ref{Corollary.known super}. By eliminating the effect of the estimator $\hat\theta_n$, the proposed test statistics achieve a parametric rate of convergence, even if the convergence rate of $\hat\theta_n$ is typically slower than $\sqrt{n}$ in the presence of measurement errors.
\subsection{Asymptotic power}
In this section, we proceed to derive the power properties of the proposed tests against a sequence of local alternatives. In contrast to \cite{otsu2021specification}, where the power of the local smoothing specification test is analyzed under local alternatives converging to the null hypothesis at a nonparametric rate, our tests have nontrivial power against local alternatives converging at a parametric rate under the following form:
\begin{align}\label{hyp.local alt}
H_{1n}:m(X) = g(X;\theta_0)+\frac{\Delta(X)}{\sqrt{n}} \quad a.s.
\end{align}
where $\Delta:\mathbb{R}\to\mathbb{R}$ is a bounded nonzero function satisfying the regularity conditions of the function $h$ mentioned in Assumptions \ref{ass.O} and \ref{ass.S}. The asymptotic local power properties of $S_{n}^{pro}(\cdot,\hat\theta_n)$ under $H_{1n}$ are established in the following theorem.
\begin{theorem}\label{theorem.known under H_1n}
Under the sequence of local alternatives $H_{1n}$ in \eqref{hyp.local alt}, we suppose that Assumption \ref{ass.D} holds. If Assumption \ref{ass.O} holds for the ordinary smooth case,
\begin{equation*}
\sqrt n S_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow S_{\infty}^{os}(\cdot,\theta_0)+\mu_\Delta(\cdot,\theta_0),
\end{equation*}
and if Assumption \ref{ass.S} holds for the supersmooth case,
\begin{equation*}
\sqrt n S_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow S_{\infty}^{ss}(\cdot,\theta_0)+\mu_\Delta(\cdot,\theta_0),
\end{equation*}
where $S_{\infty}^{os}(\cdot,\theta_0)$ and $S_{\infty}^{ss}(\cdot,\theta_0)$ are the centered Gaussian processes defined in Theorems \ref{theorem.ordinary smooth under H_0} and \ref{theorem.supersmooth under H_0}, respectively, and $\mu_\Delta(\cdot,\theta_0)$ is a deterministic shift process given by
\begin{align*}
\mu_\Delta(\xi,\theta_0) = \mathbb{E}\left\{\Delta(X)\left[{\rm e}^{{\rm i} X\xi}-\dot{g}^\top(X;\theta_0)\Delta^{-1}(\theta_0)G(\xi,\theta_0)\right]\right\}.
\end{align*}
\end{theorem}
Theorem \ref{theorem.known under H_1n} shows that under $H_{1n}$, both for the ordinary smooth and the supersmooth measurement errors, the asymptotic behavior of the $S_{n}^{pro}(\cdot,\hat\theta_n)$ consists of the limiting Gaussian process under $H_0$ and a deterministic shift term $\mu_\Delta(\cdot,\theta_0)$. By similar arguments to those of the null hypothesis, we then apply the continuous mapping theorem to obtain the convergence in distribution of $KS_n$ and $CvM_n$, which will have nontrivial power against local alternatives as described above, provided that $\mu_\Delta(\xi,\theta_0)\neq 0$ for some $\xi$.
By denoting $\theta^\ast = \operatorname*{plim}\hat{\theta}_n$ as the pseudo-true value and noting that $\theta^\ast = \theta_0$ under the null or the local alternatives, we can derive the asymptotic global power properties for our test statistics under the fixed alternative.
\begin{theorem}\label{theorem.known alternative}
Under the alternative hypothesis $H_{1}$ in \eqref{hyp.alternative1}, suppose that Assumption \ref{ass.D} holds. Under either Assumption \ref{ass.O} or Assumption \ref{ass.S},
\begin{align*}
&\sup_{\xi\in\Pi}\left\vert S_{n}^{pro}(\xi,\hat{\theta}_n) - C(\xi,\theta^\ast)\right\vert = o_p(1),
\end{align*}
where
\begin{align}
&C(\xi,\theta^\ast) = \mathbb{E}\left\{\left[m(X)-g(X;\theta^\ast)\right]\left[{\rm e}^{{\rm i} X\xi}-\dot{g}^\top(X;\theta^\ast)\Delta^{-1}(\theta^\ast)G(\xi,\theta^\ast)\right]\right\}. \label{Drift}
\end{align}
\end{theorem}
Note that $C(\cdot,\theta^\ast)$ can be understood as the projection of $m(X)-g(X;\theta^\ast)$ onto the orthogonal space of $\dot{g}(X;\theta^\ast)$, as explained in \cite{dominguez2015simple}. Under this interpretation, the consistency of our tests is guaranteed as long as for any vector $\gamma$,
\begin{align}\label{thm.known alt consis cond}
\mathbb{P}\left[m(X)-g(X;\theta^\ast) = \gamma^\top\dot{g}(X;\theta^\ast)\right]<1.
\end{align}
This ensures $C(\xi,\theta^\ast)\neq 0$ for at least some $\xi$ with a positive Lebesgue measure, thereby leading to the fact that $\sqrt nS_{n}^{pro}(\cdot,\hat{\theta}_n)$ diverges asymptotically and thus $\sqrt nKS_n$ and $nCvM_n$ will diverge to positive infinity in probability. This indicates that any fixed alternative satisfying \eqref{thm.known alt consis cond} can be detected by our proposed test with probability tending to one. Although the tests are not consistent against all possible alternatives, we do not regard those alternatives that violate this condition as a primary empirical concern, since they are nearly observationally equivalent to the null.
\section{Case of unknown measurement error distribution}\label{sec.unknown ME}
Since it is often challenging for researchers to obtain the density function of the measurement error in practical applications, we develop specification tests for settings with an unknown measurement error. As stated in \cite{delaigle2008deconvolution} and \cite{alexander2009deconvolution}, the approach typically used to address this problem is based on additional data, more specifically, repeated measurements on $X$ in the form of
\begin{align}\label{var.repeated ME}
&W^r = X + \epsilon^r,
\end{align}
where $\epsilon^r$ and $\epsilon$ are identically distributed
and $(X,\epsilon,\epsilon^r)$ are mutually independent. Repeated measurements are used to construct the following consistent estimator for $f^{ft}_\epsilon$,
\begin{align}\label{est.ME density}
&\hat{f}^{ft}_\epsilon(t) = \left\vert\frac{1}{n}\sum\limits_{i=1}^n\cos\left[t(W_i-W_i^r)\right]\right\vert^{1/2}
\end{align}
as proposed by \cite{delaigle2008deconvolution} and recommended in \cite{otsu2021specification}.
Our main contribution is to establish detailed theoretical results for ICM-type specification tests based on the repeated measurements approach described above. The theoretical justification and the practical implementation of the multiplier bootstrap for obtaining critical values are then developed and discussed in Section \ref{sec.boot}. Specifically, we construct the following empirical process based on the aforementioned estimator,
\begin{align}\label{Test stat unknown}
\hat{S}_{n}^{pro}(\xi,\hat\theta_n)=\frac{1}{n}\sum_{i=1}^n\int\left(Y_i-g(x;\hat{\theta}_n)\right)\hat{\mathcal{K}}_b\left(\frac{x-W_i}{b}\right)\hat{\mathcal{P}}_n(x;\xi,\hat{\theta}_n)\,dx,
\end{align}
where $\hat{\mathcal{K}}_b(\cdot)$ and $\hat{\mathcal{P}}_n(\cdot)$ denote the respective counterparts of $\mathcal{K}_b(\cdot)$ and $\mathcal{P}_n(\cdot)$ by replacing $f^{ft}_\epsilon(\cdot)$ with $\hat{f}^{ft}_\epsilon(\cdot)$. Using the empirical process $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ constructed above, the test statistics $CvM_n$ and $KS_n$ introduced earlier can be modified as
\begin{align*}
\widehat{CvM}_{n}=\int_\Pi\left\vert \hat{S}^{pro}_{n}(\xi,\hat{\theta}_n)\right\vert^2\,d\xi \quad\text{and}\quad \widehat{KS}_{n}=\sup_{\xi\in\Pi}\left\vert \hat{S}^{pro}_{n}(\xi,\hat{\theta}_n)\right\vert.
\end{align*}
Furthermore, to derive the asymptotic properties of the modified statistics, we need to strengthen the assumptions to address the uncertainty introduced by the estimation of the error characteristic functions, by imposing the following additional assumptions in addition to Assumption \ref{ass.D}.
\begin{assumption}\label{ass.D'}
\quad
\begin{enumerate}[label=(\roman*)]
\item $\{W_i^r\}_{i=1}^n$ is an i.i.d sample of $W^r$ satisfying \eqref{var.repeated ME}.
\item $\epsilon^r$ is i.i.d. as $\epsilon$, $\mathbb{E}\vert\epsilon\vert^{(p+1)(2+\zeta)}<\infty$ for some positive constant $\zeta$, and $f_\epsilon$ is symmetric around zero.
\end{enumerate}
\end{assumption}
Assumption \ref{ass.D'}({\romannumeral1}) is common in literature, see, e.g., \cite{delaigle2008deconvolution}, requiring independent and identically distributed repeated measurements.
Assumption \ref{ass.D'}({\romannumeral2}) is used to evaluate the asymptotic properties of the estimator $\hat{f}^{ft}_\epsilon(t)$ in \eqref{est.ME density}, as mentioned in \cite{kurisu2022uniform}. To derive the asymptotic distribution of our statistics under the ordinary smooth case, we need to impose the following assumptions.
\begin{assumption}\label{ass.O'}
\quad
\begin{enumerate}[label=(\roman*)]
\item $nb^{10\alpha+6}\log(\frac{1}{b})^{-4}\to \infty$ as $n\to\infty$.
\item Following the notation $r^{os}_{\infty}$ and the function $h(x;\theta)$ defined in Theorem \ref{theorem.ordinary smooth under H_0} and Assumption \ref{ass.O}, respectively,
\begin{align*}
&\mathbb{E}\left\{\int_{\Pi} \left[r^{os}_{\infty}(Y,W,\xi,\theta_0)+ r^{\epsilon}_{h,\infty}(Y,W,\xi,\theta_0)\right]^2\,d\xi\right\}<\infty,
\end{align*}
where
\begin{align*}
&r^{\epsilon}_{h,\infty}(Y,W,\xi,\theta_0) = \int \left[(gf_X)^{ft}(t)h^{ft}(\xi-t) - f_X^{ft}(t)(gh)^{ft}(\xi-t)\right]\Pi_{\epsilon}(t)dt,\\
&\Pi_{\epsilon}(t) = \frac{1}{2}-\frac{\cos(t(W-W^r))}{2\left\vert f^{ft}_\epsilon(t)\right\vert^2}.
\end{align*}
\end{enumerate}
\end{assumption}
In Assumption \ref{ass.O'}({\romannumeral1}), we enhance the assumption about bandwidth compared to Assumption \ref{ass.O}({\romannumeral4}).
This enhancement is necessary because in the absence of information about the measurement error distribution, the statistic is influenced by the uncertainty in estimating $f^{ft}_\epsilon$. Assumption \ref{ass.O'}({\romannumeral1}) ensures the asymptotic negligibility of such uncertainty brought by the estimator $\hat{f}^{ft}_\epsilon$. Next, compared to our Assumptions in \ref{ass.O}({\romannumeral5}), we add a moment restriction on the term brought by the estimator of the unknown distribution in Assumption \ref{ass.O'}({\romannumeral2}).
With the assumptions above, the following theorem, which characterizes the asymptotic behavior of the proposed empirical process with the unknown characteristic function replaced by its estimator, is established under the ordinary smooth case.
\begin{theorem}\label{theorem.unknown ordinary smooth under H_0}
Suppose that Assumptions \ref{ass.D}, \ref{ass.O}, \ref{ass.D'}, and \ref{ass.O'} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align}\label{result of param effect unknown ass O H0}
&\sup_{\xi\in\Pi}\left\vert \hat{S}_{n}^{pro}(\xi,\hat\theta_n)- \hat{S}_{n}^{pro}(\xi,\theta_0)\right\vert=o_p\left(n^{-\frac{1}{2}}\right).
\end{align}
Furthermore,
\begin{align}\label{result of main term unknown ass O H0}
\sqrt n \hat{S}_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow \hat{S}_{\infty}^{os}(\cdot,\theta_0),
\end{align}
where $\hat{S}_{\infty}^{os}(\cdot,\theta_0)$ is a Gaussian process with mean zero and covariance structure
\begin{align*}
&Cov\left[\hat{S}_{\infty}^{os}(\xi,\theta_0),\hat{S}_{\infty}^{os}(\xi^\prime,\theta_0)\right] = \mathbb{E}\left[\hat{r}^{os}_{\infty}(Y,W;\xi,\theta_0)\overline{\hat{r}^{os}_{\infty}}(Y,W;\xi^\prime,\theta_0)\right],
\end{align*}
where
\begin{align}\label{theorem.unknown ordinary smooth under H_0 roshat}
&\hat{r}^{os}_{\infty}(Y,W;\xi,\theta_0) = r^{os}_{\infty}(Y,W;\xi,\theta_0) + r^{\epsilon}_{\infty}(Y,W;\xi,\theta_0),
\end{align}
with
\begin{align}\label{theorem.unknown ordinary smooth under H_0 rplus}
&2\pi\cdot r^{\epsilon}_{\infty}(Y,W;\xi,\theta_0) = (gf_X)^{ft}(\xi)\Pi_{\epsilon}(\xi) - \int f_X^{ft}(t)g^{ft}(\xi-t)\Pi_{\epsilon}(t)dt\notag \\
& - \left\{\int \left[(gf_X)^{ft}(t)\dot{g}^{ft}(-t) - f_X^{ft}(t)(g\dot{g})^{ft}(-t)\right]\Pi_{\epsilon}(t)dt\right\}^\top\Delta^{-1}(\theta_0)G(\xi,\theta_0).
\end{align}
\end{theorem}
Using the continuous mapping theorem, we can readily derive the asymptotic null distributions of $\widehat{KS}_n$ and $\widehat{CvM}_n$ for the ordinary smooth case.
\begin{corollary}\label{Corollary.unknown ordinary}
Suppose that Assumptions \ref{ass.D},\ref{ass.D'},\ref{ass.O}, and \ref{ass.O'} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align*}
\sqrt{n}\widehat{KS}_{n} \stackrel{d}\longrightarrow \sup\limits_{\xi\in\Pi}\left\vert\hat{S}_{\infty}^{os}(\xi,\theta_0)\right\vert \quad\text{ and } \quad n\widehat{CvM}_{n} \stackrel{d}\longrightarrow \int_{\Pi}\left\vert\hat{S}_{\infty}^{os}(\xi,\theta_0)\right\vert^2d\xi.
\end{align*}
\end{corollary}
Theorem \ref{theorem.unknown ordinary smooth under H_0} shows that for the ordinary smooth case when the distribution of measurement error is unknown, $\hat{S}^{pro}_n(\cdot,\hat{\theta}_n)$ still converges at the $\sqrt{n}$-rate to a centered Gaussian process. But it is worth noting that the Gaussian process here is different from the one mentioned in Theorem \ref{theorem.ordinary smooth under H_0}, because the uncertainty brought by the $\hat{f}_{\epsilon}^{ft}$ (represented by the term $ r^{\epsilon}_{\infty}(Y,W;\xi,\theta_0)$) changes the covariance structure of the limiting process.
For the supersmooth case, the following assumptions are necessary in addition to Assumption \ref{ass.S} to derive asymptotic properties of our statistics.
\begin{assumption}\label{ass.S'}
\quad
\begin{enumerate}[label=(\roman*)]
\item $n{\rm e}^{-6\mu (1+b^{-1})^{2}}\log(\frac{1}{b})^{-2}\to \infty$ as $n\to\infty$.
\item Assume that
\begin{align*}
&\mathbb{E}\left\{\int_{\Pi} \left[r^{ss}_{\infty}(Y,W,\xi,\theta_0)+ r^{\epsilon}_{h,\infty}(Y,W,\xi,\theta_0)\right]^2\,d\xi\right\}<\infty,
\end{align*}
where $r^{ss}_{\infty}$ and $r^{\epsilon}_{h,\infty}$ are defined in Theorem \ref{theorem.supersmooth under H_0} and Assumption \ref{ass.O'}, respectively.
\end{enumerate}
\end{assumption}
Similar to Assumption \ref{ass.O'}, Assumption \ref{ass.S'}({\romannumeral1}) requires a stronger assumption about the bandwidth $b$ for ensuring the asymptotic negligibility of uncertainty brought by the estimation of the error characteristic functions, while Assumption \ref{ass.S'}({\romannumeral2}) ensures that the asymptotic variance of $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ is bounded through an additional moment restriction based on Assumption \ref{ass.S}({\romannumeral5}).
The following theorem characterizes the asymptotic behavior of $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ for the supersmooth case under the null hypothesis.
\begin{theorem}\label{theorem.unknown supersmooth under H_0}
Suppose that Assumptions \ref{ass.D},\ref{ass.S},\ref{ass.D'}, and \ref{ass.S'} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align}\label{result of main term unknown ass S H0}
\sqrt n \hat{S}_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow \hat{S}_{\infty}^{ss}(\cdot,\theta_0),
\end{align}
where $\hat{S}_{\infty}^{ss}(\cdot,\theta_0)$ is a Gaussian process with mean zero and covariance structure
\begin{align*}
&Cov\left[\hat{S}_{\infty}^{ss}(\xi,\theta_0),\hat{S}_{\infty}^{ss}(\xi^\prime,\theta_0)\right] = \mathbb{E}\left[\hat{r}^{ss}_{\infty}(Y,W;\xi,\theta_0)\overline{\hat{r}^{ss}_{\infty}}(Y,W;\xi^\prime,\theta_0)\right],
\end{align*}
where
\begin{align}\label{theorem.unknown supersmooth under H_0 rss}
&\hat{r}^{ss}_{\infty}(Y,W;\xi,\theta_0) = r^{ss}_{\infty}(Y,W;\xi,\theta_0) + r^\epsilon_{\infty}(Y,W;\xi,\theta_0)
\end{align}
and $r^\epsilon_{\infty}$ is defined in Theorem \ref{theorem.unknown ordinary smooth under H_0}.
\end{theorem}
\begin{corollary}\label{Corollary.unknown super}
Suppose that Assumptions \ref{ass.D}, \ref{ass.D'}, \ref{ass.S}, and \ref{ass.S'} hold. Then, under the null hypothesis $H_0$ in \eqref{hyp.null1}, we have
\begin{align*}
\sqrt{n}\widehat{KS}_{n} \stackrel{d}\longrightarrow \sup\limits_{\xi\in\Pi}\left\vert \hat{S}_{\infty}^{ss}(\xi,\theta_0)\right\vert \quad\text{ and }\quad n\widehat{CvM}_{n} \stackrel{d}\longrightarrow \int_{\Pi}\left\vert \hat{S}_{\infty}^{ss}(\xi,\theta_0)\right\vert^2\,d\xi.
\end{align*}
\end{corollary}
Theorem \ref{theorem.unknown supersmooth under H_0} shows that for the supersmooth case, our empirical process still maintains the $\sqrt{n}$-convergence. Furthermore, Corollary \ref{Corollary.unknown super} shows convergence in distribution of $\sqrt n\widehat{KS}_n$ and $n\widehat{CvM}_n$ under the null hypothesis. The subsequent analysis considers the results under the sequence of local alternatives and the fixed alternative, thereby establishing the asymptotic power properties of the tests without requiring knowledge of the measurement error distribution.
\begin{theorem}\label{theorem.unknown under H_1n}
Under the sequence of local alternatives $H_{1n}$ in \eqref{hyp.local alt}, suppose that Assumptions \ref{ass.D} and \ref{ass.D'} hold. For the ordinary smooth case, suppose that Assumptions \ref{ass.O} and \ref{ass.O'} hold,
\begin{align*}
\sqrt n \hat{S}_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow \hat{S}_{\infty}^{os}(\cdot,\theta_0)+\mu_\Delta(\cdot,\theta_0),
\end{align*}
and for the supersmooth case, suppose that Assumptions \ref{ass.S} and \ref{ass.S'} hold,
\begin{align*}
\sqrt n \hat{S}_{n}^{pro}(\cdot,\hat\theta_n)\Longrightarrow \hat{S}_{\infty}^{ss}(\cdot,\theta_0)+\mu_\Delta(\cdot,\theta_0),
\end{align*}
where $\hat{S}_{\infty}^{os}(\cdot,\theta_0)$ and $\hat{S}_{\infty}^{ss}(\cdot,\theta_0)$ are the centered Gaussian processes defined in Theorems \ref{theorem.unknown ordinary smooth under H_0} and \ref{theorem.unknown supersmooth under H_0}, respectively, and $\mu_\Delta(\cdot,\theta_0)$ is the deterministic shift process defined in Theorem \ref{theorem.known under H_1n}.
\end{theorem}
\begin{theorem}\label{theorem.unknown alternative}
Under the alternative hypothesis $H_{1}$ in \eqref{hyp.alternative1}, suppose that Assumptions \ref{ass.D} and \ref{ass.D'} hold. Whether Assumptions \ref{ass.O} and \ref{ass.O'} hold for the ordinary smooth case or Assumptions \ref{ass.S} and \ref{ass.S'} hold for the supersmooth case, we have
\begin{equation*}
\sup_{\xi\in\Pi}\left\vert \hat{S}_{n}^{pro}(\xi,\hat{\theta}_n) - C(\xi,\theta^\ast) \right\vert = o_p(1),
\end{equation*}
where $C(\cdot,\theta^\ast)$ is defined in \eqref{Drift}.
\end{theorem}
Under the sequence of local alternative hypotheses defined in \eqref{hyp.local alt}, we still have a deterministic shift term similar to Theorem \ref{theorem.known under H_1n}, signifying the nontrivial local power of our tests. We also find that under the alternative hypothesis $H_1$, our proposed specification tests share the same consistency condition as in the case where the measurement error distribution is known, see condition \eqref{thm.known alt consis cond}.
\section{Multiplier bootstrap}\label{sec.boot}
As we show in Sections \ref{sec.Asy} and \ref{sec.unknown ME}, the null limiting distributions of test statistics $KS_{n}$ and $CvM_{n}$ usually depend on the underlying DGP in a rather complicated manner. As such, the associated critical values are case-dependent, forcing people to resort to bootstrap procedures. Unfortunately, although the wild bootstrap procedure is typically employed to implement tests in the error-free case, no computationally easy and valid bootstrap procedures are currently available for tests in the presence of measurement errors, as indicated above. In particular, any residual-based bootstrap methods (e.g., the widely used wild bootstrap in the literature of specification tests) are apparently infeasible in the presence of measurement errors, given that we need the true regressors to construct the residuals, but we cannot observe the regressors directly.
For the influential local smoothing specification test proposed in \cite{otsu2021specification}, a multi-step estimation and resampling method is employed to operationalize the bootstrap procedure, a strategy also adopted in classical studies such as \cite{hall2007testing} for global smoothing tests. Specifically, they first estimate the density function of the unobservable true regressor $X$, the fitted nonparametric regression function, and the parametric fit using standard deconvolution techniques and a parametric estimator. Then, the residuals are resampled by estimating their second moments and drawing from distributions such as the two-point distribution of \cite{mammen1993bootstrap}. Based on these resampled residuals, bootstrap samples of $Y$ are generated as the sum of the resampled residuals and the estimated parametric fit evaluated at the resampled regressors drawn from the estimated density of $X$. The measurement error–contaminated bootstrap samples of variable $W$ are subsequently obtained by adding the true regressor and an independent draw from the known (or estimated) measurement error distribution. Finally, the bootstrap counterparts of the test statistics are computed using each resampled pair $\{(Y_i^\ast,W_i^\ast)^\top\}_{i=1}^n$. Furthermore, bandwidth selection is also a critical issue for local smoothing tests due to the sensitivity of local smoothers to the bandwidth choice, which is addressed by adopting the \textquotedblleft two-stage selection plug-in\textquotedblright{} method of \cite{delaigle2004practical} in the test of \cite{otsu2021specification}.
While the above approach is feasible, it still has several limitations that deserve further attention. First, the complex estimation and resampling steps complicate both theoretical justification and empirical implementation. In particular, integrals involving deconvolution kernels often lack closed-form solutions for many higher-order kernel functions, and numerical approximation may reduce accuracy. Second, within each simulation, bootstrap iteration requires repeating the estimation of functions and the computation of the bootstrap counterparts of the test statistics, making the overall procedure computationally inefficient. Third, the need to select suitable bandwidths adds further computational complexity and theoretical difficulties. In our proposed tests, we eliminate the parameter estimation effect via the projection approach, thereby facilitating a much simpler multiplier bootstrap and avoiding the aforementioned problems. This serves as one of the main contributions of the study.
\subsection{Known measurement error}\label{subsec.kme}
For the cases with known measurement error distribution, we approximate the asymptotic null distribution of a continuous functional $\Gamma(S_n^{pro})$ by that of $\Gamma(S_n^{pro,\ast})$, where $\Gamma(\cdot)$ represents the sup norm and the squared norm employed in constructing the statistics $KS_{n}$ and $CvM_{n}$, respectively, and
\begin{align}\label{bootstrap version of stat known ME}
S_{n}^{pro,\ast}(\xi,\hat\theta_n)=\frac{1}{n}\sum_{i=1}^n\int V_i\left(Y_i-g(x;\hat{\theta}_n)\right)\mathcal{K}_b\left(\frac{x-W_i}{b}\right)\mathcal{P}_n (x;\xi,\hat{\theta}_n)\,dx.
\end{align}
Here, $\{V_i\}_{i=1}^n$ is a sequence of i.i.d. random variables with zero mean, unit variance, and independent of the original sample, known as the multipliers. A popular example is i.i.d. Bernoulli variables constructed by \cite{mammen1993bootstrap}. Both for the ordinary smooth case and the supersmooth case, $S_{n}^{pro,\ast}(\cdot,\hat\theta_n)$ is expected to have the same limiting process as $S_{n}^{pro}(\cdot,\hat\theta_n)$ under the null hypothesis, due to the zero mean and unit variance properties of the multipliers. Furthermore, we expect that the multipliers can eliminate the deterministic shift term mentioned in Theorem \ref{theorem.known under H_1n} under the local alternatives, thus establishing that $S_{n}^{pro,\ast}(\cdot,\hat\theta_n)$ converges to the same limiting process as under the null hypothesis.
Defining \textquotedblleft $\overset{\ast}{\Longrightarrow}$\textquotedblright{} as weak convergence and $\mathbb{P}_n^\ast$ as the bootstrap probability under the bootstrap law, i.e., conditional on the original sample, see, e.g., Section 2.9 of \cite{van1996weak}, the validity of the proposed multiplier bootstrap is formally established by the following theorem.
\begin{theorem}\label{theorem.boot known}
Under both the null hypothesis $H_0$ in \eqref{hyp.null1} and the sequence of local alternatives $H_{1n}$ in \eqref{hyp.local alt}, suppose that Assumptions \ref{ass.D} and \ref{ass.O} hold for the ordinary smooth case, and Assumptions \ref{ass.D} and \ref{ass.S} hold for the supersmooth case. Then, we have
\begin{align}\label{result of param effect boot}
\sup_{\xi\in\Pi}\left\vert S_n^{pro,\ast}(\xi,\hat\theta_n)- S_n^{pro,\ast}(\xi,\theta_0)\right\vert=o_p\left(n^{-\frac{1}{2}}\right).
\end{align}
Furthermore, for the ordinary smooth case,
\begin{align}\label{result of main term ass O boot}
\sqrt n S_{n}^{pro,\ast}(\cdot,\hat\theta_n)\overset{\ast}{\Longrightarrow}S_{\infty}^{os}(\cdot,\theta_0),
\end{align}
and for the supersmooth case,
\begin{align}\label{result of main term ass S boot}
\sqrt n S_{n}^{pro,\ast}(\cdot,\hat\theta_n)\overset{\ast}{\Longrightarrow}S_{\infty}^{ss}(\cdot,\theta_0),
\end{align}
where $S_{\infty}^{os}(\cdot,\theta_0)$ and $S_{\infty}^{ss}(\cdot,\theta_0)$ are the centered Gaussian processes defined in Theorems \ref{theorem.ordinary smooth under H_0} and \ref{theorem.supersmooth under H_0}, respectively.
\end{theorem}
For the ordinary smooth and the supersmooth cases, respectively, the bootstrap empirical processes share the same limiting behavior under $H_0$ and $H_{1n}$ as their original sample counterparts under the null, while the latter exhibit a deterministic shift under $H_{1n}$. As a consequence of the above analysis, the asymptotic critical value at the significance level $\alpha$ is $c^\ast_{\alpha}=\inf\{c_\alpha\in[0,\infty):\lim_{n\to\infty}\mathbb{P}_n^\ast\{\Gamma(\sqrt n S_{n}^{pro,\ast})>c_\alpha\}=\alpha\}$. In practice, $c^\ast_{\alpha}$ can be approximated as $c^\ast_{n,\alpha} = \{\Gamma(\sqrt n S_{n}^{pro,\ast})\}_{B(1-\alpha)}$, the $B(1-\alpha)$-th order statistic for $B$ replicates $\{\Gamma(\sqrt n S_{n,b}^{pro,\ast})\}_{b=1}^B$ of $\{\Gamma(\sqrt n S_{n}^{pro,\ast})\}$ and we reject $H_0$ if $\Gamma(\sqrt n S_{n}^{pro})>c^\ast_{n,\alpha}$. Specifically, building upon Theorem \ref{theorem.boot known} and taking $\Gamma(\cdot)$ to be the sup norm and the squared norm, respectively, we establish the asymptotic validity of the multiplier bootstrap procedure for the $KS_n$ and $CvM_n$ statistics.
\subsection{Unknown measurement error}
As we mentioned in Section \ref{sec.unknown ME}, in the absence of information about the distribution of measurement errors, we argue that our multiplier bootstrap procedure remains feasible. We still use repeated measurements to estimate the characteristic function of the measurement error, thereby estimating the deconvolution kernel as shown in \eqref{var.repeated ME}. Then we can construct the multiplier bootstrap version of $\hat{S}_n^{pro}(\cdot,\hat{\theta}_n)$ as given by
\begin{align}\label{bootstrap version of stat unknown ME}
\hat S_{n}^{pro,\ast}(\xi,\hat\theta_n)=\frac{1}{n}\sum_{i=1}^n\int V_i\left(Y_i-g(x;\hat{\theta}_n)\right)\hat{\mathcal{K}}_b\left(\frac{x-W_i}{b}\right)\hat{\mathcal{P}}_n(x;\xi,\hat{\theta}_n)\,dx.
\end{align}
Similar to what we have shown in Theorem \ref{theorem.boot known}, we establish the validity of our bootstrap procedure by the following theorem for the unknown measurement error case.
\begin{theorem}\label{theorem.boot unknown}
Suppose that Assumptions \ref{ass.D} and \ref{ass.D'} hold. Under both the null hypothesis $H_0$ in \eqref{hyp.null1} and the sequence of local alternatives $H_{1n}$ in \eqref{hyp.local alt}, suppose that Assumptions \ref{ass.O} and \ref{ass.O'} hold for the ordinary smooth case, and Assumptions \ref{ass.S} and \ref{ass.S'} hold for the supersmooth case. Then, we have
\begin{align}\label{result of param effect unknown boot}
\sup_{\xi\in\Pi}\left\vert \hat{S}_n^{pro,\ast}(\xi,\hat\theta_n)- \hat{S}_n^{pro,\ast}(\xi,\theta_0)\right\vert=o_p\left(n^{-\frac{1}{2}}\right).
\end{align}
Furthermore, for the ordinary smooth case,
\begin{align}\label{result of main term ass O unknown boot}
\sqrt n \hat{S}_{n}^{pro,\ast}(\cdot,\hat\theta_n)\overset{\ast}{\Longrightarrow}\hat{S}_{\infty}^{os}(\cdot,\theta_0),
\end{align}
and for the supersmooth case,
\begin{align}\label{result of main term ass S unknown boot}
\sqrt n \hat{S}_{n}^{pro,\ast}(\cdot,\hat\theta_n)\overset{\ast}{\Longrightarrow}\hat{S}_{\infty}^{ss}(\cdot,\theta_0),
\end{align}
where $\hat{S}_{\infty}^{os}(\cdot,\theta_0)$ and $\hat{S}_{\infty}^{ss}(\cdot,\theta_0)$ are the centered Gaussian processes defined in Theorems \ref{theorem.unknown ordinary smooth under H_0} and \ref{theorem.unknown supersmooth under H_0}, respectively.
\end{theorem}
Theorem \ref{theorem.boot unknown} shows that the bootstrap version of the empirical process constructed by repeated measurements converges to the same limiting process under $H_0$ and $H_{1n}$, which is the same limiting process as the one under the null hypothesis mentioned in Section \ref{sec.unknown ME}, for the ordinary smooth and super smooth cases, respectively. Consequently, we can validate the multiplier bootstrap for $\widehat{KS}_n$ and $\widehat{CvM}_n$ using the continuous mapping theorem and obtain the critical values as discussed in Subsection \ref{subsec.kme}. Therefore, using the multiplier bootstrap to obtain the critical values for our tests remains feasible even when the measurement error is unknown.
\section{Numerical evidence}\label{sec.Simulation}
In this section, we investigate the finite-sample performance of the proposed test by Monte Carlo experiments. We set the unobservable regressors $\{X_i\}_{i=1}^n$ distributed as $N(0,1)$, and we use a trigonometric model and a polynomial model in our simulation, although we note that any nonlinear model can be used in our test procedure. We name $Y_i=1+X_i+\delta X_i^2+U_i$ as model one and $Y_i=1+X_i+\delta \cos(\pi X_i)+U_i$ as model two, respectively, for $i$ ranging from $1$ to $n$, where $\delta$ is a variable constant representing the nonlinearity of the model and is set to $0.5$. Additionally, $U_i\sim N(0,1/4)$. The contaminated regressor is given by $W_i=X_i+\epsilon_i$, where we use the Laplace distribution with variance of $1/12$ for the ordinary smooth case and $N(0,1/12)$ for the supersmooth case to remain consistent with the experiment in \cite{otsu2021specification}. We use the infinite-order flat-top kernel proposed by \cite{mcmurry2004nonparametric}, which has also been employed in \cite{dong2022nonparametric}, for both cases and all of our simulations, which is defined by
\begin{equation*}
K^{ft}(t) = \begin{cases}
1 & \text{if } |t| \leq 0.05, \\
\exp \left\{ \frac{-\exp(-(|t|-0.05)^{-2})}{(|t|-1)^2} \right\} & \text{if } 0.05 < |t| < 1, \\
0 & \text{if } |t| \geq 1.
\end{cases}
\end{equation*}
We choose the bandwidth according to the rules of thumb, see, e.g., \cite{delaigle2008deconvolution}. Specifically, we use $b=c(5\sigma^4/n)^{1/27}$ for the ordinary smooth case and $b = c(4\sigma^2/\log(n))^{1/2}$ for the supersmooth case, where $\sigma$ is the standard deviation of the measurement error and sensitivity indicator $c$ varies in the grid presented in the following table. For the estimator in our test, we use a polynomial estimator of degree $2$, as mentioned in \cite{cheng1998polynomial}, which is also consistent with the results in \cite{otsu2021specification}.
We report the simulation results based on $199$ bootstrap iterations and $1000$ Monte Carlo replications. Tables include the size (the proportion of rejections when the DGP is under the null hypothesis, named DGP$(0)$) and powers (the proportion of rejections when the DGP is under the alternative hypothesis model one, named DGP$(1)$, and model two, named DGP$(2)$). We report the results for two statistics $KS_n$ and $CvM_n$, two sample sizes $n=\{500,1000\}$, three nominal level $\alpha = \{1\%,5\%,10\%\}$ and varied sensitivity indicators. Constrained by the length of papers, only part of the results $c\in\{1,5,10\}$ and $\alpha\in\{5\%,10\%\}$ are shown in the main text, and the complete results are reported in Tables \ref{tab:knwon laplace 500}--\ref{tab:unknwon normal 1000} of Appendix \ref{sec.Appendix sim res}.
\begin{table}[ht!]
\centering
\caption{Results under known ordinary smooth case}
\scalebox{0.95}{
\begin{tabular}{c|c|c|cccccc}
\hline\hline
\multirow{2}{*}{$n$} & \multirow{2}{*}{$c$} & \multirow{2}{*}{level} & \multicolumn{2}{c}{DGP(0)} & \multicolumn{2}{c}{DGP(1)} & \multicolumn{2}{c}{DGP(2)}\\
\cline{4-9}
\quad & \quad & \quad & KS & CvM & KS & CvM & KS & CvM\\
\hline
\multirow{6}{*}{500} & \multirow{2}{*}{1} & 0.05 & 0.042 & 0.041 & 0.815 & 0.806 & 0.775 & 0.762\\
\quad & \quad & 0.1 & 0.115 & 0.114 & 0.898 & 0.888 & 0.864 & 0.849\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.058 & 0.054 & 0.835 & 0.818 & 0.779 & 0.769\\
\quad & \quad & 0.1 & 0.099 & 0.101 & 0.890 & 0.882 & 0.870 & 0.856\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.051 & 0.052 & 0.818 & 0.797 & 0.745 & 0.735\\
\quad & \quad & 0.1 & 0.093 & 0.093 & 0.881 & 0.870 & 0.855 & 0.843\\
\hline
\multirow{6}{*}{1000} & \multirow{2}{*}{1} & 0.05 & 0.054 & 0.052 & 0.985 & 0.979 & 0.965 & 0.959\\
\quad & \quad & 0.1 & 0.123 & 0.117 & 0.990 & 0.987 & 0.979 & 0.974\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.052 & 0.050 & 0.977 & 0.968 & 0.966 & 0.955\\
\quad & \quad & 0.1 & 0.101 & 0.099 & 0.993 & 0.988 & 0.978 & 0.973\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.049 & 0.050 & 0.969 & 0.962 & 0.969 & 0.964\\
\quad & \quad & 0.1 & 0.087 & 0.090 & 0.985 & 0.979 & 0.987 & 0.978\\
\hline\hline
\end{tabular}
}
\label{tab:knwon laplace}
\end{table}
\begin{table}[ht!]
\centering
\caption{Results under known supersmooth case}
\scalebox{0.95}{
\begin{tabular}{c|c|c|cccccc}
\hline\hline
\multirow{2}{*}{$n$} & \multirow{2}{*}{$c$} & \multirow{2}{*}{level} & \multicolumn{2}{c}{DGP(0)} & \multicolumn{2}{c}{DGP(1)} & \multicolumn{2}{c}{DGP(2)}\\
\cline{4-9}
\quad & \quad & \quad & KS & CvM & KS & CvM & KS & CvM\\
\hline
\multirow{6}{*}{500} & \multirow{2}{*}{1} & 0.05 & 0.059 & 0.061 & 0.921 & 0.920 & 0.893 & 0.885\\
\quad & \quad & 0.1 & 0.109 & 0.108 & 0.961 & 0.957 & 0.935 & 0.933\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.062 & 0.061 & 0.933 & 0.932 & 0.901 & 0.896\\
\quad & \quad & 0.1 & 0.107 & 0.109 & 0.953 & 0.948 & 0.931 & 0.929\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.042 & 0.045 & 0.953 & 0.951 & 0.901 & 0.898\\
\quad & \quad & 0.1 & 0.110 & 0.113 & 0.971 & 0.969 & 0.937 & 0.936\\
\hline
\multirow{6}{*}{1000} & \multirow{2}{*}{1} & 0.05 & 0.056 & 0.048 & 0.997 & 0.997 & 0.998 & 0.997\\
\quad & \quad & 0.1 & 0.110 & 0.111 & 0.999 & 0.999 & 0.997 & 0.997\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.050 & 0.049 & 1.000 & 1.000 & 0.993 & 0.992\\
\quad & \quad & 0.1 & 0.114 & 0.114 & 1.000 & 1.000 & 0.998 & 0.998\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.056 & 0.057 & 0.996 & 0.995 & 0.994 & 0.994\\
\quad & \quad & 0.1 & 0.110 & 0.108 & 0.999 & 0.999 & 0.998 & 0.998\\
\hline\hline
\end{tabular}
}
\label{tab:known normal}
\end{table}
\begin{table}[ht!]
\centering
\caption{Results under unknown ordinary smooth case}
\scalebox{0.95}{
\begin{tabular}{c|c|c|cccccc}
\hline\hline
\multirow{2}{*}{$n$} & \multirow{2}{*}{$c$} & \multirow{2}{*}{level} & \multicolumn{2}{c}{DGP(0)} & \multicolumn{2}{c}{DGP(1)} & \multicolumn{2}{c}{DGP(2)}\\
\cline{4-9}
\quad & \quad & \quad & KS & CvM & KS & CvM & KS & CvM\\
\hline
\multirow{6}{*}{500} & \multirow{2}{*}{1} & 0.05 & 0.060 & 0.059 & 0.814 & 0.796 & 0.777 & 0.768\\
\quad & \quad & 0.1 & 0.133 & 0.128 & 0.892 & 0.874 & 0.868 & 0.851\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.055 & 0.053 & 0.822 & 0.801 & 0.786 & 0.770\\
\quad & \quad & 0.1 & 0.116 & 0.108 & 0.891 & 0.885 & 0.844 & 0.837\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.042 & 0.041 & 0.830 & 0.812 & 0.777 & 0.766\\
\quad & \quad & 0.1 & 0.104 & 0.105 & 0.909 & 0.900 & 0.844 & 0.831\\
\hline
\multirow{6}{*}{1000} & \multirow{2}{*}{1} & 0.05 & 0.061 & 0.058 & 0.970 & 0.963 & 0.949 & 0.946\\
\quad & \quad & 0.1 & 0.118 & 0.115 & 0.989 & 0.984 & 0.969 & 0.964\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.062 & 0.060 & 0.980 & 0.970 & 0.964 & 0.956\\
\quad & \quad & 0.1 & 0.116 & 0.111 & 0.989 & 0.987 & 0.980 & 0.976\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.057 & 0.052 & 0.980 & 0.978 & 0.955 & 0.946\\
\quad & \quad & 0.1 & 0.105 & 0.104 & 0.992 & 0.988 & 0.982 & 0.973\\
\hline\hline
\end{tabular}
}
\label{tab:unknown laplace}
\end{table}
\begin{table}[ht!]
\centering
\caption{Results under unknown supersmooth case}
\scalebox{0.95}{
\begin{tabular}{c|c|c|cccccc}
\hline\hline
\multirow{2}{*}{$n$} & \multirow{2}{*}{$c$} & \multirow{2}{*}{level} & \multicolumn{2}{c}{DGP(0)} & \multicolumn{2}{c}{DGP(1)} & \multicolumn{2}{c}{DGP(2)}\\
\cline{4-9}
\quad & \quad & \quad & KS & CvM & KS & CvM & KS & CvM\\
\hline
\multirow{6}{*}{500} & \multirow{2}{*}{1} & 0.05 & 0.056 & 0.057 & 0.921 & 0.920 & 0.895 & 0.886\\
\quad & \quad & 0.1 & 0.114 & 0.115 & 0.974 & 0.973 & 0.936 & 0.933\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.057 & 0.057 & 0.927 & 0.924 & 0.901 & 0.901\\
\quad & \quad & 0.05 & 0.057 & 0.057 & 0.927 & 0.924 & 0.901 & 0.901\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.055 & 0.055 & 0.927 & 0.927 & 0.896 & 0.894\\
\quad & \quad & 0.1 & 0.092 & 0.092 & 0.973 & 0.971 & 0.932 & 0.932\\
\hline
\multirow{6}{*}{1000} & \multirow{2}{*}{1} & 0.05 & 0.046 & 0.045 & 0.998 & 0.998 & 0.995 & 0.993\\
\quad & \quad & 0.1 & 0.116 & 0.115 & 1.000 & 1.000 & 0.998 & 0.998\\
\cline{2-9}
\quad & \multirow{2}{*}{5} & 0.05 & 0.052 & 0.051 & 0.996 & 0.996 & 0.993 & 0.993\\
\quad & \quad & 0.1 & 0.097 & 0.102 & 0.999 & 0.999 & 0.998 & 0.997\\
\cline{2-9}
\quad & \multirow{2}{*}{10} & 0.05 & 0.058 & 0.057 & 1.000 & 0.999 & 0.993 & 0.992\\
\quad & \quad & 0.1 & 0.108 & 0.106 & 0.999 & 0.999 & 1.000 & 1.000\\
\hline\hline
\end{tabular}
}
\label{tab:unknown normal}
\end{table}
Tables \ref{tab:knwon laplace} and \ref{tab:known normal} show our testing results when we know the distribution of the measurement error for the ordinary smooth case and the supersmooth case, respectively. In our results, the $KS_n$ and $CvM_n$ statistics perform well
in terms of both size and power.
We observe significant improvements in size accuracy and power gain as the sample size increases.
We also observe that the size and power of the tests do not fluctuate much with changes in bandwidth values on the grid. Comparing the results between the ordinary smooth case and the supersmooth case, we find that test sizes for the supersmooth cases perform comparably to those for the ordinary smooth cases, which differs from the findings reported in \cite{otsu2021specification}. It stems from the fact that the statistics for both the ordinary smooth and supersmooth cases exhibit $\sqrt{n}$-rate convergence, another advantage of the proposed tests. Other results in the tables reveal the power properties of our tests, where DGP$(1)$ represents the polynomial model and DGP$(2)$ represents the trigonometric model. It is well known that global smoothing tests perform better in polynomial models, while local smoothing tests perform better in trigonometric models. For our results in Tables \ref{tab:knwon laplace} and \ref{tab:known normal}, the test performance under DGP$(1)$ is better than that under DGP$(2)$. We can expect better performance with higher nonlinearity and more samples.
Tables \ref{tab:unknown laplace} and \ref{tab:unknown normal} display size and power properties of the $\widehat{KS}_n$ and $\widehat{CvM}_n$ statistics in the absence of information about the distribution of the measurement error. However, because our asymptotic variance is affected by the estimation error of the characteristic functions, the results for size and power in finite samples are slightly worse but remain within the acceptable range.
We can also observe that our tests perform better under the polynomial model than under the trigonometric model. Meanwhile, our tests still retain robustness to bandwidth selection when the measurement error distribution is unknown.
Finally, we summarize the advantages of our tests, as demonstrated in the simulations, including a good level of accuracy and power properties due to the $\sqrt{n}$-convergence of the statistics, better efficiency for polynomials and low-frequency alternative hypotheses, and robustness in bandwidth selection. Additionally, our testing framework naturally accommodates the case without measurement error.
\section{Conclusion}\label{sec.Conclusion}
In this paper, we propose new ICM-type specification tests for regression models with measurement errors based on a deconvoluted residual-marked empirical process. The innovation of our method is the use of a projection to eliminate the parameter estimation effect, a particularly attractive tool in the presence of measurement errors.
We formally establish the asymptotic properties of the projection-based
test statistics and suggest a straightforward multiplier bootstrap procedure to simulate the critical values. Tests and the multiplier bootstrap procedure are also examined when the measurement error distribution is unknown. We emphasize three main advantages of the tests proposed in this paper. From a theoretical perspective, the statistics of $\sqrt{n}$-convergence for both the ordinary smooth and the supersmooth measurement errors ensure excellent size accuracy and satisfactory power properties, especially for the supersmooth cases, where slower convergence rates can typically be achieved, as mentioned in the previous literature on measurement error. From the application perspective, our proposed method enables a general convergence rate of the parametric estimator, thereby resolving the difficulty in obtaining a $\sqrt{n}$-convergence estimator in the presence of measurement errors. Additionally, our tests are more robust to bandwidth selection and, therefore, more reliable. From a computational perspective, our projection approach enables us to
simulate the critical values as accurately as desired via the multiplier bootstrap, which is easier to understand and compute.
For future research, we anticipate that the valuable combination of projection and multiplier bootstrap can be extended to other interesting testing problems in the presence of measurement errors, such as tests for heteroskedasticity and significance tests.
\bibliographystyle{apalike_revised}
\bibliography{Ref}
\newpage
\begin{center}
\Large{Specification tests for regression models with
measurement errors\\
-- Online supplementary appendix}
\end{center}
\begin{center}
\Large{Xiaojun Song and Jichao Yuan\\
Peking University}
\end{center}
\normalsize
In this appendix, we provide additional simulation results and proofs of the main theoretical results.