Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,234 characters · 13 sections · 79 citation commands
Specification tests for regression models with measurement errors
\doublespacing
There already exists a substantial body of literature on specification testing for regression models, see, e.g., bierens1982consistent, hardle1993comparing, JOHNXUZHENG1996263, stute1997nonparametric, stute1998bootstrap, stinchcombe1998consistent, stute2002model, zhu2003testing, zhu2005testing, Escanciano_2006, hall2007testing, song2008model, xu2015nonparametric, sant2019specification, otsu2021specification, and TAN2025106113. The list is undoubtedly not exhaustive. However, among these, relatively few studies have adequately developed specification tests in the presence of measurement errors. Measurement error is a common issue in data across many disciplines, including economics, finance, and medicine, see, e.g., fuller2009measurement, alexander2009deconvolution, and hu2017measurement. Unfortunately, specification tests often have incorrect size and low power in the presence of measurement errors.
In an earlier influential work, otsu2021specification proposed a specification test for regression models with measurement errors in the explanatory variables based on nonparametric local smoothing estimators. Their test belongs to the class of local smoothing tests, and so does the minimum distance test proposed by song2008model. These tests are typically constructed from the distance between the fitted nonparametric regression and the parametric fit under the null hypothesis, and have been shown to be more powerful against high-frequency alternatives. Complementary to such approaches, global smoothing tests, developed in hall2007testing, exhibit higher power against low-frequency alternatives. As emphasized in otsu2021specification, local smoothing and global smoothing tests serve as complements rather than substitutes. A similar complementarity in power properties between the two types of tests, in the absence of measurement errors, was also established in fan2000consistent.
Within the class of global smoothing tests, the integrated conditional moment (ICM)-type tests proposed by bierens1982consistent and bierens1997asymptotic are particularly important and are the main focus of this paper. These tests typically compare the integrated regression function with the integrated parametric regression function under the null hypothesis. Specifically, a residual-based empirical process is employed to construct the test statistics, and, in general, the asymptotic null distributions depend on the data-generating process, thereby motivating the use of bootstrap techniques to obtain the critical values. The ICM-type tests have been widely applied to specification testing when measurement errors are absent, as shown in su1991lack, stute1997nonparametric, and Escanciano_2006, among others. However, in the presence of measurement errors, the ICM-type tests face challenges arising from the unobservability of the true residuals, the stringent requirements for the estimators, and the computational complexity of obtaining critical values, which motivate the present study to address these issues. In this paper, we propose ICM-type tests based on a deconvoluted residual-marked empirical process (i.e., the residuals are constructed using the deconvolution kernel estimator). We employ an orthogonal projection onto the tangent space of nuisance parameters to eliminate the parameter estimation effect. As a desirable byproduct, the projection facilitates the simulation of critical values via a computationally straightforward multiplier bootstrap.
It is important to highlight our contribution regarding the simple multiplier bootstrap procedure for obtaining critical values. Conventional bootstrap methods face serious challenges in the presence of measurement errors because the true regressors are unobservable; consequently, we cannot resample them. Infeasibility of obtaining the critical values through the multiplier bootstrap procedure is due to the parameter estimation effect, initially discussed in durbin1973distribution. We introduce a novel projection to address the parameter estimation effect by imposing an orthogonality condition on the weight function of the residual-marked empirical process. The introduction of the projection also facilitates a computationally attractive multiplier bootstrap procedure to implement our tests, whose asymptotic validity can be easily justified. Notably, compared with other bootstrap methods, the multiplier bootstrap is simpler to justify theoretically, easier to implement, and computationally more efficient. To our knowledge, this is the first time a multiplier bootstrap has been proposed in the literature of specification testing in the presence of measurement errors. Furthermore, we establish a comprehensive theoretical framework and a multiplier bootstrap-based procedure for obtaining critical values when the measurement error distribution is unknown. In addition, our tests are more robust to bandwidth choices than the local smoothing approach.
The rest of this paper is organized as follows. In Section (ref), we describe the testing framework and introduce our projection method. Section (ref) discusses the asymptotic properties of the test statistics. Section (ref) addresses the case of unknown measurement error distribution. In Section (ref), we provide a multiplier bootstrap procedure and its implementation. Monte Carlo simulations are conducted in Section (ref). We present our conclusions in Section (ref). Additional simulation results and the mathematical proofs are included in the online supplementary appendix.
Let $(Y,X)^\top$ be a random vector in a two-dimensional Euclidean space, where $Y$ is an observed real-valued response variable and $X$ is an unobservable error-free explanatory variable. Consider the following regression model:
where $m(X)=\mathbb{E}[Y|X]$ is the regression function of $Y$ given $X$ and $U$ is the unpredictable part of $Y$. While the true explanatory variable $X$ is not directly observed, we instead observe a noisy variable $W$ through
where the measurement error $\epsilon\in\mathbb{R}$ is assumed to be independent of $X$ and $Y$. In addition, we assume that the density function of $\epsilon$ is known and denoted by $f_\epsilon$. The case of unknown $f_\epsilon$ but with repeated measurements available is discussed in Section (ref). Suppose researchers consider using the following parametric regression model:
where $g(X;\theta)$ is the parametric specification for the regression function $m(X)$ and $e(\theta)$ is the parametric disturbance of the model. In both econometrics and statistics, it is crucial to assess the adequacy of the putative parametric model $g(X;\theta)$. That is, our null hypothesis of interest is
The alternative hypothesis is the negation of $H_0$, i.e.,
Clearly, testing the null hypothesis in (ref) is equivalent to testing
Note that (ref) is a standard conditional moment restriction. It is well known that (ref) can be equivalently expressed as a continuum number of unconditional moment restrictions using the exponential weight function, see, for example, bierens1982consistent, bierens1990consistent, and bierens1997asymptotic. Specifically, we can rewrite (ref) as follows:
where $e(\theta)=Y-g(X;\theta)$ is the parametric error, $\Pi$ is a properly chosen compact set with nonempty interior, and ${\rm i}=\sqrt{-1}$ denotes the imaginary unit.
To provide evidence of model misspecification, we can then compare a suitable estimator for $S(\xi,\theta_0)$ in (ref) with the zero function. Unfortunately, the classical residual-marked empirical process is clearly infeasible, as the true regressor $X_i$ is unobservable due to the measurement error $\epsilon$ in (ref). Motivated by dong2022nonparametric, however, $S(\xi,\theta_0)$ can still be consistently estimated by noting that
where $f_{Y,X}(y,x)$ denotes the unknown joint density function of $\left(Y,X\right)^\top$. Given a random sample $\{(Y_i,W_i)^\top\}_{i=1}^n$ of size $n\geq 1$, and motivated by the deconvolution methods developed for density estimation [see, e.g., carroll1988optimal and stefanski1990deconvolving], we replace $f_{Y,X}(y,x)$ by a nonparametric estimator and employ a consistent estimator $\hat{\theta}_n$ for $\theta_0$. Then, $S(\xi,\theta_0)$ is estimated by
where $\hat f_{Y,X}(y,x)$ is the deconvolution kernel density estimator for $f_{Y,X}(y,x)$, i.e.,
with $K_b(a)=K(a)/b$, where $K(\cdot)$ is a symmetric kernel function and $b=b_n\in\mathbb{R}^+$ is a sequence of bandwidth parameters shrinking to zero at an appropriate rate specified later. In addition, $\mathcal{K}_b(a)$ is the univariate deconvolution kernel given by
in which $K^{\text{ft}}(\cdot)$ and $f_\epsilon^{\text{ft}}(\cdot)$ are the Fourier transforms of the kernel function $K(\cdot)$ and the measurement error density $f_\epsilon(\cdot)$, respectively. The above deconvolution kernel density estimator has been widely used to construct density and nonparametric regression estimators in the presence of measurement errors; see, e.g., alexander2009deconvolution. Such an approach is particularly important when kernel smoothers are employed, as shown in otsu2021specification. We note that, following this approach, $S_{n}(\xi,\hat{\theta}_n)$ can be further rewritten as
It can be shown that under the null hypothesis and mild regularity conditions on the smoothness of $g(x;\theta)$ with respect to $\theta$, the smoothness of $f_\epsilon$ and kernel function $K(\cdot)$, the $\sqrt n$-consistency assumption on the estimator $\hat\theta_n$ as well as the bandwidth $b$,
uniformly in $\xi\in\Pi$, where,
with $f_X(x)$ denoting the density function of $X$. Here and below, a dot denotes differentiation with respect to the variable $\theta$, i.e., $\dot{g}(x;\theta)=\partial g(x;\theta)/\partial\theta$.
However, the uniform asymptotic decomposition obtained in (ref) requires the $\sqrt n$-consistency property of $\hat\theta_n$, which may not be satisfied in the measurement error context. It is known that $\hat\theta_n$ may have a lower convergence rate, for example, $\hat\theta_n-\theta_0=O_p(n^{\delta-1/2})$ for some $0\leq\delta<1/4$ in parametric models with measurement errors, see taupin1998estimation. Even $\hat\theta_n$ is $\sqrt n$-consistent, due to the presence of $\sqrt{n}(\hat{\theta}_n-\theta_0)^\top G(\xi,\theta_0)$ (commonly known as the parameter estimation effect) in the decomposition of $\sqrt{n}S_n(\xi,\hat{\theta}_n)$, the asymptotic null distribution of $\sqrt{n}S_n(\xi,\hat{\theta}_n)$ depends on $\hat\theta_n$ and the computationally attractive multiplier bootstrap is also infeasible. A typical way to deal with the parameter estimation effect is to impose the asymptotically linear representation assumption of $\sqrt n(\hat\theta_n-\theta_0)$ as follows:
for some function $l(Y_i,W_i;\theta_0)$ with zero mean and finite variance. However, such an asymptotically linear representation is even more difficult to obtain in the context of measurement errors.
In this section, we propose a class of orthogonal projection-based test statistics with attractive theoretical and empirical properties. This approach has been used in the literature to address parameter estimation effects in various testing problems; for example, specification analysis of linear quantile models in escanciano2014specification, specification tests for the propensity score in sant2019specification, model checking in partially linear spatial autoregressive models in yang2024model, and testing for the nonparametric component in partially linear quantile regression models in song2025unified. It is worth emphasizing that the projection approach proposed here is particularly suitable for regression models with measurement errors, as it does not require assuming an asymptotically linear representation of $\sqrt n(\hat\theta_n-\theta_0)$, or even the $\sqrt n$-consistency of $\hat\theta_n$ to $\theta_0$. Specifically, we introduce a projection-based weight to eliminate the parameter estimation effect in (ref), thereby relaxing the requirement on the estimator $\hat\theta_n$ and preserving the parametric convergence advantage of our ICM-type tests relative to local-smoothing-based methods. To be precise, $n^{\delta}(\hat\theta_n-\theta_0)$ for some $\delta>1/4$ would suffice. Consequently, a computationally attractive multiplier bootstrap procedure is facilitated, which is particularly advantageous in the presence of measurement errors.
To motivate our approach, we consider a projection-based transformation of the weight function ${\rm e}^{{\rm i} x\xi}$ given by
with
A similar modification of the weight function can also be found in escanciano2014specification and sant2019specification, among others. The intuition behind (ref) is that $\Delta^{-1} \left( \theta \right)G\left(\xi,\theta \right) $ represents the vector of linear projection coefficients of regressing ${\rm e}^{{\rm i} X\xi} $ on $\dot g(X;\theta)$. Thus, it follows that $\dot g(X;\theta )^\top\Delta^{-1} \left( \theta \right) G\left(\xi,\theta \right) $ is the best linear predictor of ${\rm e}^{{\rm i} X\xi} $ given $\dot g(X;\theta )$ and
Based on the properties mentioned above, our tests are then based on continuous functionals of the following feasible projection-based deconvoluted residual-marked empirical process,
where $\hat{\theta}_n$ is a suitably consistent estimator for $\theta_{0}$ under the null hypothesis with required convergence rates specified later (not necessarily $\sqrt n$-consistent), and $\mathcal{P}_n(x;\xi,\theta)$ is the sample analog of projection $\mathcal{P}(x;\xi,\theta)$ in ((ref)); namely,
where
with
the deconvolution kernel density estimator for $f_X(x)$, where $\mathcal{K}_b(\cdot)$ is defined in (ref).
With the assistance of (ref), we can establish that under the null hypothesis of correct specification, the empirical process $S_{n}^{pro}(\cdot,\hat{\theta}_n)$ is expected to be close to the zero function. Under the alternative hypothesis of misspecification, $S_{n}^{pro}(\cdot,\hat{\theta}_n)$ tends to deviate from the zero function, and consequently $\sqrt nS_{n}^{pro}(\cdot,\hat{\theta}_n)$ diverges. It is therefore natural to construct the test statistic by measuring an appropriate distance between $S^{pro}_{n}(\cdot,\hat{\theta}_n)$ and the zero function, denoted by $\Gamma(S^{pro}_{n})$ for some continuous functional $\Gamma(\cdot)$. For example, we could consider the following two test statistics based on the popular Kolmogorov--Smirnov (KS) and Cram\'{e}r--von Mises (CvM) functionals, respectively,
In constructing the $CvM_{n}$ statistic, we adopt the uniform integrating measure on $\Pi$, which is also recommended in dong2022nonparametric. Previous studies have discussed the use of integrating measures that are absolutely continuous with respect to the Lebesgue measure on $\Pi$. However, due to the difficulty of selecting an appropriate measure and the estimation challenges that arise when the true regressor $X$ is unobserved, such measures have rarely been employed in models with measurement errors. The test statistics $KS_{n}$ and $CvM_{n}$ should be small if the null hypothesis (ref) is true, while \textquotedblleft large\textquotedblright\ values of $KS_{n}$ and $CvM_{n}$ imply the rejection of $H_{0}$. These \textquotedblleft large\textquotedblright\ values will be determined by a convenient multiplier bootstrap procedure described in Section (ref), thanks to the projection $\mathcal{P}_n(x;\xi,\hat{\theta}_n)$ that eliminates the parameter estimation effect.
In this section, we establish the asymptotic distributions of the test statistics $KS_n$ and $CvM_n$ under the null hypothesis $H_0$. Subsequently, we investigate the asymptotic power behavior of these statistics under a sequence of local alternative hypotheses $H_{1n}$ that converge to $H_0$ at the parametric rate $n^{-1/2}$, as well as under the fixed alternative. Denote $f^{ft}(t)=\int {\rm e}^{{\rm i} tx}f(x)dx$ for a generic function $f$ and let $g^{(p)}(x;\theta)$ denote the $p$-times derivative of function $g$ with respect to variable $x$. Throughout this paper, $|c|$ is used to denote the Euclidean norm of a vector $c$ and $|A|$ is used to denote the Frobenius norm of a matrix $A$.
We first impose some regularity conditions to derive the asymptotic null distributions of the test statistics $KS_n$ and $CvM_n$.
Assumption (ref)({\romannumeral1}) requires random sampling and the existence of the second moment of $Y$, both of which are frequently mentioned in the literature. We emphasize that Assumption (ref)({\romannumeral2}) relaxes the requirement on the convergence rate of the estimator $\hat\theta_n$ used in our tests, making it more suitable for cases involving measurement errors. Most estimators cannot typically be expanded as (ref) detailed in Section (ref) in the presence of measurement errors. Even worse, they cannot achieve $\sqrt{n}$-rate convergence without imposing numerous restrictions [see taupin1998estimation], making the multiplier bootstrap infeasible. Assumption (ref)({\romannumeral3}) is common and essential, as mentioned in the classical literature on measurement error [see otsu2021specification].
As shown in the classical measurement error literature, we categorize two separate cases by the decay rate of the tail of the characteristic function of the measurement errors: the ordinary smooth case and the supersmooth case. We first focus on the ordinary smooth case and impose the following assumptions.
Assumption (ref)({\romannumeral1}) imposes both requirements of the Lipschitz continuity and restrictions on the smoothness of the structural functions, mainly adopted from dong2022nonparametric. The Lipschitz continuity assumptions are necessary to analyze the moment properties of structural functions under ordinary smooth conditions, as detailed in the proofs of the lemmas. Our assumptions are slightly stronger, requiring the Lipschitz continuity of $h^{(p)}$ and $[gh]^{(p)}$, enhancements that are essential for ensuring convergence. These assumptions hold in commonly used scenarios, such as polynomial $g$, uniformly distributed $X$, normally distributed $U$, and Laplace distributed $\epsilon$. Assumption (ref)({\romannumeral2}) is commonly known as the ordinary smooth assumption, which specifies the polynomial decay rate of $f_\epsilon^{ft}$, the characteristic function of measurement error. As shown in fan1995average and dong2022nonparametric, it can be generalized to $f_\epsilon^{ft}(t)=\exp(it\zeta)/(c_0^{os}+c_{1}^{os}t+\cdots+c_{\alpha}^{os}t^{\alpha})$, including typical distributions such as the Laplace distribution. The characteristic function of the measurement error is estimated through repeated sampling when it is unknown, as outlined in Section (ref). Assumption (ref)({\romannumeral3}) concerns the order of the kernel function, which is essential for constructing the deconvolution kernel, ensuring its integration properties, and eliminating the asymptotic bias introduced by nonparametric estimators. Moreover, the construction of higher-order kernel functions is feasible as discussed in alexander2009deconvolution. Assumption (ref)({\romannumeral4}) contains the undersmoothing condition frequently used in the literature. It also guarantees the asymptotic negligibility of the bias from nonparametric estimators and establishes the integration properties of the deconvolution kernel. It is worth noting that, in the context of the specification testing problem considered in this paper, we do not impose a lower bound on the bandwidth to ensure the existence of the variance, as is commonly done in studies using kernel-based estimators. In particular, we verify in Appendix (ref) that, under the Lipschitz continuity of the underlying functions and an undersmoothing bandwidth condition, the use of kernel smoothing does not introduce any additional non-negligible variance. Consequently, it suffices to assume the boundedness of the asymptotic variance to guarantee the existence of the limiting process, as stated in Assumption (ref)({\romannumeral5}). Similar assumptions are made in fan1995average and dong2022nonparametric.
Using the assumptions above, we can characterize the limiting behavior of $S_{n}^{pro}(\cdot,\hat\theta_n)$ for the ordinary smooth case. Let \textquotedblleft $\Longrightarrow$ \textquotedblright denote weak convergence on $(l^{\infty}(\Pi),\mathcal{B}_{\infty})$ in the sense of Hoffmann--J{\o}rgensen, where $\mathcal{B}_{\infty}$ denotes the corresponding Borel $\sigma$-algebra, see, e.g., Definition 1.3.3 in van1996weak. Subsequently, we derive the asymptotic null distributions of the statistics $KS_n$ and $CvM_n$, as stated in the following theorem and corollary.
Based on Theorem (ref) and the continuous mapping theorem, see, e.g., van1996weak, we can further derive the asymptotic null distributions of $KS_n$ and $CvM_n$ for the ordinary smooth case.
Theorem (ref) shows that $S_{n}^{pro}(\cdot,\hat\theta_n)$ converges weakly to a centered Gaussian process with its covariance structure depending on the data-generating process (DGP) and the parametric model $g(x;\theta_0)$ under the null. Subsequently, Corollary (ref) shows that the $KS_n$ and $CvM_n$ statistics converge to the sup norm and the squared norm of the aforementioned Gaussian process, respectively. We note that $S_{n}^{pro}(\cdot,\hat\theta_n)$ exhibits $\sqrt{n}$-rate convergence for the ordinary smooth case and is asymptotically independent of the chosen bandwidth $b$, improving the convergence results of local smoothing tests.
For the supersmooth case where the characteristic function of measurement error decays at an exponential rate, e.g., the distribution of $\epsilon$ is normal, regularity conditions underlying the derivation of the asymptotic distribution of our proposed test statistics under the null are given as follows.
Assumption (ref)({\romannumeral1}) requires structural functions and the distribution of latent variables to be infinitely smooth. Although assumptions about conditional mean functions are restrictive, commonly used functions, such as polynomials, circular functions, exponentials, and sums or products of such functions, satisfy the requirement. Additionally, the construction of the class of infinitely differentiable density functions is also mentioned in alexander2009deconvolution. Assumption (ref)({\romannumeral2}) is a supersmooth condition. Assuming the function $f^{ft}_\epsilon$ has an exponential decay rate further enhances Assumption (ref)({\romannumeral2}), with the Gaussian distribution being a typical example. The supersmoothness of measurement error and structural functions necessitate the use of an infinite-order kernel in Assumption (ref)({\romannumeral3}) and integral properties to derive the asymptotic behavior of the statistic. Assumption (ref)({\romannumeral4}) is simply a trivial bandwidth requirement, because the infinite-order smoothness condition removes the bias of the proposed empirical process under the null hypothesis, so that the undersmoothing condition required in Assumption (ref)({\romannumeral4}) is no longer needed. Assumption (ref)({\romannumeral5}) pertains to the boundedness of the asymptotic variance of our statistics, analogous to Assumption (ref)({\romannumeral5}).
Using these assumptions, we derive the asymptotic behavior of the empirical process $S_{n}^{pro}(\cdot,\hat\theta_n)$ for the supersmooth case in the following theorem.
Based on the theorem above, we use the continuous mapping theorem to derive the asymptotic null distributions of $KS_n$ and $CvM_n$ for the supersmooth case.
The projection plays a crucial role in Corollaries (ref) and (ref). By eliminating the effect of the estimator $\hat\theta_n$, the proposed test statistics achieve a parametric rate of convergence, even if the convergence rate of $\hat\theta_n$ is typically slower than $\sqrt{n}$ in the presence of measurement errors.
In this section, we proceed to derive the power properties of the proposed tests against a sequence of local alternatives. In contrast to otsu2021specification, where the power of the local smoothing specification test is analyzed under local alternatives converging to the null hypothesis at a nonparametric rate, our tests have nontrivial power against local alternatives converging at a parametric rate under the following form:
where $\Delta:\mathbb{R}\to\mathbb{R}$ is a bounded nonzero function satisfying the regularity conditions of the function $h$ mentioned in Assumptions (ref) and (ref). The asymptotic local power properties of $S_{n}^{pro}(\cdot,\hat\theta_n)$ under $H_{1n}$ are established in the following theorem.
Theorem (ref) shows that under $H_{1n}$, both for the ordinary smooth and the supersmooth measurement errors, the asymptotic behavior of the $S_{n}^{pro}(\cdot,\hat\theta_n)$ consists of the limiting Gaussian process under $H_0$ and a deterministic shift term $\mu_\Delta(\cdot,\theta_0)$. By similar arguments to those of the null hypothesis, we then apply the continuous mapping theorem to obtain the convergence in distribution of $KS_n$ and $CvM_n$, which will have nontrivial power against local alternatives as described above, provided that $\mu_\Delta(\xi,\theta_0)\neq 0$ for some $\xi$.
By denoting $\theta^\ast = \operatorname*{plim}\hat{\theta}_n$ as the pseudo-true value and noting that $\theta^\ast = \theta_0$ under the null or the local alternatives, we can derive the asymptotic global power properties for our test statistics under the fixed alternative.
Note that $C(\cdot,\theta^\ast)$ can be understood as the projection of $m(X)-g(X;\theta^\ast)$ onto the orthogonal space of $\dot{g}(X;\theta^\ast)$, as explained in dominguez2015simple. Under this interpretation, the consistency of our tests is guaranteed as long as for any vector $\gamma$,
This ensures $C(\xi,\theta^\ast)\neq 0$ for at least some $\xi$ with a positive Lebesgue measure, thereby leading to the fact that $\sqrt nS_{n}^{pro}(\cdot,\hat{\theta}_n)$ diverges asymptotically and thus $\sqrt nKS_n$ and $nCvM_n$ will diverge to positive infinity in probability. This indicates that any fixed alternative satisfying (ref) can be detected by our proposed test with probability tending to one. Although the tests are not consistent against all possible alternatives, we do not regard those alternatives that violate this condition as a primary empirical concern, since they are nearly observationally equivalent to the null.
Since it is often challenging for researchers to obtain the density function of the measurement error in practical applications, we develop specification tests for settings with an unknown measurement error. As stated in delaigle2008deconvolution and alexander2009deconvolution, the approach typically used to address this problem is based on additional data, more specifically, repeated measurements on $X$ in the form of
where $\epsilon^r$ and $\epsilon$ are identically distributed and $(X,\epsilon,\epsilon^r)$ are mutually independent. Repeated measurements are used to construct the following consistent estimator for $f^{ft}_\epsilon$,
as proposed by delaigle2008deconvolution and recommended in otsu2021specification.
Our main contribution is to establish detailed theoretical results for ICM-type specification tests based on the repeated measurements approach described above. The theoretical justification and the practical implementation of the multiplier bootstrap for obtaining critical values are then developed and discussed in Section (ref). Specifically, we construct the following empirical process based on the aforementioned estimator,
where $\hat{\mathcal{K}}_b(\cdot)$ and $\hat{\mathcal{P}}_n(\cdot)$ denote the respective counterparts of $\mathcal{K}_b(\cdot)$ and $\mathcal{P}_n(\cdot)$ by replacing $f^{ft}_\epsilon(\cdot)$ with $\hat{f}^{ft}_\epsilon(\cdot)$. Using the empirical process $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ constructed above, the test statistics $CvM_n$ and $KS_n$ introduced earlier can be modified as
Furthermore, to derive the asymptotic properties of the modified statistics, we need to strengthen the assumptions to address the uncertainty introduced by the estimation of the error characteristic functions, by imposing the following additional assumptions in addition to Assumption (ref).
Assumption (ref)({\romannumeral1}) is common in literature, see, e.g., delaigle2008deconvolution, requiring independent and identically distributed repeated measurements. Assumption (ref)({\romannumeral2}) is used to evaluate the asymptotic properties of the estimator $\hat{f}^{ft}_\epsilon(t)$ in (ref), as mentioned in kurisu2022uniform. To derive the asymptotic distribution of our statistics under the ordinary smooth case, we need to impose the following assumptions.
In Assumption (ref)({\romannumeral1}), we enhance the assumption about bandwidth compared to Assumption (ref)({\romannumeral4}). This enhancement is necessary because in the absence of information about the measurement error distribution, the statistic is influenced by the uncertainty in estimating $f^{ft}_\epsilon$. Assumption (ref)({\romannumeral1}) ensures the asymptotic negligibility of such uncertainty brought by the estimator $\hat{f}^{ft}_\epsilon$. Next, compared to our Assumptions in (ref)({\romannumeral5}), we add a moment restriction on the term brought by the estimator of the unknown distribution in Assumption (ref)({\romannumeral2}).
With the assumptions above, the following theorem, which characterizes the asymptotic behavior of the proposed empirical process with the unknown characteristic function replaced by its estimator, is established under the ordinary smooth case.
Using the continuous mapping theorem, we can readily derive the asymptotic null distributions of $\widehat{KS}_n$ and $\widehat{CvM}_n$ for the ordinary smooth case.
Theorem (ref) shows that for the ordinary smooth case when the distribution of measurement error is unknown, $\hat{S}^{pro}_n(\cdot,\hat{\theta}_n)$ still converges at the $\sqrt{n}$-rate to a centered Gaussian process. But it is worth noting that the Gaussian process here is different from the one mentioned in Theorem (ref), because the uncertainty brought by the $\hat{f}_{\epsilon}^{ft}$ (represented by the term $ r^{\epsilon}_{\infty}(Y,W;\xi,\theta_0)$) changes the covariance structure of the limiting process.
For the supersmooth case, the following assumptions are necessary in addition to Assumption (ref) to derive asymptotic properties of our statistics.
Similar to Assumption (ref), Assumption (ref)({\romannumeral1}) requires a stronger assumption about the bandwidth $b$ for ensuring the asymptotic negligibility of uncertainty brought by the estimation of the error characteristic functions, while Assumption (ref)({\romannumeral2}) ensures that the asymptotic variance of $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ is bounded through an additional moment restriction based on Assumption (ref)({\romannumeral5}).
The following theorem characterizes the asymptotic behavior of $\hat{S}_{n}^{pro}(\cdot,\hat\theta_n)$ for the supersmooth case under the null hypothesis.
Theorem (ref) shows that for the supersmooth case, our empirical process still maintains the $\sqrt{n}$-convergence. Furthermore, Corollary (ref) shows convergence in distribution of $\sqrt n\widehat{KS}_n$ and $n\widehat{CvM}_n$ under the null hypothesis. The subsequent analysis considers the results under the sequence of local alternatives and the fixed alternative, thereby establishing the asymptotic power properties of the tests without requiring knowledge of the measurement error distribution.
Under the sequence of local alternative hypotheses defined in (ref), we still have a deterministic shift term similar to Theorem (ref), signifying the nontrivial local power of our tests. We also find that under the alternative hypothesis $H_1$, our proposed specification tests share the same consistency condition as in the case where the measurement error distribution is known, see condition (ref).
As we show in Sections (ref) and (ref), the null limiting distributions of test statistics $KS_{n}$ and $CvM_{n}$ usually depend on the underlying DGP in a rather complicated manner. As such, the associated critical values are case-dependent, forcing people to resort to bootstrap procedures. Unfortunately, although the wild bootstrap procedure is typically employed to implement tests in the error-free case, no computationally easy and valid bootstrap procedures are currently available for tests in the presence of measurement errors, as indicated above. In particular, any residual-based bootstrap methods (e.g., the widely used wild bootstrap in the literature of specification tests) are apparently infeasible in the presence of measurement errors, given that we need the true regressors to construct the residuals, but we cannot observe the regressors directly.
For the influential local smoothing specification test proposed in otsu2021specification, a multi-step estimation and resampling method is employed to operationalize the bootstrap procedure, a strategy also adopted in classical studies such as hall2007testing for global smoothing tests. Specifically, they first estimate the density function of the unobservable true regressor $X$, the fitted nonparametric regression function, and the parametric fit using standard deconvolution techniques and a parametric estimator. Then, the residuals are resampled by estimating their second moments and drawing from distributions such as the two-point distribution of mammen1993bootstrap. Based on these resampled residuals, bootstrap samples of $Y$ are generated as the sum of the resampled residuals and the estimated parametric fit evaluated at the resampled regressors drawn from the estimated density of $X$. The measurement error–contaminated bootstrap samples of variable $W$ are subsequently obtained by adding the true regressor and an independent draw from the known (or estimated) measurement error distribution. Finally, the bootstrap counterparts of the test statistics are computed using each resampled pair $\{(Y_i^\ast,W_i^\ast)^\top\}_{i=1}^n$. Furthermore, bandwidth selection is also a critical issue for local smoothing tests due to the sensitivity of local smoothers to the bandwidth choice, which is addressed by adopting the \textquotedblleft two-stage selection plug-in\textquotedblright method of delaigle2004practical in the test of otsu2021specification.
While the above approach is feasible, it still has several limitations that deserve further attention. First, the complex estimation and resampling steps complicate both theoretical justification and empirical implementation. In particular, integrals involving deconvolution kernels often lack closed-form solutions for many higher-order kernel functions, and numerical approximation may reduce accuracy. Second, within each simulation, bootstrap iteration requires repeating the estimation of functions and the computation of the bootstrap counterparts of the test statistics, making the overall procedure computationally inefficient. Third, the need to select suitable bandwidths adds further computational complexity and theoretical difficulties. In our proposed tests, we eliminate the parameter estimation effect via the projection approach, thereby facilitating a much simpler multiplier bootstrap and avoiding the aforementioned problems. This serves as one of the main contributions of the study.
For the cases with known measurement error distribution, we approximate the asymptotic null distribution of a continuous functional $\Gamma(S_n^{pro})$ by that of $\Gamma(S_n^{pro,\ast})$, where $\Gamma(\cdot)$ represents the sup norm and the squared norm employed in constructing the statistics $KS_{n}$ and $CvM_{n}$, respectively, and
Here, $\{V_i\}_{i=1}^n$ is a sequence of i.i.d. random variables with zero mean, unit variance, and independent of the original sample, known as the multipliers. A popular example is i.i.d. Bernoulli variables constructed by mammen1993bootstrap. Both for the ordinary smooth case and the supersmooth case, $S_{n}^{pro,\ast}(\cdot,\hat\theta_n)$ is expected to have the same limiting process as $S_{n}^{pro}(\cdot,\hat\theta_n)$ under the null hypothesis, due to the zero mean and unit variance properties of the multipliers. Furthermore, we expect that the multipliers can eliminate the deterministic shift term mentioned in Theorem (ref) under the local alternatives, thus establishing that $S_{n}^{pro,\ast}(\cdot,\hat\theta_n)$ converges to the same limiting process as under the null hypothesis.
Defining \textquotedblleft $\overset{\ast}{\Longrightarrow}$\textquotedblright as weak convergence and $\mathbb{P}_n^\ast$ as the bootstrap probability under the bootstrap law, i.e., conditional on the original sample, see, e.g., Section 2.9 of van1996weak, the validity of the proposed multiplier bootstrap is formally established by the following theorem.
For the ordinary smooth and the supersmooth cases, respectively, the bootstrap empirical processes share the same limiting behavior under $H_0$ and $H_{1n}$ as their original sample counterparts under the null, while the latter exhibit a deterministic shift under $H_{1n}$. As a consequence of the above analysis, the asymptotic critical value at the significance level $\alpha$ is $c^\ast_{\alpha}=\inf\{c_\alpha\in[0,\infty):\lim_{n\to\infty}\mathbb{P}_n^\ast\{\Gamma(\sqrt n S_{n}^{pro,\ast})>c_\alpha\}=\alpha\}$. In practice, $c^\ast_{\alpha}$ can be approximated as $c^\ast_{n,\alpha} = \{\Gamma(\sqrt n S_{n}^{pro,\ast})\}_{B(1-\alpha)}$, the $B(1-\alpha)$-th order statistic for $B$ replicates $\{\Gamma(\sqrt n S_{n,b}^{pro,\ast})\}_{b=1}^B$ of $\{\Gamma(\sqrt n S_{n}^{pro,\ast})\}$ and we reject $H_0$ if $\Gamma(\sqrt n S_{n}^{pro})>c^\ast_{n,\alpha}$. Specifically, building upon Theorem (ref) and taking $\Gamma(\cdot)$ to be the sup norm and the squared norm, respectively, we establish the asymptotic validity of the multiplier bootstrap procedure for the $KS_n$ and $CvM_n$ statistics.
As we mentioned in Section (ref), in the absence of information about the distribution of measurement errors, we argue that our multiplier bootstrap procedure remains feasible. We still use repeated measurements to estimate the characteristic function of the measurement error, thereby estimating the deconvolution kernel as shown in (ref). Then we can construct the multiplier bootstrap version of $\hat{S}_n^{pro}(\cdot,\hat{\theta}_n)$ as given by
Similar to what we have shown in Theorem (ref), we establish the validity of our bootstrap procedure by the following theorem for the unknown measurement error case.
Theorem (ref) shows that the bootstrap version of the empirical process constructed by repeated measurements converges to the same limiting process under $H_0$ and $H_{1n}$, which is the same limiting process as the one under the null hypothesis mentioned in Section (ref), for the ordinary smooth and super smooth cases, respectively. Consequently, we can validate the multiplier bootstrap for $\widehat{KS}_n$ and $\widehat{CvM}_n$ using the continuous mapping theorem and obtain the critical values as discussed in Subsection (ref). Therefore, using the multiplier bootstrap to obtain the critical values for our tests remains feasible even when the measurement error is unknown.
In this section, we investigate the finite-sample performance of the proposed test by Monte Carlo experiments. We set the unobservable regressors $\{X_i\}_{i=1}^n$ distributed as $N(0,1)$, and we use a trigonometric model and a polynomial model in our simulation, although we note that any nonlinear model can be used in our test procedure. We name $Y_i=1+X_i+\delta X_i^2+U_i$ as model one and $Y_i=1+X_i+\delta \cos(\pi X_i)+U_i$ as model two, respectively, for $i$ ranging from $1$ to $n$, where $\delta$ is a variable constant representing the nonlinearity of the model and is set to $0.5$. Additionally, $U_i\sim N(0,1/4)$. The contaminated regressor is given by $W_i=X_i+\epsilon_i$, where we use the Laplace distribution with variance of $1/12$ for the ordinary smooth case and $N(0,1/12)$ for the supersmooth case to remain consistent with the experiment in otsu2021specification. We use the infinite-order flat-top kernel proposed by mcmurry2004nonparametric, which has also been employed in dong2022nonparametric, for both cases and all of our simulations, which is defined by
We choose the bandwidth according to the rules of thumb, see, e.g., delaigle2008deconvolution. Specifically, we use $b=c(5\sigma^4/n)^{1/27}$ for the ordinary smooth case and $b = c(4\sigma^2/\log(n))^{1/2}$ for the supersmooth case, where $\sigma$ is the standard deviation of the measurement error and sensitivity indicator $c$ varies in the grid presented in the following table. For the estimator in our test, we use a polynomial estimator of degree $2$, as mentioned in cheng1998polynomial, which is also consistent with the results in otsu2021specification.
We report the simulation results based on $199$ bootstrap iterations and $1000$ Monte Carlo replications. Tables include the size (the proportion of rejections when the DGP is under the null hypothesis, named DGP$(0)$) and powers (the proportion of rejections when the DGP is under the alternative hypothesis model one, named DGP$(1)$, and model two, named DGP$(2)$). We report the results for two statistics $KS_n$ and $CvM_n$, two sample sizes $n=\{500,1000\}$, three nominal level $\alpha = \{1\%,5\%,10\%\}$ and varied sensitivity indicators. Constrained by the length of papers, only part of the results $c\in\{1,5,10\}$ and $\alpha\in\{5\%,10\%\}$ are shown in the main text, and the complete results are reported in Tables (ref)--(ref) of Appendix (ref).
Tables (ref) and (ref) show our testing results when we know the distribution of the measurement error for the ordinary smooth case and the supersmooth case, respectively. In our results, the $KS_n$ and $CvM_n$ statistics perform well in terms of both size and power. We observe significant improvements in size accuracy and power gain as the sample size increases. We also observe that the size and power of the tests do not fluctuate much with changes in bandwidth values on the grid. Comparing the results between the ordinary smooth case and the supersmooth case, we find that test sizes for the supersmooth cases perform comparably to those for the ordinary smooth cases, which differs from the findings reported in otsu2021specification. It stems from the fact that the statistics for both the ordinary smooth and supersmooth cases exhibit $\sqrt{n}$-rate convergence, another advantage of the proposed tests. Other results in the tables reveal the power properties of our tests, where DGP$(1)$ represents the polynomial model and DGP$(2)$ represents the trigonometric model. It is well known that global smoothing tests perform better in polynomial models, while local smoothing tests perform better in trigonometric models. For our results in Tables (ref) and (ref), the test performance under DGP$(1)$ is better than that under DGP$(2)$. We can expect better performance with higher nonlinearity and more samples.
Tables (ref) and (ref) display size and power properties of the $\widehat{KS}_n$ and $\widehat{CvM}_n$ statistics in the absence of information about the distribution of the measurement error. However, because our asymptotic variance is affected by the estimation error of the characteristic functions, the results for size and power in finite samples are slightly worse but remain within the acceptable range. We can also observe that our tests perform better under the polynomial model than under the trigonometric model. Meanwhile, our tests still retain robustness to bandwidth selection when the measurement error distribution is unknown.
Finally, we summarize the advantages of our tests, as demonstrated in the simulations, including a good level of accuracy and power properties due to the $\sqrt{n}$-convergence of the statistics, better efficiency for polynomials and low-frequency alternative hypotheses, and robustness in bandwidth selection. Additionally, our testing framework naturally accommodates the case without measurement error.
In this paper, we propose new ICM-type specification tests for regression models with measurement errors based on a deconvoluted residual-marked empirical process. The innovation of our method is the use of a projection to eliminate the parameter estimation effect, a particularly attractive tool in the presence of measurement errors. We formally establish the asymptotic properties of the projection-based test statistics and suggest a straightforward multiplier bootstrap procedure to simulate the critical values. Tests and the multiplier bootstrap procedure are also examined when the measurement error distribution is unknown. We emphasize three main advantages of the tests proposed in this paper. From a theoretical perspective, the statistics of $\sqrt{n}$-convergence for both the ordinary smooth and the supersmooth measurement errors ensure excellent size accuracy and satisfactory power properties, especially for the supersmooth cases, where slower convergence rates can typically be achieved, as mentioned in the previous literature on measurement error. From the application perspective, our proposed method enables a general convergence rate of the parametric estimator, thereby resolving the difficulty in obtaining a $\sqrt{n}$-convergence estimator in the presence of measurement errors. Additionally, our tests are more robust to bandwidth selection and, therefore, more reliable. From a computational perspective, our projection approach enables us to simulate the critical values as accurately as desired via the multiplier bootstrap, which is easier to understand and compute.
For future research, we anticipate that the valuable combination of projection and multiplier bootstrap can be extended to other interesting testing problems in the presence of measurement errors, such as tests for heteroskedasticity and significance tests.
In this appendix, we provide additional simulation results and proofs of the main theoretical results.