Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
A Consistent ICM-based $^2$ Specification Test
\affil[1]{Department of Statistics and Data Science, Fudan University, Shanghai, China}
\affil[2]{Department of Economics, Finance and Legal Studies, Culverhouse College of Business, University of Alabama}
center[center omitted — 71 chars of source]
refsection\begin{abstract}
In spite of the omnibus property of Integrated Conditional Moment (ICM) specification tests, they are not commonly used in empirical practice owing to, e.g., the non-pivotality of the test and the high computational cost of available bootstrap schemes, especially in large samples. This paper proposes specification and mean independence tests based on ICM metrics. The proposed test exhibits consistency, asymptotic $\chi^2$-distribution under the null hypothesis, and computational efficiency. Moreover, it demonstrates robustness to heteroskedasticity of unknown form and can be adapted to enhance power towards specific alternatives. A power comparison with classical bootstrap-based ICM tests using Bahadur slopes is also provided. Monte Carlo simulations are conducted to showcase the excellent size control and competitive power of the proposed test.
Keywords: specification test, mean independence, omnibus, pivotal
JEL classification: C12, C21, C52
\end{abstract}
\section{Introduction}
Model misspecification is a major source of misleading inference in empirical work. This issue is further compounded when various competing models are available. It is thus imperative that model-based statistical inference be accompanied by proper model checks such as specification tests stute1997nonparametric.
Existing tests in the specification testing literature can be categorized into three classes, namely, conditional moment (CM) tests, non-parametric tests, and integrated conditional moment (ICM) tests. The class of CM tests, such as those proposed by newey1985maximum,tauchen1985diagnostic, is not consistent as it relies on only a finite number of moment conditions implied by the null hypothesis bierens1990consistent.
The class of non-parametric tests is therefore proposed as a remedy, see e.g. wooldridge1992test,yatchew1992nonparametric,hardle1993comparing,hong1995consistent,zheng1996consistent,li1998simple,fan2000consistent,su2007consistent,li2022learning. The key idea of these test statistics is to nonparametrically estimate the conditional moments -- such as through local smoothing techniques -- and then compare them to their parametric counterparts under the null hypothesis.
The class of non-parametric tests may encounter challenges such as non-parametric smoothing and suboptimal performance stemming from over-fitting the non-parametric alternative. In contrast, the class of ICM tests, such as those introduced by bierens1982consistent,bierens1990consistent,delgado1993testing,bierens1997asymptotic,stute1997nonparametric,delgado2006consistent,escanciano2006consistent,dominguez2015simple,su2017martingale,antoine2022identification, has gained popularity due to its ability to avoid these issues and detect local alternatives at faster rates. ICM metrics, on which ICM tests are based, also appear in other contexts: martingale difference hypothesis tests escanciano2009lack, joint coefficient and specification tests antoine2022identification, model-free feature screening zhu2011model, shao2014martingale,li2023generalized, model estimation escanciano2018simple,tsyawo2023feasible, specification tests of the propensity score sant2019specification, and tests of the instrumental variable (IV) relevance condition in ICM estimators escanciano2018simple,tsyawo2023feasible.
Despite their advantages, ICM tests are not widely used in empirical research escanciano2009simple,dominguez2015simple. First, ICM test statistics are not pivotal under the null hypothesis, thus critical values cannot be tabulated analytically bierens1997asymptotic,dominguez2015simple. Second, ICM tests—when implemented via the wild bootstrap—tend to be computationally costly, as they require estimations on resampled data to compute $p$-values. { Moreover, the wild bootstrap-based ICM test is arguably unsuitable for limited dependent outcome variable models such as the logit, because bootstrap replicates of the outcome may fail to respect the outcome variable's limited support.} While the more recent multiplier bootstrap approach to ICM specification testing escanciano2009simple,li-song-consistent-2022,escanciano-2024-gaussian offers a substantial computational advantage relative to the wild bootstrap by avoiding model re-estimation, it still entails non-negligible computational complexity due to re-sampling and the removal of the effect of estimating nuisance parameters. Third, although ICM tests are omnibus bierens1982consistent,stute1997nonparametric,dominguez2015simple, they only have substantial local power against alternatives in a finite-dimensional space escanciano2009lack. Moreover, it is not obvious how to leverage prior knowledge of potential directions under the alternative to enhance the power of existing ICM tests.
This paper proposes a consistent $\chi^2$-test for the unified ICM framework of mean independence and specification testing. The key idea is to augment the ICM metric with a user-specified non-degenerate transformation of the conditioning covariates as in CM tests, which removes the first-order degeneracy inherent in classical ICM tests and yields a pivotal test statistic.
The approach accommodates endogenous regressors, instrumental variables (IV), and heteroskedasticity of unknown form in both linear and non-linear models. As an ICM-based test, it is omnibus and capable of detecting a wide range of model misspecifications, including violations of IV exogeneity. Similar to the multiplier bootstrap method escanciano2009simple, our approach is suited for models with possibly non-additively separable errors or non-continuous outcomes.
Compared to existing ICM tests, our test has three advantages. First, it can be implemented as a $\chi^2$- or two-sided $t$-test, which can be interpreted more easily when compared to bootstrap-based tests. Second, the proposed test does not require bootstrap calibration of critical values; hence, it is computationally fast and remains feasible even in very large samples. Third, although our proposed test is not optimal, its power can be enhanced with the knowledge of directions under the alternative or, more generally, with directions the researcher may have in mind, whereas ICM tests lack this property. Therefore, we consider our test as a bridge between CM tests and ICM tests. We also acknowledge that the implementation of our test involves a tuning parameter, which is used in computing the generalized inverse that forms the Wald-type test statistic.
This paper is not the first attempt at circumventing the non-pivotality of ICM test statistics. bierens1982consistent approximates the critical values of the ICM specification test using Chebyshev's inequality for first moments under the null hypothesis, which is subsequently improved upon by bierens1997asymptotic. bierens1982consistent also proposes a $\chi^2$-test based on two estimates of Fourier coefficients and a carefully chosen tuning parameter. Simulation evidence therein shows a high level of sensitivity to the tuning parameter. Besides, estimating Fourier coefficients no longer makes the test ICM as the test statistic is no longer “integrated". Another attempt in the literature is the conditional Monte Carlo approach of hansen1996inference,de1996bierens. These are, however, computationally costly as bierens1997asymptotic notes. {Recent studies have also revisited the goal of constructing pivotal and computationally feasible specification tests within the conditional moment framework. raiola2024testing draws inspiration from the classical Pearson’s $\chi^2$ test by partitioning the sample space into cells. Although this test enjoys pivotality and excellent computational scalability, its power is inherently tied to the chosen partition and is therefore not fully omnibus: smooth model deviations that do not necessarily manifest as changes in constructed cells may go undetected. li2025powerful leverages sample splitting and learns the optimal projection direction in the reproducing kernel Hilbert space via support vector machines (SVM). This means that it can achieve high local power when the learned projection aligns with the alternative; however, sample splitting and hyper-parameter tuning (in both kernels and SVM) may introduce finite-sample variability.}
The rest of the paper is organized as follows. (ref) provides a brief presentation of ICM metrics, a new metric, and its omnibus property. (ref) proposes the $\chi^2$-test statistic and derives its limiting distribution under the null, local, and alternative hypotheses within the unified framework of mean independence and specification tests. Monte Carlo simulations in (ref) compare the empirical size and power of the $ \chi^2 $ specification test to the multiplier and wild bootstrap-based ICM specification tests, and (ref) concludes. All technical proofs and additional simulation results are relegated to the supplementary material.
\paragraph{Notation:} For $a\in\mathbb{R}^p$, we denote its transpose by $a^{\top}$, and its Euclidean norm as $\|a\|$. We denote $\mathrm{i}$ as the imaginary unit which satisfies $\mathrm{i}^2=-1$. “$\overset{p}{\rightarrow}$" and “$\overset{d}{\rightarrow}$" denote convergence in probability and distribution, respectively. Throughout the paper, for a random vector $W$, we denote $W^\dagger$ as its independent and identically distributed ($iid$) copy, and write $\mathbb{E}_n W=n^{-1}\sum_{i=1}^n W_i$ as the empirical mean for $iid$ copies $\{W_i\}_{i=1}^n$ of $W.$ To cut down on notational clutter, $\widetilde{W}$ is sometimes used to denote the centered version of a random variable $W$, i.e., $\widetilde{W}:=W-\mathbb{E} W.$
\section{The Framework}
In this section, we briefly discuss ICM metrics, their omnibus property, and a new omnibus metric that characterizes mean independence.
\subsection{ICM Metrics}
For a random variable $U\in\mathbb{R}$ and a random vector $Z\in \mathbb{R}^{p_z},$ we say $U$ is mean-independent of $Z$, if
\begin{equation}
\mathbb{E}[U \mid Z]=\mathbb{E}[U] \ almost \ surely\ (a.s.),
\end{equation}
otherwise, $U$ is mean-dependent on $Z$. To characterize the relationship (ref), the existing literature puts much effort into studying ICM mean dependence metrics of the form:
\begin{flalign*}
T(U \mid Z;\nu)=\int_{\Pi} \Big|\mathbb{E}[(U-\mathbb{E} U)w(s,Z)]\Big|^2\,\nu(ds)
\end{flalign*}
where $w(s,Z)$ is a weight function indexed by $s \in \Pi$, and $\nu$ is a measure on $\Pi$. A notable feature of $T(U \mid Z;\nu)$ is its omnibus property, namely, $T(U \mid Z;\nu)=0$ if and only if (ref) holds, see e.g., shao2014martingale. The omnibus property guarantees the consistency of ICM tests. Therefore, a larger value of $T(U \mid Z;\nu)$ indicates a stronger mean dependence of $U$ on $Z$.
Under suitable conditions, the ICM metric has the general form
\begin{align*}
T(U \mid Z;\nu)
&= \int_{\Pi} \Big(\mathbb{E}[(U-\mathbb{E} U)w(s,Z)]\Big)\overline{\Big(\mathbb{E}[(U-\mathbb{E} U)w(s,Z)]\Big)}\nu(ds)\\
&= \mathbb{E}\Big[(U-\mathbb{E} U)(U^\dagger-\mathbb{E} U)\int_{\Pi} \big(w(s,Z)\overline{w(s,Z^\dagger)}\big)\nu(ds) \Big]\\
&:= \mathbb{E}\Big[(U-\mathbb{E} U)(U^\dagger-\mathbb{E} U)K(Z,Z^\dagger)\Big]\\
&:= \mathrm{ICM}(U \mid Z),
\end{align*}where $\overline{w(\cdot,\cdot)}$ denotes the complex conjugate of $w(\cdot,\cdot)$ and $K(\cdot,\cdot)$ denotes the kernel. (ref) provides a few examples of commonly used ICM kernels, see li2023generalized for a recent review.
\begin{table}[H]
\caption{Examples of ICM Kernels}
\begin{tabular}{@llll@}
\toprule
$w(s,Z)$ & $\nu(ds)$ & $K(Z,Z^\dagger)$ & Reference \\
\midrule
$\exp( \mathrm{i} Z^\top s)$ & $ \prod_{l=1}^{p_Z}\phi(s_l)ds_l$ & $\exp\left(-0.5 \| Z - Z^\dagger\|^2\right)$ & bierens1982consistent \\
$1(Z \le s)$ & $dF_Z(s)$ & $1 - F_Z(Z \lor Z^\dagger)$ & stute1997nonparametric \\
$\exp( \mathrm{i} Z^\top s)$ & $\big(c_{p_Z}||s||^{1+p_Z}\big)^{-1} ds$ & $ -\| Z-Z^\dagger\| $ & shao2014martingale \\
\bottomrule
\end{tabular}
Notes: $\phi(x) = \frac{1}{\sqrt{2\pi}} e^{-\frac{1}{2}x^2} $ is the standard normal probability density function, and $ c_p = \frac{\pi^{(1+p)/2}}{\Gamma((1+p)/2)}, \ p \geq 1 $ where $ \Gamma(\cdot) $ is the complete gamma function.
\end{table}
Examples of weight functions from the literature include the step function $w(s,Z) = \mathrm{I}(Z\leq s) $, e.g., stute1997nonparametric,dominguez2004consistent,delgado2006consistent,escanciano2006goodness,zhu2011model; a one-dimensional projection in the step function $w(s,Z) =\mathrm{I}(Z^{\top}s_{-1}\leq s_1)$, e.g., escanciano2006consistent,kim2020robust; the real exponential $w(s,Z) = \exp(Z^{\top}s)$, e.g., bierens1990consistent; and the complex exponential $w(s,Z)=\exp(\mathrm{i}Z^{\top}s)$, e.g., bierens1982consistent,shao2014martingale,antoine2022identification. The space $\Pi$ in escanciano2006consistent,kim2020robust is given by $\Pi = \mathbb{R}\times \mathbb{S}_{p_z}$ where $\mathbb{S}_{p_z}$ denotes the space of $p_z\times 1$ vectors with unit Euclidean norm while $\Pi=\mathbb{R}^{p_z}$ for the other works mentioned above. See escanciano2006goodness for a general characterization of ICM weight functions.
The omnibus property of ICM metrics thus translates as
\begin{equation}
\mathrm{ICM}(U \mid Z)=0 \quad if and only if \mathbb{E}[U \mid Z]=\mathbb{E}[U] \ almost \ surely \ (a.s).
\end{equation}
When \( U \) and \( Z \) are observed, a natural empirical estimator of \( \mathrm{ICM}(U \mid Z) \) is given by the (modified) U-statistic
\begin{equation*}
\widehat{\mathrm{ICM}}_n(U \mid Z) = \frac{1}{n(n-1)} \sum_{i \neq j} (U_i - \mathbb{E}_n[U])(U_j - \mathbb{E}_n[U]) K(Z_i, Z_j).
\end{equation*}
As with most existing ICM-based tests, the asymptotic null distribution of the corresponding test statistic (under (ref)) is non-pivotal bierens1997asymptotic:
\begin{equation}
n\widehat{\mathrm{ICM}}_n(U \mid Z) \overset{d}{\rightarrow} \sum_{k=1}^{\infty}\lambda_k G_k^2,
\end{equation} granted $\mathbb{E} [U^2U^{\dagger 2}K^2(Z,Z^{\dagger})]<\infty$. Here, $\{G_k\}_{k=1}^{\infty}$ is a sequence of $iid$ standard Gaussian random variables, and $\{\lambda_k\}_{k=1}^{\infty}, \ \lambda_k\geq 0$, is a sequence of non-increasing coefficients that depend on the distribution of $[U,Z^\top]^\top$ and the kernel $K(\cdot,\cdot)$. Thus, the limiting distribution of \( n\widehat{\mathrm{ICM}}_n(U \mid Z) \) under (ref) is non-pivotal, and re-sampling/bootstrap techniques are usually required to obtain valid \( p \)-values for inference.
\subsection{A new characterization of mean independence}
Consider the following test hypotheses of mean independence:
\begin{flalign}
\begin{split}
\mathbb{H}_o&: \mathbb{E}[U \mid Z] - \mathbb{E}[U] = 0 \ a.s.;\\
\mathbb{H}_a&: \mathbb{P}\big(\mathbb{E}[U \mid Z]=\mathbb{E}[U]\big) < 1.
\end{split}
\end{flalign}
The above hypotheses of interest, in view of the omnibus property (ref), can be restated as testing $ \mathrm{ICM}(U \mid Z)=0.$ Indeed, by the Law of Iterated Expectations (LIE),
\begin{equation}
\begin{split}
\mathrm{ICM}(U \mid Z)=& \mathbb{E}\left\{K(Z,Z^\dagger) (U^\dagger-\mathbb{E} U) \mathbb{E}[(U-\mathbb{E} U) \mid Z,Z^\dagger,U^\dagger]\right\} \\
=& \mathbb{E}\left\{K(Z,Z^\dagger)(\underbrace{\mathbb{E}[U \mid Z]-\mathbb{E} U}_{=0 \, a.s. under \mathbb{H}_o})(U^\dagger-\mathbb{E} U)\right\}.
\end{split}
\end{equation}
Under $\mathbb{H}_o$, $\mathbb{E}[U \mid Z]$ is degenerate, i.e., $\mathbb{E}[U \mid Z]-\mathbb{E} U=0\ a.s.$ This is the first-order degeneracy problem in U/V-statistic-based estimator for (ref) that accounts for the non-pivotal limiting distribution in (ref). In this regard, we propose to replace the role of $U$ in the ICM metric with a random vector $V\in \mathbb{R}^{p_v}$ constructed such that $\mathbb{E}[V \mid Z]$ is non-degenerate under both the null and alternative hypotheses while preserving the omnibus property (ref). In other words, we propose to measure the mean independence of $U$ on $Z$ with the assistance of a random vector $V$ by
$$
\delta(V):= \mathbb{E}\left[K(Z,Z^\dagger) (U^\dagger-\mathbb{E} U) (V - \mathbb{E} V) \right].
$$
The resulting test relies on the empirical estimator of $\delta(V)$.
The two key properties---\textit{first-order non-degeneracy} and \textit{omnibusness}---ensure a pivotal limiting distribution under the null and the consistency of the test, respectively. Specifically, we require (1) $V$ to be mean dependent on $Z$ to rule out first-order degeneracy, and (2) $V$ to be chosen such that $\delta(V)\neq 0$ under $\mathbb{H}_a$. The first condition is easy to satisfy by selecting $V$ to be a measurable function of $Z$. In contrast, the second condition is non-trivial: randomly selecting $V$ may lead to the same limitation as the classical CM test if it happens to be orthogonal to the direction of departure under $\mathbb{H}_a$. This trade-off mirrors the strengths and weaknesses of existing methods: while the CM test is first-order non-degenerate, it may lack consistency under $\mathbb{H}_a$; conversely, the ICM test is consistent under $\mathbb{H}_a$, but suffers from first-order degeneracy, resulting in a non-pivotal limiting distribution under $\mathbb{H}_o$.
To this end, the proposed testing procedure integrates the respective strengths of two existing approaches in the literature: the pivotality of the CM test and the omnibus property of the ICM test.
In particular, we focus on the quantity \begin{equation}
V_h = \big[ h(Z), \ U - h(Z) \big]^\top,
\end{equation}such that
\begin{equation}
\delta_h:=\delta(V_h)= \begin{bmatrix}
\mathbb{E}[K(Z,Z^\dagger) (h(Z)-\mathbb{E}[h(Z)]) ({U}^\dagger-\mathbb{E} U)]\\
\mathrm{ICM}(U \mid Z) - \mathbb{E}[K(Z,Z^\dagger) (h(Z)-\mathbb{E}[h(Z)]) ({U}^\dagger-\mathbb{E} U)]
\end{bmatrix}.
\end{equation}
The two components in $V_h$, namely $h(Z)$ and $U - h(Z) $, serve different purposes. First, one can view $h(Z)$ as the assistant function where prior information about the alternative can be incorporated, as done in the classical CM test. Second, the inclusion of $ U $ in the second element draws inspiration from ICM tests, which gives the omnibus property of our test, that is, it has non-trivial power in all possible directions of the alternative. This structure acts as a safeguard in scenarios where the chosen assistant $h(Z)$ is (nearly) orthogonal to the alternative, a situation in which traditional CM tests suffer from trivial power. Moreover, the $-h(Z)$ in the second element ensures that $\delta_h$ is first-order non-degenerate and that the proposed test statistic is pivotal ($\chi^2$ distributed) under the null hypothesis. We remark that the choice of
$V$ can be readily extended beyond two dimensions, although in this paper we focus on the specific two-dimensional case for clarity and analytical tractability.
The following fundamental result establishes the omnibus and first-order non-degeneracy properties of $\delta_h $.
\begin{lemma}
Let \( h(Z) \) denote an arbitrary measurable and non-degenerate function of $Z$, in the sense that \( \mathbb{P}\big( h(Z) = \mathbb{E}[h(Z)] \big) < 1 \). Then:
(a) \( \delta_h \) defined in (ref) satisfies the \emph{omnibus property}, that is, \( \|\delta_h\| = 0 \) if and only if (ref) holds; (b) \( \delta_h \) exhibits \emph{first-order non-degeneracy}.
\end{lemma}
\begin{proof}
Let $\delta_h^{(1)}=\mathbb{E}[K(Z,Z^{\dagger})(h(Z) - \mathbb{E}[h(Z)])({U}^\dagger-\mathbb{E} U)]$ and $\delta_h^{(2)}=\mathrm{ICM}(U \mid Z)-\delta_h^{(1)}$. Then from (ref) we have that $\delta_h=[\delta_h^{(1)},\delta_h^{()}]^{\top}$, and
$\delta_h^{(1)} + \delta_h^{(2)} = \mathrm{ICM}(U \mid Z) > 0$ under $\mathbb{H}_a$. This implies that under $\mathbb{H}_a$, at least one element in $\delta_h$ is strictly positive (hence strictly different from zero); $\delta_h$ satisfies the omnibus property. Note that $h(Z)$ is non-degenerate, and that under $\mathbb{H}_o$, $\mathbb{E} \big[V_h - \mathbb{E}[V_h] \mid Z\big] = \big[(h(Z) - \mathbb{E}[h(Z)]),\ \mathbb{E} [U \mid Z] - \mathbb{E} U - (h(Z) - \mathbb{E}[h(Z)]) \big]^{\top} = (h(Z) - \mathbb{E}[h(Z)])[1,-1]^{\top}\neq 0 $ with probability one. It follows that $\delta_h$ avoids first-order degeneracy.
\end{proof}
(ref) implies that, with our construction of $V_h$ in (ref), $\delta_h$ is non-zero under the alternative. Therefore, $\|\delta_h\|$ can be interpreted as a metric for conditional mean independence, i.e. $\|\delta_h\|=0$ if and only if $\mathbb{E}[U\mid Z]=\mathbb{E}[U] \ a.s. $ While the performance of $\delta_h$ depends on the choice of $h(\cdot)$, we note that similar concerns arise in ICM tests, particularly regarding the selection of kernels. A practitioner who is agnostic about the alternative can use $h(Z) = \exp\big(Z^{\top}\mathbbm{1}_{p_z}/\sqrt{p_z}\big) $ where $\mathbbm{1}_{p}$ denotes a $p\times 1$ vector of ones.\footnote{The specification is appropriate when \( p_z \) is small. For larger values of \( p_z \), it is advisable to de-mean each element of \( Z \) to ensure the variance of \( h(Z) \) is robust to \( p_z \).} The Maclaurin's series expansion, namely $\exp(z) = \sum_{l=0}^{\infty} z^l/{l!}$ shows the potential of $h(Z) = \exp\big(Z^{\top}\mathbbm{1}_{p_z}/\sqrt{p_z}\big) $ to capture different directions under the alternative.\footnote{ { The covariance matrix of $Z$ ought to be positive definite, i.e. $\mathbb{P}(Z^{\top}a=c)<1$ $\forall\, a\in\mathbb{R}^{p_z}$ and $ \forall \, c\in\mathbb{R}$, to avoid possible degeneracy in the linear combination $Z^\top\mathbbm{1}_{p_z}$.}}
A key message of (ref) is that symmetrizing any non-degenerate measurable function $h(\cdot)$ of $Z$ about $U$ guarantees the omnibus and first-order non-degeneracy properties of $\delta_h$.
The requirement on $h(Z)$ in (ref) can hardly be termed a “condition," as it imposes only measurability and non-degeneracy. The proof of (ref) provides an insight into the inclusion of $U$ linearly in the construction of $V$. This brings in the term $\mathrm{ICM}(U \mid Z)$ in $\delta_h$ that is strictly positive under $\mathbb{H}_a$ and contributes power under the alternative. The proposed test thus draws its consistency from the ICM metric.
Since $h(Z)$ can be chosen arbitrarily, the practitioner does not bear the burden of “carefully" selecting functions such as polynomials that provide power by approximating $\mathbb{E}[U \mid Z]$ under $\mathbb{H}_a$, as required in non-parametric specification tests, e.g., wooldridge1992test,yatchew1992nonparametric,zheng1996consistent. In addition, the user-specified $h(Z)$ provides extra flexibility as the practitioner can use it to augment the power of the test in given directions, unlike existing ICM-based bootstrap tests.
\begin{remark}
We assume $p_z$ is fixed in this study. In high-dimensional settings where $p_z \rightarrow \infty $, zhang2018conditional demonstrates that $\mathrm{ICM}$ criteria only capture linear dependence. Our $\chi^2$-test derives its consistency property and part of its power from the $\mathrm{ICM}$ metric so it cannot be expected to perform better in high dimensions. Recent studies have explored the challenge of testing independence between high-dimensional random vectors using distance covariance-based statistics szekely2007measuring, revealing that the performance of such tests depends critically on the relative growth rates of the dimensions of the random variables and the sample size $n$ zhu2020distance,chakraborty2021new,gao2021asymptotic,zhang2024statistical. Consequently, extending the current testing procedure to high dimensions requires a separate study that is deferred to
future research.
\end{remark}
\subsection{Extension to Testing the Nullity of $\mathbb{E}[U \mid Z]$}
In some applications, such as specification testing, one may be interested in the almost sure nullity of $\mathbb{E}[U \mid Z]$ directly, i.e.
\begin{flalign}
\begin{split}
\mathbb{H}_o^*&: \mathbb{E}[U \mid Z]=0 \ a.s.;\quad
\mathbb{H}_a^* : \mathbb{P}(\mathbb{E}[U \mid Z]=0) <1.
\end{split}
\end{flalign}
This is an augmented version of (ref) which further imposes $\mathbb{E}[U]=0$, i.e., a joint hypothesis of conditional mean independence and nullity of the unconditional mean. To this end, we follow su2017martingale by augmenting $\delta_h$ with an additional quantity that accounts for $\mathbb{E}[U] = 0$ under $\mathbb{H}_o^*$. In particular, one may consider the metric
$$
\delta_h^*=\delta_h + \mathbb{E}|K(Z,Z^{\dagger})|\mathbb{E} U \mathbb{E}\big[h(Z), \ U - h(Z)) \big]^\top.
$$
The following result shows that ${\delta}_h^*$ has the omnibus property with respect to (ref).
\begin{lemma}
${\delta}_h^*=0$ if and only if $\mathbb{E}[U \mid Z]=0 \ a.s. $
\end{lemma}
To conserve space, the discussion below focuses on $\delta_h$. Theoretical properties of the $\chi^2$-test based on $\delta_h^*$ are presented in Section (ref) of the supplementary material.
\section{Test Statistic and Theoretical Properties}
In this section, we construct the test statistic based on $\delta_h$. In Section (ref), we present a framework that unifies the tests of mean independence and model specification. We analyze the asymptotic behavior of $\widehat{\delta}_h$, the empirical estimator for $\delta_h$, and the $\chi^2$-distributed test statistic under the null hypothesis, as well as under local and fixed alternatives in Sections (ref) and (ref), respectively. Section (ref) uses Bahadur slopes to compare the power of the $\chi^2$-test and bootstrap-based $\mathrm{ICM}$ tests.
\subsection{A unified framework}
The mean independence testing problem and the specification testing problem are two closely related problems in econometric analyses. Here, we present a unified framework for both tasks. Let $X\in\mathbb{R}^{p_x}$ and $Z\in\mathbb{R}^{p_z}$ be two random vectors. We consider an econometric model defined by the following conditional moment restriction \begin{equation}
\mathbb{E}[U(X;\theta_o)|Z]=0\quad a.s.,
\end{equation}
for a unique model parameter $\theta_o\in\Theta \subset \mathbb{R}^k$. Here, $U:\mathbb{R}^{p_x}\times \Theta\mapsto \mathbb{R}^{p_u}$ is the econometric model that is assumed to be known, and $\Theta$ is the parameter space. Given $iid$ observations $\{X_i,Z_i\}_{i=1}^n$, and an empirical estimator $\widehat{\theta}_n$ for $\theta_o$, we are interested in assessing the model specification in (ref).
The above framework encompasses a wide range of applications, including treatment effect analyses (e.g., callaway2022treatment,sant2019specification), Euler and Bellman equations hansen1982generalized, escanciano2018simple, the hybrid New Keynesian Phillips curve choi2021generalized, forecast rationality hansen1980forward, and tests of conditional equal and predictive ability giacomini2006tests; see li2022learning for a comprehensive discussion.
In the mean independence testing problem outlined in Section (ref), the expressions for \( \mathrm{ICM}(U \mid Z) \) and \( \delta_h \) involve the nuisance parameter \( \mathbb{E} U \). In this case, we can simply set $X=U$, $\theta_o=\mathbb{E}U$, and $U(X;\theta)=X-\theta$, then testing for (ref) reduces to the hypothesis (ref). In the class of nonlinear models with additively separable errors $X^{(1)} = g(X^{(-1)};\beta_o) + U $, where $X^{(1)}$ denotes the first element of $X$, $X^{(-1)}$ is the sub-vector of $X$ that excludes $X^{(1)}$, and $U$ is the model error (up to a location shift), we let
\begin{align*}
U(X;\theta) &= X^{(1)} - g(X^{(-1)};\beta) - \theta_c,
\end{align*}
where $\theta = [ \theta_c,\ \beta^\top ]^\top $, and $\theta_c$ is the location parameter such that $\mathbb{E}[X_1-g(X_{-1};\beta)-U]=\theta_c$. Including the location parameter $\theta_c$ in $\theta$ becomes necessary as the ICM metric and $\delta_h$ can only identify mean independence up to an unknown nuisance mean shift. Nullity testing is discussed in Section (ref).
In practice, a natural estimator for $\delta_h$ in (ref) is given by
\begin{equation}
\widehat{\delta}_h = \frac{1}{n(n-1)}\sum_{i=1}^n\sum_{j\neq i}K(Z_i,Z_j)U(X_i;\widehat{\theta}_n)\big[(h(Z_j) - \mathbb{E}_n[h(Z)]) , \ U(X_j;\widehat{\theta}_n) - (h(Z_j) - \mathbb{E}_n[h(Z)]) \big]^\top.
\end{equation}
We remark that, in the mean independence testing (ref), we have $U(X_i;\widehat{\theta}_n)=U_i-\widehat{\theta}_n$ with $\widehat{\theta}_n=\mathbb{E}_nU.$
\subsection{Asymptotic Distribution}
In this section, we analyze the asymptotic behavior of \( \widehat{\delta}_h \) under the null hypothesis, local alternatives, and fixed alternatives. Let \( D := [U, X^\top, Z^\top]^\top \), and define \( \widetilde{U} := U(X; \theta_o) \) such that $\mathbb{E}\widetilde{U}=0.$ We consider the following sequences of local alternatives: Pitman local alternatives \(
\mathbb{H}_{an}: \mathbb{E}[\widetilde{U} \mid Z] = n^{-1/2} a(Z)\) and milder (non-Pitman) local alternatives \(
\mathbb{H}_{an}': \mathbb{E}[\widetilde{U} \mid Z] = n^{-1/4} a(Z)
\), where \( a(Z) \) is a non-degenerate measurable function of \( Z \) satisfying \( \mathbb{E}[a(Z)] = 0 \). For completeness, we impose the following regularity conditions.
\begin{assumption}
$\{D_i\}_{i=1}^n$ are independently and identically distributed ($iid$).
\end{assumption}
\begin{assumption}
$\mathbb{E}\big[K(Z,Z^{\dagger})^4\big] + \mathbb{E}\big[U^4\big] + \mathbb{E}\big[h(Z)^4\big] < \infty$.
\end{assumption}
\begin{assumption}
\( U(X; \theta) \) is differentiable in \( \theta \) and admits the following first-order expansion:
\[
U(X; \theta) = \widetilde{U} + \frac{\partial U(X; \bar{\theta})}{\partial \theta'} (\theta - \theta_o) := \widetilde{U} - G(X; \bar{\theta})^\top (\theta - \theta_o),
\]
where \( G(x; \theta) \) is continuous in \( \theta \) for all \( x \) in the support of \( X \), \( \mathbb{E}\big[\|G(X; \theta)\|^2\big] <\infty \) for all $\theta$, and \( \bar{\theta} \) lies on the line segment between \( \theta \) and \( \theta_o \), i.e., \( \| \bar{\theta} - \theta_o \| \leq \| \theta - \theta_o \| \).
\end{assumption}
\begin{assumption}
$\widehat{\theta}_n$ is a $\sqrt{n}$-consistent estimator of $\theta_o$ with the asymptotically linear representation
\[
\sqrt{n}(\widehat{\theta}_n - \theta_o) = \frac{1}{\sqrt{n}}\sum_{i=1}^n \xi_{\theta,i} + o_p(1)
\]where $ \xi_{\theta,i}\in\mathbb{R}^k$ satisfies that $\mathbb{E}[\xi_{\theta,i}]=\bf{0}$, $ \mathbb{E}\big[\|\xi_{\theta,i}\|^2\big] < \infty $.
\end{assumption}
The assumption of $iid$ observations in (ref) simplifies the theoretical analyses.\footnote{Our results are extensible to weak temporal dependence, panel data, and clustered data settings, but this lies beyond the scope of the current paper.} (ref) is standard, e.g., serfling1980; it is needed to establish the asymptotic normality of $\sqrt{n}(\widehat{\delta}_h-\delta_h)$. Although the moment restriction on $U$ cannot be relaxed, that of $K(\cdot)$ can be relaxed through the choice of bounded kernels or suitable transformations of $Z$. (ref) regulates the smoothness and moment conditions of the possibly nonlinear parameterized error $U(X;\theta)$. (ref) is akin to escanciano2009lack; it assumes the estimator $\widehat{\theta}_n$ is consistent with respect to its probability limit $\theta_o$. The Bahadur linear representation holds for a broad class of estimators, including commonly used estimators such as (nonlinear) least squares, maximum likelihood, and the (generalized) method of moments, under standard regularity conditions.
The result on the limiting behavior of the proposed test statistic relies on the following asymptotically linear representation of
$\widehat{\delta}_h$ in the following lemma. Define $\widetilde{h}(Z):=h(Z)-\mathbb{E}h(Z)$.
\begin{lemma}
Under (ref), $\widehat{\delta}_h$ has the following asymptotically linear representation:
\begin{align*}
\sqrt{n}\big(\widehat{\delta}_h - \delta_h\big)
= \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\xi_h(D_i) + o_p(1),
\end{align*}
where $ \xi_h(D_i):= 2\big[\xi_{h,1}(D_i), \ \xi_{h,2}(D_i) - \xi_{h,1}(D_i) \big]^\top $, $
\delta_h = \big[\delta_h^{(1)},\delta_h^{(2)}\big]^{\top}
$,
\begin{align*}
\xi_{h,1}(D_i): &= \psi_h^{(1)}(D_i) - \delta_h^{(1)} - H_1^\top[\xi_{\theta,i}^\top, \widetilde{h}(Z_i)]^\top , \\
\xi_{h,2}(D_i): &= \psi_U^{(1)}(D_i)-\mathrm{ICM}(U\mid Z) - H_2^\top\xi_{\theta,i},\\
\psi_h^{(1)}(D_i) : &=\mathbb{E}[\psi_h(D_i,D_j)|D_i],\\
\psi_U^{(1)}(D_i): &=\mathbb{E}[\psi_U(D_i,D_j)|D_i],
\end{align*}
$\delta_h^{(1)}=\mathbb{E}[K(Z,Z^{\dagger})\widetilde{h}(Z)\widetilde{U}^\dagger], $ $\delta_h^{(2)}=\mathrm{ICM}(U \mid Z)-\delta_h^{(1)}$, $ \psi_h(D_i,D_j):= K(Z_i,Z_j)\big\{ \widetilde{h}(Z_j)\widetilde{U}_i + \widetilde{h}(Z_i)\widetilde{U}_j \big\}/2 $, $\psi_U(D_i,D_j) := K(Z_i,Z_j)\widetilde{U}_i\widetilde U_j$,
$ H_1= (1/2)\Big[ \mathbb{E}\big[K(Z, Z^\dagger)\widetilde{h}(Z^\dagger)G(X;\theta_o)\big]^\top
,\ \mathbb{E}\big[K(Z,Z^\dagger)\widetilde{U}\big]
\Big]^\top $, and
$H_2= \mathbb{E}\big[K(Z,Z^\dagger){\widetilde{U}^\dagger}G(X;\theta_o)\big] $.
\end{lemma}
The proof of (ref) exploits the classical Hoeffding decomposition serfling1980 for U-statistics, and provides intuition for the two different rates of local alternatives. If the user-defined assistant $h(Z)$ is non-orthogonal to the alternative direction $\mathbb{E}\big[a(Z^{\dagger})K(Z,Z^{\dagger}) \mid Z\big]$ under the Pitman local alternatives $ \mathbb{H}_{an}: \ \mathbb{E}[\widetilde{U}|Z]=n^{-1/2}a(Z)$, then $\sqrt{n}\delta_h^{(1)}$ is non-diminishing, i.e.
$$
\sqrt{n}\delta_h^{(1)} = \mathbb{E}\big[K(Z,Z^{\dagger})\widetilde{h}(Z)\sqrt{n}\mathbb{E}(\widetilde{U}^\dagger\mid Z^{\dagger})\big] = \mathbb{E}[K(Z,Z^{\dagger})\widetilde{h}(Z)a(Z^{\dagger})] \not \to 0,
$$
which drives the power in a manner similar to the classical CM test. If instead $h(Z)$ is (nearly) orthogonal to the alternative with $\sqrt{n}\delta_h^{(1)}\to 0$, then for the milder local alternatives $ \mathbb{H}_{an}': \ \mathbb{E}[\widetilde{U}|Z] = n^{-1/4}a(Z)$, $$\sqrt{n}\delta_h^{(2)}\to \sqrt{n}\mathrm{ICM}(U\mid Z)= \sqrt{n}\mathbb{E}[K(Z,Z^{\dagger})\mathbb{E}(\widetilde{U}|Z)\mathbb{E}(\widetilde{U}^{\dagger}|Z^{\dagger})]=\mathbb{E}[K(Z,Z^{\dagger})a(Z)a(Z^{\dagger})]>0.$$
In this case, the classical ICM test drives the power. As such, it provides robustness to the user’s choice of direction, though this comes at the cost of a slower detection rate under local alternatives.
To study the limiting behavior of $\sqrt{n}(\widehat{\delta}_h-\delta_h)$, we
define $$\Omega_h:= \operatorname{Var}[\xi_h(D_i)].$$ We distinguish $\Omega_{h,o}$ and $\Omega_{h,a}$, corresponding to specific expressions of $\Omega_h$ under $\mathbb{H}_o$ and $\mathbb{H}_a$, respectively. Both are generically referred to as $\Omega_h$ whenever the distinction is not necessary. In particular, under $\mathbb{H}_o$, $\psi_U^{(1)}(D_i)=\mathrm{ICM}(U\mid Z)=H_2^\top\xi_{\theta,i}=0 \ a.s.$ by the LIE, and $\sqrt{n}\big(\widehat{\delta}_h - \delta_h\big)$ reduces to $ \sqrt{n}\widehat{\delta}_h = [1, -1]^\top \frac{2}{\sqrt{n}}\sum_{i=1}^{n} \xi_{h,1}(D_i) + o_p(1) $, with
\begin{equation}
\Omega_{h,o}: = 4\mathbb{E}\big[\xi_{h,1}(D)^2\big]\times \begin{bmatrix}
1 & -1\\ -1 & 1
\end{bmatrix}.
\end{equation}
The following result builds upon (ref).
\begin{theorem}
Suppose (ref) hold, then
\begin{enumerate}[(i)]
• under $\mathbb{H}_o$,
$
\sqrt{n}\widehat{\delta}_h\overset{d}{\rightarrow}\mathcal{N}(0,\Omega_{h,o}) $;
• under $\mathbb{H}_{an}$,
$
\sqrt{n}\widehat{\delta}_h\overset{d}{\rightarrow}\mathcal{N}(a_o,\Omega_{h,o})
$;
• under $\mathbb{H}_{an}'$,
$
\sqrt{n}\widehat{\delta}_h\overset{d}{\rightarrow}\mathcal{N}(a_o',\Omega_{h,o})
$ if $\mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})]=0$; and
• under $\mathbb{H}_a$,
$
\sqrt{n}(\widehat{\delta}_h-\delta_h) \overset{d}{\rightarrow}\mathcal{N}(0,\Omega_{h,a}); $
\end{enumerate}
where $a_o := [1,-1]^{\top}\mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})]$ and $a_o':= \big[ 0, \, \mathrm{ICM}\big(a(Z)\mid Z\big) \big]^\top $.
\end{theorem}
(ref) offers a comprehensive asymptotic analysis of $\widehat{\delta}_h$ under null, local alternatives, and fixed alternatives. In particular, part (i) shows that $\sqrt{n}\widehat{\delta}_h$ is non-degenerate and converges in distribution, which forms the basis of our $\chi^2$-distributed test. Under the Pitman local alternatives $\mathbb{H}_{an}$, the test is only non-trivial in directions where the user-defined assistant $h(Z)$ is non-orthogonal to $\mathbb{E}\big[a(Z^{\dagger})K(Z,Z^{\dagger}) \mid Z\big]$, i.e., $\mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})] \neq 0$. This is expected, as the use of the assistant $h(Z)$ draws inspiration from classical conditional moment (CM) tests. Part (iii) of (ref) partially resolves this limitation: the local power becomes non-trivial in all directions due to the presence of $\mathrm{ICM}\big(a(Z)\mid Z\big)$ in the term $a_o' := [0, \ \mathrm{ICM}\big(a(Z)\mid Z\big)]^\top$ under a sequence of alternatives converging to the null at the $n^{-1/4}$ rate. In other words, the test becomes omnibus for alternative directions that converge to the null at a rate slower than $n^{-1/4}$, thereby enhancing its ability to detect a broader class of local alternatives. The two types of local alternatives stem from our novel construction of the test, which combines components that operate at two different convergence rates. Part (iv) suggests the resulting test is consistent under the fixed alternative.
\begin{remark}
The power performance of the proposed $\chi^2$-test under the Pitman local alternative $\mathbb{H}_{an}$ is sensitive to the choice of $h(Z)$ and the ICM kernel $K(Z,Z^\dagger)$. A potential way to enhance the local power of the proposed test is to extend the estimand $\delta_h$ to a vector
$\big\{ \delta_{h,K} \big\}_{\{h,K\} \in \mathcal{H} \times \mathcal{K}}$, where the subscript in $\delta_{h,K}$ emphasizes the dependence of the parameter on both the function $h$ and the ICM kernel $K$. Here, $\mathcal{H}$ and $\mathcal{K}$ are finite sets of measurable, non-degenerate functions of $Z$ and valid ICM kernels, respectively.\footnote{Another possibility is to extend $V_h$ with a vector comprising a dictionary of transformations, e.g., polynomials of $Z$ -- see (ref) for some simulation evidence. A treatment of these extensions is omitted for considerations of space and scope.}
\end{remark}
\subsection{Test Statistic}
Theorem (ref) justifies the asymptotic normality of $\widehat{\delta}_h$, even under $\mathbb{H}_o$, which naturally motivates the Wald test statistic:
\begin{equation}
\widetilde{T}_{V,n}=n\widehat{\delta}_h^{\top}\widetilde{\Omega}_{h,n}^{-1}\widehat{\delta}_h,
\end{equation}
where $\widetilde{\Omega}_{h,n}$ is a consistent estimator of $\Omega_h$ under both $\mathbb{H}_o$ and $\mathbb{H}_a$. This testing procedure is valid when $\Omega_h$ is positive definite. However, from (ref), $\Omega_h$ is singular under $\mathbb{H}_o$ with $\mathrm{rank}(\Omega_{h,o})=1$.
Although it is tempting to replace the inverse matrix $\widetilde{\Omega}_{h,n}^{-1}$ in (ref) with a generalized inverse matrix $\widetilde{\Omega}_{h,n}^{-}$, e.g., the Moore-Penrose inverse, the resulting Wald statistic may still not have an asymptotic $\chi^2$ distribution unless the rank condition, $\mathbb{P}\big(\mathrm{rank}(\widetilde{\Omega}_{h,n})=\mathrm{rank}(\Omega_h)\big) \to 1 \quad \text{as }n\to\infty$ is satisfied andrews1987asymptotic. To ensure the rank condition is satisfied and thereby address this problem, we adopt the thresholding technique described in lutkepohl1997modified (see also duchesne2015multivariate,dufour2016rank). Let \begin{equation*}
\widetilde{\Omega}_{h,n}:=\frac{1}{n-1}\sum_{i=1}^n \widehat{\xi}_h(D_i)\widehat{\xi}_h(D_i)^{\top},
\end{equation*}
be a consistent estimator of $\Omega_h$ sen1960some, where $\widehat{\xi}_h(D_i)$ denotes the sample analog of $\xi_h(D_i)$ with parameters replaced by estimates.
By the singular value decomposition,
\begin{equation} \widetilde{\Omega}_{h,n}=\widetilde{\Gamma}_n\widetilde{\Lambda}_n\widetilde{\Gamma}_n^{\top},
\end{equation}
where $\widetilde{\Lambda}_n = \mathrm{diag}(\widetilde{\lambda}_1,\widetilde{\lambda}_2)$ is the diagonal matrix comprising the eigenvalues $\widetilde{\lambda}_1\geq \widetilde{\lambda}_2 \geq 0$ of $\widetilde{\Omega}_{h,n}$, and the columns of $\widetilde{\Gamma}_n$ are the corresponding eigenvectors. Let \( c_n = C n^{-1/2 + \iota} \), where \( \iota \in (0, 1/2) \) is a small positive constant and \( C \in (0, \infty) \) is a fixed constant, as in dufour2016rank. Since \( \mathrm{rank}(\Omega_h) \geq 1 \) regardless of whether the null hypothesis \( \mathbb{H}_o \) holds, we define the regularized estimator of \( \Omega_h \) as
\begin{equation}
\widehat{\Omega}_{h,n} := \widetilde{\Gamma}_n \widehat{\Lambda}_{n,c_n} \widetilde{\Gamma}_n^{\top},
\quad \text{where} \quad
\widehat{\Lambda}_{n,c_n} = \mathrm{diag}\big( \widetilde{\lambda}_1, \widetilde{\lambda}_2 \mathbf{1}(\widetilde{\lambda}_2 > c_n) \big).
\end{equation}
The corresponding Moore–Penrose inverse is defined as
\[
\widehat{\Omega}_{h,n}^{-} := \widetilde{\Gamma}_n \widehat{\Lambda}_{n,c_n}^{-} \widetilde{\Gamma}_n^{\top},
\quad \text{where} \quad
\widehat{\Lambda}_{n,c_n}^{-} = \mathrm{diag}\big( \widetilde{\lambda}_1^{-1}, \widetilde{\lambda}_2^{-1} \mathbf{1}(\widetilde{\lambda}_2 > c_n) \big).
\]
Finally, we define the regularized Wald test statistic:
\begin{equation}
T_{h,n} := n\widehat{\delta}_h^{\top}\widehat{\Omega}_{h,n}^{-}\widehat{\delta}_h.
\end{equation}
In practice, $\mathbb{H}_o$ is rejected when $T_{h,n} > \chi^2_{1,1-\alpha}$, where $\chi^2_{1,1-\alpha}$ denotes the $1-\alpha$ quantile of the $\chi^2_1$ distribution at a pre-specified significance level $\alpha\in(0,1)$. We remark that the ranks of $\Omega_{h,o}$ and $\Omega_{h,a}$ can differ: $\mathrm{rank}(\Omega_{h,o}) = 1$, while $\mathrm{rank}(\Omega_{h,a}) = 2$ if $\mathbb{P}(a(Z)= c \cdot \widetilde{h}(Z))<1$ for any $c \in \mathbb{R}$, and $\mathrm{rank}(\Omega_{h,a}) = 1$ otherwise. We use $c_n=\widetilde{\lambda}_1 n^{-1/3}$ following lutkepohl1997modified and dufour2016rank.\footnote{For the robustness of the $\chi^2$-test to variations of \( c_n = \widetilde{\lambda}_1 n^{-\iota} \), \( \iota \in (0, 1/2) \), and other common truncation criteria, see Section (ref) in the supplement.}
\begin{remark}
$\mathrm{rank}(\Omega_{h,o}) = 1$ and $\mathrm{ICM}_n(U \mid Z) = O_p(n^{-1}) $ under $\mathbb{H}_o$. From the proof of part (i) of (ref), $ \sqrt{n}\widehat{\delta}_h = \sqrt{n} \widehat{\delta}_h^{(1)} [1,\ -1]^{\top} + O_p(n^{-1/2}) $, and
\[ T_{h,n} = n\widehat{\delta}_h^{\top}\widehat{\Omega}_{h,n}^{-}\widehat{\delta}_h = \Bigg(\frac{\widehat{\delta}_h^{(1)}}{\sqrt{\operatorname{Var}(\widehat{\delta}_h^{(1)})}}\Bigg)^2 + o_p(1) \xrightarrow{d} \chi_1^2 \quad \text{as}\quad n\rightarrow \infty
\]under $\mathbb{H}_o$. This implies $\sqrt{T_{h,n}}$ converges in distribution to the half standard normal under $\mathbb{H}_o$, and shares the interpretability (without a formal hypothesis test) of a two-sided $t$-test.
\end{remark}
The next theorem justifies using (ref) under the null, local alternative, and fixed alternative hypotheses.
\begin{theorem}
If $\,\mathbb{E}\left|\xi_h(D)\right|^{4+\varepsilon}<\infty$ for some $\varepsilon>0$, and let $c_n= C n^{-1/2+\iota} $ for some constants $\iota\in(0,1/2)$ and $C\in(0,\infty)$ independent of $n$, then
\begin{enumerate}[(i)]
• under $\mathbb{H}_o$, $
\widehat{\Omega}_{h,n}^-\overset{p}{\rightarrow} \Omega_{h,o}^-,
$ and
$$
T_{h,n}\overset{d}{\rightarrow} \chi^2_1;
$$
• under $\mathbb{H}_{an}$, $
\widehat{\Omega}_{h,n}^-\overset{p}{\rightarrow} \Omega_{h,o}^-
$, and the asymptotic local power is given by
$$
\lim_{n\to\infty}\mathbb{P}(T_{h,n}>\chi^2_{1,1-\alpha})=\mathbb{P}\left(\chi^2_{1}(b_o)>\chi^2_{1,1-\alpha}\right),
$$
where $ \displaystyle b_o:=a_o^{\top}\Omega_{h,o}^-a_o= \frac{(\mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})])^2}{4\mathbb{E}[\xi_{h,1}(D)^2]} $, $a_o$ is defined in (ref), and $\chi^2_1(b_o)$ is a non-central $\chi_1^2$ random variable;
• under $\mathbb{H}_{an}'$, $
\widehat{\Omega}_{h,n}^-\overset{p}{\rightarrow} \Omega_{h,o}^-
$, and the asymptotic local power is given by
$$
\lim_{n\to\infty}\mathbb{P}(T_{h,n}>\chi^2_{1,1-\alpha})=\mathbb{P}\left(\chi^2_{1}(b_o')>\chi^2_{1,1-\alpha}\right),
$$
if $\mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})]=0$ where $ \displaystyle b_o':=a_o^{'\top}\Omega_{h,o}^-a_o' = \frac{\mathrm{ICM}^2\big( a(Z) \mid Z \big)}{16\mathbb{E}[\xi_{h,1}(D)^2]} >0$, and $a_o'$ is defined in (ref); and
• under $\mathbb{H}_a$, $
\widehat{\Omega}_{h,n}^-\overset{p}{\rightarrow} \Omega_{h,a}^-,
$
and if $\delta_h\not\in\mathcal{M}_0$, where $\mathcal{M}_0$ is the eigenspace associated with the null eigenvalue of $\Omega_{h,a}$, $$
\lim_{n\to\infty}\mathbb{P}(T_{h,n}>\chi^2_{1,1-\alpha})=1.
$$
\end{enumerate}
\end{theorem}
(ref) establishes the consistency of the thresholded estimator of the Moore–Penrose inverse of the covariance matrix, and characterizes the asymptotic behavior of the resulting test statistic, in parallel with the analysis in (ref). In particular, under the local alternatives $\mathbb{H}_{an}$ of part (ii), when $a_0 = \mathbb{E}[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})] = 0$, the test exhibits trivial power because $b_o = 0$. However, this limitation is partially remedied in part (iii), where the modified limiting direction $b_o'$ remains positive.
\begin{remark}
Under $\mathbb{H}_a$, $\delta_h\not\in\mathcal{M}_0$ always holds for the mean independence test. The proof is provided in (ref) of the supplementary material.
\end{remark}
\subsection{Power Comparison with the bootstrap-based ICM test}
In this section, we examine the efficiency of the proposed test statistic compared with the bootstrap-based ICM test by adopting the approach of bahadur1960stochastic. Under a fixed alternative, both tests demonstrate consistency as $n$ tends to infinity. The Bahadur slope in bahadur1960stochastic allows us to further compare the rate of convergence of \( p \)-values to zero as $n$ increases.
Let \( \displaystyle S_G(t):=\mathbb{P}\Big( \sum_{k=1}^{\infty}\lambda_k G_k^2>t \Big) \) and \( \displaystyle S_T(t):=\mathbb{P}\big( \chi^2_{1} > t \big) \) be the survival functions of the asymptotic null distributions of the ICM test statistic $n\mathrm{ICM}_n(U \mid Z)$, and the pivotal test statistic, $T_{h,n}$, respectively.
Their Bahadur slopes are respectively given by
$$
c_G= \lim_{n\to\infty}-\frac{2}{n} \log S_G(n\mathrm{ICM}_n(U \mid Z)) \ \text{ and } \ c_T= \lim_{n\to\infty}-\frac{2}{n} \log S_G(T_{h,n}).
$$
\begin{theorem}
Suppose $
\widehat{\Omega}_{h,n}^-\overset{p}{\rightarrow} \Omega_{h,a}^-,
$ and the conditions of (ref) hold, then under $\mathbb{H}_a$, the (approximate) Bahadur slopes of the ICM test statistic $n\mathrm{ICM}_n(U \mid Z)$ and the pivotal test statistic $T_{h,n}$ are, respectively, given by
$$
c_G= \frac{\mathrm{ICM}(U \mid Z)}{\lambda_1} \quad \text{ and } c_T= \delta_a^{\top}\Omega_{h,a}^{-}\delta_a
$$
where $\lambda_1$ is the leading eigenvalue associated with the limiting null distribution in (ref), and
$$
\delta_a = \Big[\mathbb{E}\big[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})\big], \ \mathrm{ICM}(U \mid Z) - \mathbb{E}\big[\widetilde{h}(Z)a(Z^{\dagger})K(Z,Z^{\dagger})\big] \Big]^\top.
$$
\end{theorem}
The approximate Bahadur slopes presented in Theorem (ref) are primarily of theoretical interest. Conducting a comprehensive comparison of these slopes is challenging as they depend on data-dependent quantities such as $\lambda_1$, $a(Z)$, and user-specified variables such as $h(Z)$ and the kernel $K(\cdot,\cdot)$.
\begin{remark}
The main implication of (ref) is that, even with a fixed ICM kernel, neither test is uniformly more powerful across all data-generating processes satisfying (ref). This limitation arises because the true alternative is unknown and the function \( h(\cdot) \) is user-specified. (ref) presents a numerical illustration of this result.
\end{remark}
\section{Monte Carlo Experiments - Specification Test}
This section examines the size and power of the proposed $\chi^2$ specification test compared to bootstrap-based ICM procedures via simulations. The specification $\widehat{V}_h = [h(Z),\ \widehat{U}-h(Z)]^\top$ with $h(Z) = \exp\big(Z^\top\mathbbm{1}_{p_z}/\sqrt{p_z} \big) $. For the proposed $\chi^2$-test, the regularized inverse in (ref) is computed using $c_n=\widetilde{\lambda}_1 n^{-1/3}$ where $\widetilde{\lambda}_1$ is the leading eigenvalue of $\widetilde{\Omega}_{h,n}$. We consider two commonly used ICM kernels: the Gaussian $K(Z,Z^\dagger) = \exp\left(-0.5 \|Z - Z^\dagger\|^2\right) $ , e.g., bierens1982consistent and the negative Euclidean $K(Z,Z^\dagger) = -\|Z - Z^\dagger\|$ shao2014martingale. For a given ICM kernel, the proposed $\chi^2$-test is compared to ICM tests based on the multiplier bootstrap (MB) of escanciano-2024-gaussian and the standard wild bootstrap, e.g., escanciano2006consistent,su2017martingale. The empirical size and power curves are based on $1000$ Monte Carlo replicates. $999$ bootstrap samples are used for the bootstrap procedures.\footnote{Both bootstrap procedures are conducted using the mammen1993bootstrap two-point distributed auxiliary variables.} Other simulation results on specification and mean independence tests are available in Section (ref) of the supplementary material.
\subsection{Specifications}
We consider the linear model
$$
Y= \theta_c + \sum_{l=1}^5X_l\theta_l + U,$$
with and without excluded instruments.\footnote{{(ref) provides simulation results on the specification of propensity scores using the logit model.}} The data-generating processes (DGPs) are varied through $U$:
\begin{enumerate}[label=(\roman*)]
• $\displaystyle U = \frac{\mathcal{E}}{\sqrt{1+Z_1^2}} $;
• $ \displaystyle U = \frac{ \gamma }{5\sqrt{2}} \sum_{l=1}^5 Z_l^2 + \frac{\mathcal{E}}{\sqrt{1+Z_1^2}} $;
• $ \displaystyle U = \gamma \sum_{l=1}^5 \frac{\cos(2Z_l)}{\sqrt{2(1-\exp(-8))}} + \frac{\mathcal{E}}{\sqrt{1+Z_1^2}} $; and
• $ \displaystyle U = \gamma\sum_{l=1}^5 \big(\exp(-Z_l^2/3) - \sqrt{3/5}\big) + \frac{\mathcal{E}}{\sqrt{1+Z_1^2}} $
\end{enumerate}
where $ \mathcal{E} \sim \mathcal{N}(0,1)$ independently of $Z$ which follows the multivariate normal distribution with mean zero and covariance matrix $\Sigma_{ll'} = 0.25^{|l-l'|}$, $ [\theta_c,\theta_1,\ldots,\theta_5]^\top = \mathbbm{1}_6 $, and $\gamma\in[0,1]$ tunes the deviation away from $\mathbb{H}_o$. $X=Z$ in DGP LS1 and $X_1 = (1.5Z_1 + \widetilde{\mathcal{E}}) / \sqrt{3.25} $ with $\widetilde{\mathcal{E}}\sim \mathcal{N}(0,1)$ and $\mathrm{Cov}(\widetilde{\mathcal{E}},{\mathcal{E}})=0.25$ in DGPs LS2 through LS4.\footnote{The non-parametric quantity $\mathbb{E}[X_1 \mid Z_1]$ is needed for the multiplier bootstrap under endogeneity -- see escanciano-2024-gaussian. It is estimated using a third-order polynomial.} Thus, $X$ is exogenous under DGP LS1 and endogenous under DGPs LS2 through LS4. LS1 is estimated via Ordinary Least Squares (OLS) while the remaining DGPs are estimated using the Instrumental Variable (IV) estimator with $Z$ as the instrument. Heteroskedasticity of arbitrary form is imposed in all DGPs.
Different sample sizes $n \in \{200, 400, 600, 800\}$ serve to examine the empirical size and local power of the test using DGPs LS1 and LS2 (with $\gamma=0$). For the analyses of power in DGPs LS2 through LS4, the sample size is kept at $n=400$ while $\gamma$ is varied in order to study the power of the proposed $\chi^2$-test in comparison to the bootstrap-based ICM procedures.
\subsection{Empirical Size and Power}
(ref) presents the empirical sizes corresponding to DGPs LS1 (under strictly exogenous $X$) and LS2 (under endogenous $X$ instrumented with $Z$) in addition to the local power under LS2 at three nominal levels: 10%, 5%, and 1%. One observes comparably good size control of the proposed $\chi^2$-test and the bootstrap-based procedures across all the sample sizes considered. This provides evidence in support of the proposed test's validity under both exogenous $X$ and endogenous $X$ with valid instruments $Z$.
\begin{table}[!htbp]
\caption{Empirical Size & Local Power}
\begin{tabular}{lcccccccccc}
\toprule
& & \multicolumn{3}{c}{$ 10\% $} & \multicolumn{3}{c}{$ 5\% $} & \multicolumn{3}{c}{$ 1\% $} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}\cmidrule(lr){9-11}
$n$ & Kernel & $\chi^2$ & MB & WB & $\chi^2$ & MB & WB & $\chi^2$ & MB & WB \\ \midrule
\multicolumn{1}{c}{\textcolor{blue}{LS1}} & & \multicolumn{9}{c}{Empirical Size} \\
\cmidrule(lr){1-1} \cmidrule(lr){3-11}
200 &Gauss &0.096 &0.103 &0.103 &0.053 &0.049 &0.052 &0.007 &0.012 &0.012 \\
&Euclid &0.101 &0.072 &0.057 &0.048 &0.031 &0.021 &0.008 &0.003 &0.003 \\ \cmidrule[0.25pt](lr){2-11}
400 &Gauss &0.103 &0.098 &0.096 &0.046 &0.048 &0.046 &0.007 &0.009 &0.006 \\
&Euclid &0.104 &0.074 &0.056 &0.053 &0.032 &0.029 &0.005 &0.003 &0.004 \\ \cmidrule[0.25pt](lr){2-11}
600 &Gauss &0.089 &0.102 &0.106 &0.053 &0.050 &0.047 &0.009 &0.011 &0.008 \\
&Euclid &0.100 &0.084 &0.085 &0.047 &0.036 &0.034 &0.006 &0.006 &0.005 \\ \cmidrule[0.25pt](lr){2-11}
800 &Gauss &0.090 &0.104 &0.097 &0.039 &0.053 &0.055 &0.008 &0.007 &0.006 \\
&Euclid &0.098 &0.088 &0.078 &0.042 &0.040 &0.041 &0.006 &0.004 &0.006 \\
\midrule
\multicolumn{1}{c}{\textcolor{blue}{LS2}} & & \multicolumn{9}{c}{Empirical Size} \\
\cmidrule(lr){1-1} \cmidrule(lr){3-11}
200 &Gauss &0.095 &0.093 &0.100 &0.055 &0.052 &0.053 &0.007 &0.011 &0.014 \\
&Euclid &0.097 &0.070 &0.055 &0.045 &0.032 &0.021 &0.008 &0.003 &0.002 \\ \cmidrule[0.25pt](lr){2-11}
400 &Gauss &0.102 &0.103 &0.099 &0.046 &0.051 &0.048 &0.007 &0.007 &0.006 \\
&Euclid &0.103 &0.079 &0.056 &0.056 &0.031 &0.028 &0.005 &0.005 &0.003 \\ \cmidrule[0.25pt](lr){2-11}
600 &Gauss &0.089 &0.094 &0.107 &0.052 &0.044 &0.048 &0.008 &0.011 &0.008 \\
&Euclid &0.100 &0.090 &0.083 &0.046 &0.040 &0.032 &0.006 &0.010 &0.006 \\ \cmidrule[0.25pt](lr){2-11}
800 &Gauss &0.090 &0.099 &0.095 &0.039 &0.051 &0.055 &0.009 &0.008 &0.006 \\
&Euclid &0.099 &0.089 &0.079 &0.042 &0.041 &0.037 &0.006 &0.004 &0.006 \\
\midrule
\multicolumn{1}{c}{\textcolor{blue}{LS2}} & & \multicolumn{9}{c}{Local Power: $\gamma = 5/\sqrt{n}$} \\
\cmidrule(lr){1-1} \cmidrule(lr){3-11}
200 &Gauss &0.549 &0.486 &0.480 &0.418 &0.391 &0.371 &0.164 &0.223 &0.201 \\
&Euclid &0.883 &0.656 &0.625 &0.803 &0.543 &0.507 &0.535 &0.295 &0.262 \\ \cmidrule[0.25pt](lr){2-11}
400 &Gauss &0.605 &0.545 &0.522 &0.459 &0.413 &0.390 &0.202 &0.226 &0.182 \\
&Euclid &0.887 &0.707 &0.689 &0.813 &0.601 &0.575 &0.588 &0.376 &0.341 \\ \cmidrule[0.25pt](lr){2-11}
600 &Gauss &0.592 &0.521 &0.509 &0.461 &0.423 &0.408 &0.221 &0.237 &0.215 \\
&Euclid &0.896 &0.723 &0.714 &0.830 &0.614 &0.621 &0.606 &0.410 &0.398 \\ \cmidrule[0.25pt](lr){2-11}
800 &Gauss &0.606 &0.547 &0.539 &0.473 &0.441 &0.428 &0.243 &0.246 &0.225 \\
&Euclid &0.912 &0.743 &0.742 &0.840 &0.651 &0.644 &0.632 &0.420 &0.422 \\
\bottomrule
\end{tabular}
\end{table}
To ensure that the good size control of the $\chi^2$-test in (ref) is not achieved at the expense of power under the alternative, we consider its performance under both local and fixed alternatives. The results, presented in the third panel of (ref), indicate that the local power of the proposed $\chi^2$-test, along with the bootstrap-based procedures, is non-trivial. (ref) present power curves corresponding to DGPs LS2 through LS4 using the Gaussian and Negative Euclidean kernels. One observes that all tests demonstrate non-trivial power against the fixed alternative for variations of $\gamma$ away from zero. The proposed test exhibits highly competitive power performance across all three DGPs and both kernels considered. Notably, the relative power of the $\chi^2$-test compared to the bootstrap-based procedures demonstrates variability depending on the specific DGP and kernel, as highlighted in (ref).
\begin{figure}[H]
\caption{DGP LS2 -- Gaussian Kernel -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
\begin{figure}[H]
\caption{DGP LS2 -- Negative Euclidean -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
\begin{figure}[H]
\caption{DGP LS3 -- Gaussian Kernel -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
\begin{figure}[H]
\caption{DGP LS3 -- Negative Euclidean -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
\begin{figure}[H]
\caption{DGP LS4 -- Gaussian Kernel -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
\begin{figure}[H]
\caption{DGP LS4 -- Negative Euclidean -- $n=400$.}
\begin{subfigure}{0.32\textwidth}
\caption{10%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{5%}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{1%}
\end{subfigure}
\end{figure}
The $\chi^2$-test proves useful when researchers seek a measure of correct model specification or mean independence without resorting to a formal hypothesis test, owing to the pivotality of its test statistic. In summary, our simulations demonstrate that the proposed $\chi^2$-test, in addition to its enhanced interpretability compared to bootstrap-based ICM specification tests, exhibits good size control and comparable power performance.
\subsection{Running Time}
One key advantage of a pivotal test is its computational efficiency relative to bootstrap-based procedures, which becomes particularly critical in large samples. Accordingly, a comparison of the running times of competing tests is in order. (ref) presents a comparison between the proposed pivotal $\chi^2$-test and the bootstrap-based ICM tests in terms of computational time for fixed ICM kernels.\footnote{Computations were performed on a 2.8 GHz Quad-Core Intel Core i7 with 16 GB of RAM, running on a MacBook Pro.}
\begin{table}[!htbp]
{4pt}
\caption{Running Time - Specification Test - DGP LS1}
\begin{tabular}{lcccccc}
\toprule \toprule
& & \multicolumn{3}{c}{Average Running Time (seconds)} & \multicolumn{2}{c}{Median Relative Time} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-7}
n & Kernel & $\chi^2$ & MB & WB & MB & WB \\ \midrule
& Gauss & 0.004 & 0.114 & 0.965 & 29.000 & 248.000 \\
200 & & (0.002) & (0.021) & (0.063) & (16.833) & (142.208) \\
& Euclid & 0.004 & 0.111 & 0.993 & 35.000 & 315.667 \\
& & (0.001) & (0.006) & (0.043) & (15.425) & (127.733) \\ \midrule
& Gauss & 0.011 & 0.440 & 1.100 & 47.222 & 116.422 \\
400 & & (0.003) & (0.022) & (0.054) & (18.856) & (41.898) \\
& Euclid & 0.009 & 0.435 & 1.095 & 58.786 & 145.571 \\
& & (0.002) & (0.016) & (0.055) & (23.858) & (55.071) \\ \midrule
& Gauss & 0.021 & 1.137 & 1.254 & 59.833 & 65.526 \\
600 & & (0.004) & (0.035) & (0.048) & (18.407) & (18.479) \\
& Euclid & 0.016 & 1.125 & 1.247 & 76.857 & 85.000 \\
& & (0.003) & (0.040) & (0.051) & (25.869) & (27.041) \\ \midrule
& Gauss & 0.038 & 2.427 & 1.552 & 65.859 & 41.789 \\
800 & & (0.005) & (0.057) & (0.048) & (13.833) & (7.420) \\
& Euclid & 0.030 & 2.471 & 1.584 & 87.554 & 55.930 \\
& & (0.004) & (0.053) & (0.040) & (22.298) & (13.505) \\ \bottomrule
\end{tabular}
\textit{Notes: The second row (in parentheses) for each sample size and kernel includes standard deviations and inter-quartile ranges for the average running times and relative times the $\chi^2$-test as the benchmark), respectively.}
\end{table}
(ref) illustrates the substantial computational advantage of employing the $\chi^2$ specification test. The average running times for the $\chi^2$-test are negligible compared to those of the bootstrap-based procedures across all considered sample sizes. A striking difference emerges in terms of median relative computational time (with the $\chi^2$-test serving as a benchmark). While the multiplier bootstrap procedure is generally faster than the wild bootstrap procedure, our findings indicate that the median relative computational time of the multiplier bootstrap tends to increase with the sample size, whereas that of the wild bootstrap appears to decrease. At particularly large sample sizes ($n=600$ and $n=800$), the computational gain offered by the $\chi^2$-test becomes considerable.
\section{Conclusion}
Despite the 40-year history of ICM tests, dating back to bierens1982consistent and encompassing numerous interesting contributions, a \emph{bona fide} pivotalized ICM specification test remains lacking. This paper achieves the objective of proposing an omnibus $\chi^2$-test of specification and mean independence based on ICM metrics.
The proposed $\chi^2$-test complements existing ICM tests by overcoming certain limitations of the latter. The test statistic can be constructed with functional forms that boost power in the direction of alternatives the researcher may have in mind. The test is computationally more efficient than commonly used bootstrap-based ICM tests and remains computationally viable even in large samples. In addition to providing a reliable pivotal test that derives its omnibus property from ICM metrics, the test statistic offers an easily interpretable metric of model specification and mean dependence, thereby obviating formal hypothesis tests.
In conclusion, we highlight several potential extensions to the current work. While this paper offers a viable solution for pivotalizing ICM tests, it is noteworthy that this approach necessitates regularization under the null hypothesis to address the ill-posed inverse problem. Future research could explore alternative formulations of $V$ that may yield solutions with more desirable theoretical properties. Furthermore, extending our analysis from the current fixed-dimensional setting to a high-dimensional framework presents an interesting avenue for research. Moreover, leveraging the projection approach proposed by escanciano2009simple,escanciano-2024-gaussian, which eliminates the estimation effect, could be advantageous in dealing with estimators whose limiting distributions are difficult to characterize, for example, the LASSO. Finally, extending our methodology to encompass time series data, clustered data, and multiple-equation models with (non)-smooth objective functions warrants future exploration.
\printbibliography
\setcounter{page}{1}
refsection