Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
54,714 characters · 11 sections · 22 citation commands
In regression analysis, conditional distributions play a fundamental role in constructing economic models. They fully characterize the relationship between a response variable and its covariates while enabling insights of crucial statistical indicators—including expectations, variance, quantiles, and hazard rates. These distributions form the core of numerous econometric frameworks, from categorical or discrete choice models (e.g., (ordered) logit/probit) to count data regression models (Poisson, negative binomial, hurdle, and zero-inflation models) and duration analyses (Cox proportional hazards and accelerated failure time models). In these settings, the conditional distribution is typically specified up to a finite-dimensional parameter. The validity of estimation and inference procedures crucially depends on the correct specification of this parametric form, making model checking essential for their reliable use.
Let $Y$ be a real-valued response variable and $X$ a vector of covariates. This paper investigates specification tests for assessing whether the conditional distribution of $Y$ given $X$ belongs to a parametric family indexed by an unknown parameter vector $\theta$. Formally, we consider testing
against its negation. The parameter space $\Theta$ is assumed to be a subset of $\mathbb{R}^p$, where $p$ is a fixed positive integer.
Although there is an extensive literature on consistent testing of conditional expectations—see, for example, bierens1982consistent,bierens1990consistent,bierens1997asymptotic,stute1997nonparametric,stinchcombe1998consistent,delgado2001significance,fan2000consistent,escanciano2006consistent,hardle1993comparing,horowitz1994testing,hong1995consistent,zheng1996consistent,lavergne2000nonparametric—the literature on consistent specification testing for conditional distributions is comparatively limited. Only a few papers have addressed this problem, including andrews1997conditional,zheng2000consistent,zheng2012testing,bierens2012integrated,amengual2020testing. Among these, andrews1997conditional and bierens2012integrated employ a similar methodology, which is based on transforming the conditional moment restriction into a continuum of unconditional moment conditions. This methodology is commonly known as the Integrated Conditional Moment (ICM) approach. In parallel, kernel-based tests have been developed by zheng2000consistent,zheng2012testing, which compare the parametric model to nonparametric or semiparametric estimates using kernel smoothing techniques.
Compared to kernel-based tests, a key advantage of the ICM approach is its ability to address the curse of dimensionality in covariates. Kernel-based tests are particularly vulnerable to the curse of dimensionality, as they require nonparametric kernel density estimation, which becomes increasingly unreliable in high dimensions. In contrast, the ICM framework naturally accommodates dimension reduction techniques. For instance, escanciano2006consistent proposes using projection weights based on indicator functions in the transformed unconditional moments, while bierens1982consistent advocates for characteristic function weighting.
However, the performance of ICM tests are still unsatisfactory as they often exhibit limited power in practice. While theoretically consistent—being omnibus tests capable of detecting all possible deviations—they generally lack detection power. A more critical limitation is that although these tests can identify general model misspecification, they fail to reveal which specific aspects of the data lead to rejection. This naturally raises an important question: Can the omnibus test be decomposed into components, each capturing a distinct aspect of the data? If possible, evaluating both the overall test and the significance of individual components would provide more informative diagnostics and deeper insights into the nature of detected deviations.
The idea of decomposing omnibus tests into orthogonal components originates from the work of durbin1972components, who introduced Principal Component Analysis (PCA) for the classical empirical process in the context of unconditional distribution goodness-of-fit testing. Through functional PCA, they achieved an orthogonal decomposition of the omnibus test statistic into principal components (PCs), with each PC capturing a distinct direction of departure from the null hypothesis. This approach was later extended by durbin1975components to composite hypothesis testing with estimated parameters. Building on these foundations, stute1997nonparametric further demonstrated the significant application of PCA in goodness-of-fit tests by discussing the issue of checking the mean regression model. He derived the PCs of the residual marked process, which serve as a basis for constructing more efficient tests, such as Neyman-type smooth tests and optimal directional tests.
While PCA is a powerful tool for enhancing test power, only a handful of studies have successfully derived the PCs of omnibus test statistics in explicit form. Most existing work merely establishes the existence of an orthogonal decomposition without providing its concrete representation. Following the foundational work of Durbin and Stute, we develop an extension of PCA application for general conditional distribution testing. However, previous PCA approaches—designed for univariate stochastic processes—prove inadequate in conditional distribution case. The introduction of covariates in distribution functions fundamentally alters the empirical process's structure, transforming it from univariate to multivariate. The resulting complex dependence structure in the multivariate process presents the primary challenge for implementing PCA in this setting.
To address this challenge, we propose a novel methodology called conditional Principal Component Decomposition (PCD). Our approach demonstrates that the residual-marked empirical process can be decomposed into a series of asymptotically independent component processes. Each component process consists of individual PCs and serves as the foundation for our new class of specification tests. Notably, the dimension reduction strategies developed for ICM tests can be seamlessly incorporated into our PCD framework, enabling effective handling of high-dimensional covariates.
PCA also provides a straightforward explanation for the underperformance of omnibus tests. In omnibus statistics, the contribution of each PC is weighted by its corresponding eigenvalue, which typically decays rapidly. As a result, higher-order components—those most sensitive to high-frequency departures from the null—are heavily downweighted and thus contribute little to the overall test statistic. This makes omnibus tests inherently less effective at detecting such high-frequency alternatives. By contrast, once the PCs are explicitly identified, we gain the flexibility to assign weights as desired: individual components can be used as targeted tests for specific types of deviations, while reweighted combinations (smooth tests) can be constructed to enhance power against a broader class of alternatives. In this sense, PCs serve as fundamental building blocks for more powerful and informative tests.
Despite this flexibility, a practical challenge remains: there is little theoretical guidance on how to optimally select and combine components for a given dataset. Traditionally, researchers have relied on the heuristic that the first few components capture most of the relevant departures from the null, but this approach may be unreliable in complex settings—such as those involving high-dimensional covariates—where the signal may be spread across many components. Consequently, smooth tests that simply aggregate the leading components can exhibit substantial variation in power. This underscores the need for data-driven methods to identify and utilize the most informative components, thereby making full use of the diagnostic potential offered by PCA.
In this paper, we proposes a component selection procedure based on a “learning before testing” strategy using sample splitting. In the learning sample, we estimate the p-values of individual component tests to gauge their significance, selecting those with the smallest p-values as the most informative. These selected components are then aggregated—typically with equal weights—into a smooth test statistic in the testing sample. Our simulation studies demonstrate that even this straightforward approach can substantially enhance testing power. Further research on optimal weighting schemes for the selected components is left for future work.
The remainder of the paper is organized as follows. In Section 2, after a brief reviewing of existing omnibus ICM tests for conditional distributions, we introduce our conditional PCD methodology and present the construction of component-based test statistics, and establishes their asymptotic properties. Section 3 describes the proposed component selection procedure. Section 4 investigates the finite-sample performance of the new tests through Monte Carlo simulations. We also apply the proposed methodology to a real data example in this section. Section 5 concludes.
Consider a sample $\{Y_i, X_i\}_{i=1}^n$ consisting of i.i.d. realizations of the random vector $(Y, X)$, where $Y$ is the real-valued response variable and $X$ is the $d$-dimensional covariate vector. Our objective is to assess the conditional distribution of $Y$ given $X$. To test $H_0$, the ICM approach introduces a suitable weighting function $h(v, X)$, allowing the conditional moment condition $$ \mathbb{E}\left[\boldsymbol{1}_{\{Y \leq y\}} - F(y|X, \theta_0) \mid X\right] = 0, \quad \text{a.s.} $$ to be equivalently reformulated as a continuum of unconditional moment conditions: $$ \mathbb{E}\left[\left(\boldsymbol{1}_{\{Y \leq y\}} - F(y|X, \theta_0)\right) h(v, X)\right] = 0, \quad \forall v \in \Pi, $$ for some index set $\Pi$. Several families of weighting functions satisfy this equivalence, with the most common choices being the characteristic function (see bierens1982consistent) and the indicator function (see andrews1997conditional, stute1997nonparametric). While the literature often uses $x$ as the argument to denote the weighting function's independent variable, we use $v$ to distinguish it from the actual values of the covariates, which will be important in later sections. We assume $v$ is $d'$-dimensional.
To test the unconditional moment conditions, which is considerably easier, we introduce the empirical process $$R_{n}(y,v):=n^{-1/2}\sum_{i=1}^{n}h(v,X_i) M_i(y), \eqno{(2.1)}$$ where $${M}_i(y)=\boldsymbol{1}_{\{Y_i\le y\}} - F \left(y | X_i, {\theta}\right). \eqno{(2.2)}$$ Here, the “residual” process $M_i$ is written separately, as it will play a central role in our subsequent decomposition. By analyzing the behavior of the ${R}_n$ process, we can make statistical inferences: if the process deviates sufficiently from zero, we reject the null hypothesis.
In fact, all existing ICM tests for conditional distributions are fundamentally based on the $R_{n}$ process, with the primary distinction being the choice of weighting function $h$. For instance, if $h$ is selected as the indicator function $\boldsymbol{1}_{\{X \leq v\}}$ with $v \in \mathbb{R}^d$, we obtain the process introduced by Andrews (1997): $${\alpha}_{n}(y,v) := n^{-1/2} \sum_{i=1}^{n} \boldsymbol{1}_{\{X_i \leq v\}} M_i(y). \eqno{(2.3)}$$ Its empirical version $$\hat{\alpha}_n(y,v) := n^{-1/2} \sum_{i=1}^{n} \boldsymbol{1}_{\{X_i \leq v\}} \hat{M}_i(y),$$ where $$\hat{M}_i(y) = \boldsymbol{1}_{\{Y_i \leq y\}} - F\big(y \mid X_i, \hat{\theta}\big),$$ and $\hat{\theta}$ is a consistent estimator of $\theta,$ is used to built Andrews' conditional Kolmogorov-Smirnov (CK) test, namely, $$CK = \sup_{y, v} \big| \hat{\alpha}_n(y, v) \big|. \eqno{(2.4)}$$
However, Andrews' test suffers from the curse of dimensionality due to the use of the indicator function. To address this, escanciano2006consistent proposed replacing the weighting function $\boldsymbol{1}_{\{X\le v\}}$ with projection weights, which mitigates the high-dimensionality problem. Although originally developed for regression models, this approach is also applicable to conditional distributions. Specifically, the empirical process becomes $$\hat{\zeta}_{n}(\beta,z,y)=n^{-1/2}\sum_{i=1}^{n}\boldsymbol{1}_{\{\beta^\top X_i\le z\}} \hat{M}_i(y),$$ where the weighting function is $\boldsymbol{1}_{\{\beta^\top X\le z\}}$ with $v=(\beta;z)\in \mathbb{R}^{d+1}$. A Cramér-von Mises (CvM) type statistic can be defined as $$CvM_{es}=\int \big(\hat{\zeta}_{n}(\beta,z,y)\big)^2 F_{n, Y}(dy) F_{n,\beta}(dz)d\beta, \eqno{(2.5)}$$ where $F_{n, Y}(y)$ and $F_{n,\beta}(z)$ are the empirical distributions of $Y$ and $\beta^\top X$, respectively, and $d\beta$ denotes the uniform measure on the unit sphere.
Another widely used weighting function, introduced by bierens1982consistent, is the characteristic function $\exp(i v^{\top} X)$ with $v \in \mathbb{R}^d$. In the context of conditional distributions, the corresponding CvM statistic is $$CvM_{ch}=\int\big(\hat{\eta}(y,v)\big)^2 F_{n,Y}(dy) \Upsilon(v)dv, \eqno{(2.6)}$$ where $$ \hat{\eta}(y,v) = n^{-1/2}\sum_{i=1}^{n}\exp(i v^{\top} X_i) \hat{M}_i(y), $$ and $\Upsilon(v)$ is the standard multivariate normal density.
All these tests share the ICM framework and are consistent against all possible alternatives, but the choice of weighting function can significantly affect their power performance. Escanciano (2006) argues that his projection-based test combines the advantages of both Andrews' and Bierens' approaches. It preserves the convenience of having a natural integration measure for CvM statistics offered by the indicator weighting function, and also has the ability to handle high-dimensional problems associated with exponential weighting.
Another concern of ICM tests lies in their low power against high-frequency alternatives. As fan2000consistent demonstrated, these tests can be viewed as kernel tests with a fixed bandwidth. In such cases, the fixed bandwidth may oversmooth the conditional distribution, masking important features of the alternative.
In this paper, we address the power limitations of omnibus tests for high-frequency alternatives from a different perspective. Rather than focusing solely on the choice of weighting function, we develop a functional PCA for the $R_{n}$ process, decomposing it into an infinite series of asymptotically independent component processes. Each component process captures deviations in a specific frequency direction, providing a foundation for constructing more powerful tests. Importantly, this decomposition is independent of the choice of $h$, thereby preserving the same flexibility in selecting weighting functions for the component processes as in the original $R_{n}$. As a result, our approach enables effective handling of both high-dimensional covariates and high-frequency alternatives. The details of the decomposition are presented in the next section.
Before proceeding, we introduce the key assumptions regarding the conditional distribution model and the estimation procedure.
(A1). The function $F(y| X_i,\theta)$ is differentiable with respect to $\theta$ in a neighborhood of $\theta_0$ for all $i \geq 1$.
for all sequences of positive constants $\{r_n : n\ge 1\}$ with $r_n\to 0$, where $\Delta_0(y,v)=\int (\partial / \partial \theta)F(y| s,\theta_0)h(v,s)\,dF_X(s)$ and $F_X$ denotes the marginal distribution of $X$.
(A3). $\sup_{(y,v)\in \mathbb{R}^{d'+1}} \| \Delta_0(y,v) \| <\infty$ and $\Delta_0(\cdot)$ is uniformly continuous on $\mathbb{R}^{d'+1}$.
(A4). The parametric estimator admits the expansion $$ n^{1/2}\big(\hat{\theta}-\theta_0\big) = n^{-1/2}\sum_{i=1}^{n}l(X_i,Y_i,\theta_0) + o_p(1), $$ for some function $l$ satisfying $\mathbb{E}[l(X,Y,\theta_0)] = 0$, and $L(\theta_0) = \mathbb{E}\left[l(X,Y,\theta_0)l^{\top}(X,Y,\theta_0)\right]$ exists and is positive definite.
Assumptions (A1)-(A3) are standard in the literature; see, for example, andrews1997conditional. Assumption (A4) requires that the maximum likelihood estimator satisfies the usual regularity conditions, which is also commonly imposed; see, for instance, bierens1982consistent,bierens1997asymptotic,bierens2012integrated,delgado2001significance.
Since the process $R_n$ is multivariate and involves dependent components $y$ and $v$, it does not admit a direct Karhunen-Loève representation. Nevertheless, thanks to the martingale structure of each ${M}_i$, a frequency-domain decomposition is still possible. Our conditional PCD approach leverages this fact that each ${M}_i$ is a centered process with a Brownian Bridge type covariance structure, hence admits an explicit PCD. The spectral decomposition of $R_n$ is then obtained by aggregating the PCs of all ${M}_i$ processes. Before presenting the details, we introduce some necessary notation.
Let $$\mu_j = \frac{1}{(\pi j)^2}, \qquad \varphi_j(y) = \sqrt{2}\sin(j\pi y), \quad j=1,2,\ldots \eqno{(2.7)}$$ be the eigenvalues and eigenfunctions of the standard Brownian Bridge $B(t)$, whose covariance kernel is $K(s,t) = s \wedge t - st$, with $s \wedge t = \min(s,t)$. For each $x$, introduce the transformation $$T(y, x) := F(y \mid X = x, \theta_0), \eqno{(2.8)}$$ where $\theta_0$ is the true parameter value. The function $T$ is non-decreasing in $y$, with $T(-\infty, x) = 0$ and $T(\infty, x) = 1$. Applying this transformation to $\varphi_j$, we define compound functions $$f_j(y,x):=\varphi_j(T(y,x)), \quad j=1,2,\ldots$$ They are actually the eigenfunctions of transformed Brownian Bridge $B(T(y,x))$ with fixed $x$. Also, their cosine counterpart are important, thus we define $$ g_j(y,x):=\sqrt{2}\cos(j\pi T(y,x)).$$
For each fixed $x$, the collection $\{f_j(\cdot, x)\}_{j=1}^\infty$ forms an orthonormal basis for a subspace of $L^2(\mathbb{R}, T(\cdot, x))$, the Hilbert space of square-integrable functions on $\mathbb{R}$ with an inner product $$\langle \rho, g \rangle_x = \int_{\mathbb{R}} \rho(y) g(y) T(dy, x).$$ Indeed,
Moreover, $\{f_j(\cdot, x)\}_{j=1}^\infty$ are the eigenfunctions of $M(y)=\boldsymbol{1}_{\{Y \leq y\}} - F\big(y \mid X, {\theta_0}\big)$ conditional on $X = x$, with corresponding eigenvalues $\{\mu_j\}_{j=1}^\infty$. This result follows from the conditional covariance kernel of $M(y)$:
It is identical to the covariance kernel of $B(T(y,x))$, therefore the eigenfunctions of $M(y)$ are given by the transformation of $\varphi_j$.
By Mercer's theorem, the covariance kernel of $M(y)$ admits the following spectral decomposition: $$K(T(y_1, x), T(y_2, x)) = \sum_{j=1}^{\infty} \mu_j f_j(y_1, x) f_j(y_2, x). \eqno{(2.9)}$$ The PCD of each $M_i$ is then straightforward. By substituting all these PCDs into $R_n$, we obtain a decomposition of $R_n$ itself. This result is formalized in the following proposition, with the proof provided in the appendix.\\
We now focus on the component processes $c_{n,j}$, as they will form the basis for our proposed tests. Each $c_{n,j}$ is a sum of i.i.d. terms involving the weighting function $h(v, X_i)$, the eigenfunction $f_j(y, X_i)$, and a random variable $g_j(Y_i, X_i)$. The variable $g_j(Y_i, X_i)$ is the $j$-th normalized principal component of $M_i$ conditional on $X_i$, and thus possesses the standard orthogonality properties of principal components, but conditional on $X_i$. Specifically, for each $j$ and $j \neq h$,
The i.i.d. structure of the component processes $c_{n,j}$ is also useful for deriving their asymptotic properties. Each $c_{n,j}$ is a sum of i.i.d. centered random functions with variance $$ H_j(y,v) := \mathbb{E}\left[ h^2(v,X) f_j^2(y,X) \right] = \int h^2(v,s) f_j^2(y,s) F_X(ds), $$ where $F_X(\cdot)$ denotes the distribution function of $X$. The following theorem describes their asymptotic behavior.
The component processes derived above form the foundation for a new class of specification tests for conditional distributions. Our proposed tests are based on their empirical counterparts, defined as $$ \hat{c}_{n,j}(y,v) := n^{-1/2} \sum_{i=1}^{n} h(v, X_i) \hat{f}_j(y, X_i) \hat{g}_j(Y_i, X_i). \eqno{(2.11)} $$ Here, $\hat{f}_j(y, x) = \sqrt{2} \sin\big(j\pi \hat{T}(y, x)\big)$ and $\hat{g}_j(y, x) = \sqrt{2} \cos\big(j\pi \hat{T}(y, x)\big)$, where $\hat{T}(y, x)$ is an estimator of $T(y, x)$. A natural and consistent choice is $$ \hat{T}(y, x) = F\big(y \mid X = x, \hat{\theta}\big). $$
Given a specific weighting function $h$ in (2.11), we can construct CvM type statistics for each component process. Specifically, we introduce our component test statistic as $$ CvM_{n,j} = \int \left(\hat{c}_{n,j}(y, v)\right)^2 \Psi_{n}(dy, dv), ~~j = 1, 2, \ldots, \eqno{(2.12)} $$ where $\Psi_{n}(dy, dv)$ is the empirical measure determined by the choice of $h$ in $\hat{c}_{n,j}$. For example, if Escanciano's projection weight $\boldsymbol{1}_{\{\beta^\top X_i \leq z\}}$ is used, $\Psi_{n}$ corresponds to $F_{n, Y}(dy) F_{n, \beta}(dz) d\beta$ as in (2.5); if Bierens' characteristic function $\exp(i v^{\top} X_i)$ is used, it becomes $F_{n, Y}(dy) \Upsilon(v) dv$ as in (2.6).
Beyond individual component tests, one can construct test statistics by combining several component processes. Specifically, consider the first $m$ component processes, reweighted by a sequence $w = (w_1, w_2, \ldots, w_m)$. The corresponding CvM type statistic\footnote{Analogously, KS type statistics can be defined as $$KS_{n,j} = \sup_{y,v} \big| \hat{c}_{n,j}(y,v) \big|$$ and $$KS_{n,m}^w = \sup_{y,v} \left| \sum_{j=1}^{m} w_j \hat{c}_{n,j}(y,v) \right|.$$} is given by $$ CvM_{n,m}^w = \int \left( \sum_{j=1}^{m} w_j \hat{c}_{n,j}(y,v) \right)^2 \Psi_{n}(dy, dv). \eqno{(2.13)} $$ A smooth test, in the spirit of Neyman's smooth test, can be constructed by setting $w = (1, \ldots, 1)$, i.e., by equally weighting the first $m$ components. Smooth tests offer a balance between omnibus tests and individual component tests, and are often effective against a broad class of alternatives. However, the choice of the weight vector $w$ is critical in practical applications, an issue we will address in later section.
We finish this part by providing the asymptotic theories of the estimated component processes and their reweighted combinations. Formal proofs are in the Appendix.\\
For the reweighted process $\sum_{j=1}^{m} w_j \hat{c}_{n,j}(y,v)$, the asymptotic distribution coincides with that of $\sum_{j=1}^{m} w_j \tilde{c}_{n,j}(y,v)$. Its limiting behavior follows directly.\\
The convergence of our CvM test statistics then follows immediately by the continuous mapping theorem. In practice, following andrews1997conditional, we estimate the critical values of these tests using a parametric bootstrap procedure.
A crucial remark that need to be made is that our proposed tests—including the component tests $CvM_{n,j}$ and the smooth tests $CvM_{n,m}^w$—are generally not consistent, as they are only parts of the omnibus tests in certain spectral directions. However, in practice they often outperform the omnibus tests. As equation (2.10) shows, the weight $(\pi j)^{-1}$ assigned to the $j$-th component process decays rapidly with increasing $j.$ Consequently, higher-order components contribute progressively less to $R_n$, making the omnibus test less sensitive to deviations captured by these components. Importantly, these higher-order components detect higher-frequency departures, which explains the omnibus test's limited power against high-frequency alternatives. In contrast, individual component tests are specifically designed to detect such particular types of deviations. To demonstrate this specialized detection capability, we present an example in the following section that illustrates how different components serve as targeted detectors for specific types of departures from the null.
In the context of unconditional normal distributions, durbin1972components observed that different PCs are sensitive to departures in specific moments: the first PC is most responsive to mean shifts, the second to variance, the third to skewness, and the fourth to kurtosis. This insight led to the recommendation that each PC be examined individually to better understand the nature of any detected deviation.
To investigate whether this correspondence between PCs and moments extends to conditional distributions, we consider a conditional normal model with a linear conditional mean and constant variance: $$ H_0: Y \mid X \sim \mathcal{N}(\theta_0 + \theta_1 X, \theta_2), $$ where $\theta = (\theta_0, \theta_1, \theta_2)$. For simplicity, we let $X$ be real-valued and construct several data generating processes (DGPs) that introduce deviations in mean, variance, skewness, and kurtosis, respectively.
The first two DGPs are specified as follows:
For skewness and kurtosis, we use a conditional distribution function of the form $$ F(y \mid x, \gamma_3, \gamma_4) = \Phi(y - 1 - x) + \gamma_3 \sin(3\pi \Phi(y - 1 - x)) + \gamma_4 \sin(4\pi \Phi(y - 1 - x)), $$ where $\Phi(\cdot)$ denotes the standard normal cumulative distribution function. This construction, adapted from (8.4) in durbin1975components, allows $\gamma_3$ and $\gamma_4$ to control the degree of skewness and kurtosis, respectively:
We evaluate the performance of the first four component tests, using Andrews' indicator weighting function (appropriate for the univariate covariate case). Critical values are obtained via standard parametric bootstrap, as detailed in Section 5. Table (ref) reports the power of each component test, with the highest power for each scenario highlighted in bold.
We observe the same relationship pattern between PCs and moments in the conditional normal case as that recorded in durbin1972components for the unconditional setting. The first four component tests are each most sensitive to deviations in the first four moments, respectively. These component tests serve as experts for detecting moment departures from conditional normality.
Moreover, in many practical situations, deviations tend to be concentrated in the first few components, making smooth tests that assemble these leading components particularly effective. However, this experience and the above pattern may fail in high-dimensional settings—even when using projection weights (Escanciano) or characteristic function weights (Bierens). In such cases, the first few components may not be the most informative, and the clear correspondence between PCs and moments, as seen in the normal case, can break down. For example, Table (ref) in Appendix B shows that for two high-dimensional normal DGPs with different mean shift scales, the third and eighth components are most powerful for detecting mean deviations.
Consequently, when prior knowledge about the roles of PCs is unreliable, identifying which specific components capture the most significant deviations in a given dataset becomes crucial, as this can substantially improve testing efficiency. To address this challenge, we introduce a “learning then testing” procedure in the next section, which provides a data-driven way to identify the most informative components, thereby guiding the construction of more efficient tests.
In this section, we present a straightforward, data-driven component selection strategy. The idea is to split the full sample into two parts: a learning sample and a testing sample. In the learning sample, we estimate the p-values for each component test by applying a parametric bootstrap to approximate the empirical null distribution. The components are then ranked according to their p-values, and the $m$ components with the smallest p-values are selected for further analysis. These selected components are subsequently used to construct the test statistic in the testing sample. The following pseudo-code outlines this procedure.
During the learning procedure, it is important to assess the stability of the selected components, particularly the variability of the estimated p-values. Several strategies can be employed to enhance the robustness of the learning process. A straightforward approach is to increase the size of the learning sub-sample, which may reduce the variance of the p-values. However, this comes at the cost of a smaller testing sub-sample, potentially diminishing the power of the final test. Determining the optimal split between learning and testing samples is a nontrivial problem and is left for future research.
Another strategy is to use bagging. This involves repeatedly resampling the learning sub-sample and applying Algorithm (ref) to obtain multiple sets of p-values for each component. Averaging these p-values across resamples can reduce their variability and improve the reliability of component selection. The main drawback of this approach is the increased computational burden.
In the testing procedure, a key unresolved issue is the choice of weights assigned to each component in the smooth test statistic. In this paper, we adopt the simple approach of assigning equal weights to all selected components. Ideally, components with smaller p-values—indicating stronger evidence against the null—should receive larger weights, so that their influence on the test statistic is amplified. However, the choice of weights also affects the variance of the test statistic, and inappropriate weighting may actually reduce power. At present, there is limited theoretical guidance on optimal weight selection, and further research is needed to address this question.
In this section, we evaluate the performance of our proposed test statistics using Monte Carlo simulations. The data-generating processes (DGPs) considered are as follows:
For all DGPs, the covariates $X$ are generated independently from the uniform distribution $U(0,1)$, with dimension $d=15$. All coefficients are set to one, i.e., $\beta = (1,\ldots,1)^\top$.
DGP1 examines deviations in both mean and variance, while DGP2 introduces a distributional deviation. DGPs 3--5 focus on mean deviations with increasing frequency, and DGPs 6--8 are designed to illustrate the benefits of the learning-based procedure in detecting more complex alternatives.
We compare the proposed smooth test, both with and without the learning-then-testing procedure. The characteristic function is used as the weighting function $h$. The default smooth test, denoted $Smooth\_m$, is constructed by equally weighting the first $m$ components from the full sample. Alternatively, $Smooth\_Learn\_m$ is computed in the testing sub-sample, using the $m$ components identified in the learning phase as most indicative of deviation. In both cases, components are equally weighted.
For benchmarking, we also report results for three omnibus test statistics: (i) Andrew's test ($CK$), (ii) the CvM test with Escanciano's projection weighting ($CvM_{es}$), and (iii) the CvM test with characteristic function weighting ($CvM_{ch}$), as defined in equations (2.4)--(2.6).
Table (ref) reports the estimated empirical size at the 95% confidence level for sample sizes ranging from 100 to 300. The results indicate that the proposed component-based tests, including those using the learning procedure, exhibit only minor size distortions, which diminish as the sample size increases.
Table (ref) summarizes the power results, from which several key observations emerge.
First, component-based smooth tests generally outperform omnibus tests, but their effectiveness depends critically on the choice of component weights. For instance, in DGP3-5, all three omnibus tests ($CK$, $CvM_{ch}$, and $CvM_{es}$) exhibit negligible power, whereas the smooth tests—whether constructed from the first 5 or 10 components or using the learning procedure—demonstrate substantially higher power. In contrast, for DGP6-8, the omnibus tests ($CvM_{ch}$ and $CvM_{es}$) outperform the smooth tests, which show relatively poor power. This pattern can be explained by examining the power of individual component tests (see Appendix Tables (ref)--(ref)): for these DGPs, only the first two components capture significant deviations. As a result, omnibus tests, which assign higher weights to the leading components, are more powerful than smooth tests that distribute equal weights across a larger set of components.
Second, incorporating the learning procedure generally improves power. Although the testing sample is reduced in size, selecting components based on their performance in the learning sample leads to higher power than simply using the first 5 or 10 components. This improvement is especially pronounced in DGP3-8, where the underlying alternatives are more complex and not well detected by the default leading components.
Finally, the selection and weighting of components play a pivotal role, as illustrated by the results for DGP6-8 in Table (ref). In these scenarios, the first two components exhibit markedly lower p-values than the others, indicating that they capture the primary deviations from the null hypothesis. This observation motivates the construction of the $Smooth\_Learn\_2$ statistic, which aggregates only these two components. The simulation results confirm that focusing on the most informative components can substantially enhance power.
Moreover, the choice of weights assigned to each component is equally important. A comparison between the power of $Smooth\_Learn\_2$ and the individual two-component tests reveals that the weighting scheme can significantly affect test performance. While equal weighting is a practical default, identifying optimal weights remains a challenging and open problem for future research.
We apply the proposed component-based testing strategies to a real-world dataset from the UCI Machine Learning Repository lichman2017uci. The “Student Performance” dataset contains information on student dropout, including 36 covariates for 4,424 observations. The response variable is binary, indicating dropout status. We utilize a probit model to estimate the dropout probability conditional on the covariates.
Due to the large sample size, we randomly partition the dataset into a learning sample (40% of the data) and a testing sample (60% of the data). We consider the first 10 components. Using the learning sample, we identify the 5 components exhibiting the smallest p-values. These selected components are then used to construct the smooth test statistic within the testing sample. For comparison, we also report results for the smooth test constructed without the learning procedure, using the first 5 and the first 10 components, respectively. Additionally, we apply the CK test and the CvM test that utilizes the characteristic function as the weighting function. The CvM test with Escanciano's projection weighting is not included, as it is rather time consuming for this dataset.
The empirical results are summarized in Table (ref). The component-based test statistics consistently reject the null hypothesis, exhibiting p-values effectively equal to zero. In contrast, the non-component-based tests fail to reject the null, with p-values close to unity. This outcome highlights the superior power performance of the proposed component-based methods relative to the non-component alternatives for this dataset. Interestingly, the learning procedure offered no significant improvement to the smooth test's power. As indicated by the component analysis in Table (ref), the primary deviation signals appear concentrated within the first 10 components, thus limiting the potential benefit of the component selection step.
This paper focuses on two central questions in the context of testing conditional distributions: (i) whether omnibus statistics can be decomposed into a set of individual components, each capturing a distinct aspect of model fit; and (ii) how to identify which components are most informative about departures from the null hypothesis. A third, related question—how to optimally combine these components into a new, more powerful test statistic—is left for future research.
We address the first two questions by developing a conditional principal component analysis (PCA) of the omnibus test statistic. The resulting components provide a principled basis for constructing more informative specification tests. To determine the relative importance of each component, we propose a simple, data-driven learning procedure: we compute directional test statistics for each component and use bootstrap p-values to assess their significance. Components with small p-values are interpreted as carrying strong signals of model misspecification, while those with large p-values are likely to reflect noise.
Ideally, one would assign weights to each component proportional to the strength of its signal. However, determining optimal weights is a challenging problem that lies beyond the scope of this paper. As a practical solution, we adopt equal weighting of the selected components, which, despite its simplicity, yields substantial improvements in testing power in our empirical studies.
\vskip 0.2in