Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,498 characters · 15 sections · 74 citation commands
Generalized Spectral Testing with Sample Splitting
{\it Keywords:} Goodness-of-fit; Sample-splitting; Generalized spectral test; Bootstrap.
Diagnostic checking is a central step in time series analysis. Once a parametric dynamic model has been fitted, any subsequent inference, forecasting exercise, or structural interpretation rests on the maintained assumption that the conditional mean has been adequately specified. If the fitted model leaves systematic predictability in the residuals, parameter estimates may still be numerically stable, but the resulting inferential and forecasting conclusions can be misleading. This motivates formal goodness-of-fit tests that examine whether the residuals satisfy the martingale difference property implied by a correctly specified conditional mean.
We consider a strictly stationary and ergodic time series $\{(Y_{t}, \mathbf{Z}_{t-1}^{\prime})\}_{t \in \mathbb{Z}}$ defined on the probability space $(\Omega, \mathcal{F}, P)$. Here $Y_t\in\mathbb R$ denotes the response variable, and $\mathbf Z_{t-1}=(Y_{t-1},\mathbf X_{t-1}')'\in\mathbb R^m$ collects the lagged response and other predetermined explanatory covariates. Let $\mathbf I_{t-1}=(\mathbf Z_{t-1}',\mathbf Z_{t-2}',\ldots)'$ denote the full information set available at time $t-1$, and let $\mathcal F_{t-1}=\sigma(\mathbf I_{t-1})=\sigma(\mathbf Z_s:s\le t-1)$ be the $\sigma$-field generated by $\mathbf{I}_{t-1}$. Under the integrability of $Y_{t}$, the model can be written as
where $m(\mathbf{I}_{t-1})=E[Y_{t} \mid \mathbf{I}_{t-1}]$ is the conditional mean almost surely (a.s.) given the conditioning set $\mathbf{I}_{t-1}$, and $\varepsilon_{t}=Y_{t}-E[Y_{t} \mid \mathbf{I}_{t-1}]$ is a martingale difference sequence (MDS) with respect to $\mathbf{I}_{t-1}$ as the error term.
In this paper, we ask whether there exists a parametric family of functions $\mathcal{M}=\{f(\cdot, \boldsymbol{\theta}): \boldsymbol{\theta} \in \Theta \subset \mathbb{R}^{p} \}$, such that
or, equivalently, \[ H_0:\quad E\{e_t(\boldsymbol{\theta}_0)\mid \mathbf{I}_{t-1}\}=0\ \text{a.s.}, \] where $e_{t}(\boldsymbol{\theta})=Y_{t}-f(\mathbf{I}_{t-1},\boldsymbol{\theta})$ denotes the parameter induced model error.
Thus, goodness-of-fit can be formulated as a martingale difference hypothesis for the model errors. This formulation is especially natural in dynamic regression settings, where the relevant null is not merely the absence of linear autocorrelation, but the absence of any remaining conditional mean predictability.
Classical diagnostic tools are often based on residual autocorrelations and portmanteau statistics. These procedures remain attractive because they are simple, familiar, and easy to implement; see, among others, BoxPierce1970, Hosking1980, and LiMcLeod1981. Their scope, however, is limited. Portmanteau-type checks are primarily geared toward linear serial dependence, and their standard reference distributions are often justified under innovation assumptions that are stronger than the null one would actually like to test. In particular, when the residuals are uncorrelated but still dependent, as may happen under conditional heteroskedasticity, the usual chi-squared calibration can fail; see Romano1996. This mismatch is important in many nonlinear and heteroskedastic time series models, where the martingale difference null is the appropriate benchmark.
These concerns have led to a rich literature on omnibus specification testing for time series models. Hong1999 introduced generalized spectral methods for detecting serial dependence through empirical characteristic functions, and HongLee2005 extended this line of work to conditional mean models with conditional heteroskedasticity of unknown form. Particularly relevant for the present paper is the bandwidth-free generalized spectral test of escanciano2006goodness, which provides a Cram\'er-von Mises type diagnostic for linear and nonlinear conditional mean models by integrating a continuum of pairwise moment restrictions. Subsequent work has broadened the scope of these ideas to joint and marginal checks for conditional mean and variance models Escanciano2008, as well as to specification testing for conditional distribution models ChenHong2014. More recently, Escanciano2024 has recast model checking in a broader conditional moment restriction framework. In a related direction, WangZhuShao2022 developed MDDM (Martingale Difference Divergence Matrix) based tests for the martingale difference hypothesis in multivariate time series models, and also provided a detailed discussion about the similarity and distinction between the MDDM-based approach and the one by escanciano2006goodness.
A recurring technical difficulty in this literature is the estimation effect. Because the residuals used for diagnostic checking are computed from the same data that are used to estimate the unknown parameter, residual-based empirical processes do not, in general, behave as if the true innovations were observed. For many non-smoothing-based tests, this first-order estimation effect is asymptotically non-negligible, which leads to nonpivotal null distributions and complicates implementation. Existing remedies typically rely either on bootstrap procedures that repeatedly generate pseudo-data and refit the model, or on martingale-transform arguments designed to remove nuisance-parameter effects; see, for example, KoulStute1999, KhmaladzeKoul2004, and Bai2003. While these approaches are theoretically powerful, they may be computationally demanding and can make routine model checking less convenient in practice.
Recent work by davis2025sample shows that sample splitting offers an appealing alternative in time series goodness-of-fit problems. They generalize the half-sample splitting device proposed by durbin1976kolmogorov from an independent data setting to the time series setting. The idea is to estimate the model parameters on one subsample and to assess goodness-of-fit on another, while allowing for a carefully chosen asymptotic overlap between the two parts of the sample. For diagnostics based on the autocorrelation function and the auto-distance correlation function, they show that the resulting residual-based statistics can have the same asymptotic behavior as in the infeasible oracle case, where the innovations are observed if the split-ratio between the fitting sample length and testing sample length is 1/2. This insight is important, but their framework is tailored to tests of residual serial independence and, in the portmanteau setting, it works under an independently and identically distributed (i.i.d.) innovation paradigm that is stronger than the martingale difference null. From the standpoint of conditional mean adequacy, it is therefore natural to ask whether the same sample-splitting idea can be combined with a diagnostic whose null hypothesis is aligned with the martingale difference property itself.
The goal of this paper is to answer that question. We develop a sample-splitting-based generalized spectral test for the goodness-of-fit of parametric conditional mean models. More specifically, we combine the integrated generalized spectral/Cram\'er-von Mises framework of escanciano2006goodness with the sample-splitting strategy of davis2025sample. One part of the sample is used for estimation, while a checking/testing subsample is used to form residuals based on the fitted parameter and then to evaluate pairwise conditional mean restrictions of the form \[ E\{e_t(\boldsymbol{\theta}_0)w(\mathbf Z_{t-j},\mathbf x)\}=0, \qquad j\geq 1, \] using appropriately chosen weighting functions $w(\mathbf Z_{t-j},\mathbf x)$ indexed by lagged explanatory variables $\mathbf{Z}_{t-j}$ and a continuum of values $\mathbf{x}$. In this way, the proposed procedure preserves the omnibus and bandwidth-free nature of generalized spectral diagnostics, while using sample splitting to separate model fitting from model assessment. As in escanciano2006goodness, the resulting test targets pairwise conditional mean dependence rather than full joint conditional mean dependence on the entire past.
The paper makes several contributions. First, it develops a general split-sample generalized spectral testing framework for linear and nonlinear time series models with parametric conditional means. Second, under suitable regularity conditions and an overlap condition linking the fitting and checking subsamples, we show that the split-sample generalized spectral process has the same asymptotic null distribution as the corresponding infeasible oracle process based on the true errors. Thus, sample splitting eliminates the first-order parameter-estimation effect from the limiting null law. Third, although the limiting null distribution remains nonpivotal, it can be consistently approximated by a simple multiplier bootstrap applied directly to the split-sample residuals. This avoids generating bootstrap time series and re-estimating the model in each bootstrap replication, and therefore leads to a markedly lighter computational procedure. This is a major difference from escanciano2006goodness, and the computational gain can be substantial. Similar to the latter paper, our procedure is also tuning parameter-free. Fourth, we establish consistency against fixed alternatives, derive nontrivial power against local alternatives, and verify the required conditions for important model classes, including ARMA and GARCH models.
The proposed methodology complements several existing strands of literature. Relative to davis2025sample, our focus is not residual serial independence per se, but the martingale difference hypothesis implied by a correctly specified conditional mean. Relative to escanciano2006goodness, the contribution is not a new generalized spectral metric, but rather a new inferential device that uses sample splitting to simplify the effect of parameter estimation and to facilitate bootstrap implementation. In this sense, the paper should be viewed as complementary to both earlier MDS testing work and the recent sample-splitting literature.
The rest of the paper is organized as follows. Section (ref) proposes the split-sample generalized spectral process and the associated Cram\'er-von Mises statistic. Section (ref) develops the asymptotic theory under the null hypothesis, fixed alternatives, and local alternatives. Section (ref) presents the multiplier bootstrap approximation algorithm and its theoretical justification. Section (ref) reports extensive simulation studies on the finite-sample performance of the proposed tests, including linear ARMA models, nonlinear volatility models, and threshold autoregressive models. Section (ref) presents empirical illustrations based on S&P 500 dynamics and annual Wolf's sunspot data, demonstrating the validity and computational advantage of the proposed tests. Section (ref) concludes.
This section constructs the sample-splitting generalized spectral test for (ref). The population restrictions and their integrated spectral aggregation follow the generalized spectral approach of escanciano2006goodness: the null is represented by a continuum of pairwise conditional-mean restrictions, and the test measures whether the corresponding integrated spectral distribution is flat across the frequency index. The implementation, however, is different in an important way. In escanciano2006goodness, the parameter is estimated from the full sample, and the same full sample is then used to construct the residual-based test statistic. Here, the parameter is estimated on a fitting subsample, whereas the generalized spectral process is formed from residuals on a checking subsample.
We first introduce the generalized spectral test in escanciano2006goodness. Under $H_0$, the model error satisfies $e_t(\boldsymbol{\theta}_0)=\varepsilon_t$ and hence has zero conditional mean given the full past. In particular, for every lag $j\ge 1$,
Thus, a correctly specified conditional mean implies an infinite collection of pairwise restrictions, one for each lagged vector $\mathbf Z_{t-j}$. To make these restrictions operational, let $\{w(\mathbf Z_{t-j},\mathbf x):\mathbf x\in\Upsilon\subset\mathbb R^s\}$ be a family of bounded weighting functions such that the conditional restriction in (ref) is equivalently characterized by
As discussed later, usual choices of weight functions $w$ include $w (\mathbf{Z}_{t-j}, \mathbf{x})=\mathbbm{1} (\mathbf{Z}_{t-j} \leq \mathbf{x} )$ with $\mathbf{x} \in[-\infty, \infty]^{m}$, where $\mathbbm{1}(A)$ denotes the indicator function of the event $A$, and $w (\mathbf{Z}_{t-j}, \mathbf{x} )=\exp (i \mathbf{x}^{\prime} \mathbf{Z}_{t-j} )$ with $\mathbf{x} \in \mathbb{R}^{m}$, where $i=\sqrt{-1}$ is the imaginary unit.
The sequence $\{\gamma_{j,w}(\cdot,\boldsymbol{\theta}_0):j\ge1\}$ measures departures from the pairwise martingale-difference implications across all lags. To aggregate these restrictions, define $\gamma_{-j,w}(\cdot,\boldsymbol{\theta}_0)=\gamma_{j,w}(\cdot,\boldsymbol{\theta}_0)$ for $j\ge1$ and consider the generalized spectral transformation
Therefore, the tests are based on an integrated Fourier transform of the measures $ \{\gamma_{j, w} (\cdot, \boldsymbol{\theta}_{0} ) \}_{j=1}^{\infty}$. Under $H_0$, all nonzero-lag coefficients vanish, so $$f_w(u,\mathbf x,\boldsymbol{\theta}_0) \equiv f_{0, w} (\mathbf{x}, \boldsymbol{\theta}_{0} )=(2 \pi)^{-1} \gamma_{0, w} (\mathbf{x}, \boldsymbol{\theta}_{0} ).$$ Equivalently, the integrated generalized spectral distribution
reduces to $\gamma_{0,w}(\mathbf x,\boldsymbol{\theta}_0)\lambda$ under $H_0$. Therefore, any departure from this constant function implies a violation of at least one of the pairwise restrictions.
We now describe the proposed sample-splitting implementation. Suppose the observed sample is $\{Y_t,\widehat{\mathbf I}_{t-1}\}_{t=1}^n$, where $\widehat{\mathbf I}_{t-1}=(\mathbf Z_{t-1}',\mathbf Z_{t-2}',\ldots,\mathbf Z_0')'$ is the observed information set and may contain initial values. Let $f_n$ and $l_n$ be two non-decreasing sequences diverging with $n$. The first $f_n$ observations are used for estimation, and the last $l_n$ observations are used for model checking:
Let $\widehat{\boldsymbol{\theta}}_{f_n}$ denote the estimator computed from the fitting sample. On the checking sample, define the split-sample residuals \[ \widehat e_t\equiv \widehat e_t(\widehat{\boldsymbol{\theta}}_{f_n}) =Y_t-f(\widehat{\mathbf I}_{t-1},\widehat{\boldsymbol{\theta}}_{f_n}), \qquad t=n-l_n+1,\ldots,n. \] Note that, for lag $j$, the effective sample size is $n_j=l_n-j+1.$ The empirical lag-$j$ weighted moment is
Hence, the sample analogue of the generalized spectral distribution function (ref) is given by
where $ (n_{j} / l_n )^{1 / 2}$ is a finite-sample correction factor that does not affect the asymptotic theory and delivers a better finite-sample performance of the test procedure.
Under $H_0$, the target function is $H_w(\lambda,\mathbf x,\boldsymbol{\theta}_0) =\gamma_{0,w}(\mathbf x,\boldsymbol{\theta}_0)\lambda$ and thus tests can be based on the discrepancy between $\widehat{H}_{w} (\lambda, \mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} )$ and $\widehat{H}_{0, w} (\lambda, \mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} ) \equiv \widehat{\gamma}_{0, w} (\mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} ) \lambda$. Define the centered process
Finally, we aggregate the process over $\Pi=[0,1]\times\Upsilon$ using a Cram\'er-von Mises norm:
where $W$ is an integrating distribution for the index $\mathbf x$. The test rejects $H_0$ for large values of $D_{n,w}^2(\widehat{\boldsymbol{\theta}}_{f_n})$. Because all lags in the checking sample are included with deterministic spectral weights, the procedure does not require choosing a truncation lag or bandwidth.
This section studies the large-sample behavior of the split-sample generalized spectral process introduced in Section (ref). The main message is that, under a simple score-alignment condition (ref) for the data generating process and a half split-ratio condition on the length of the fitting and checking samples, the residual-based process has the same null limit as the infeasible oracle process that would be computed from the true innovations. We then show that the corresponding Cram\'er-von Mises statistic is consistent against fixed pairwise alternatives and has nontrivial power against local alternatives.
To begin with, we introduce the notation used throughout the asymptotic theory. Let \[ \mathbf{g}_{t}(\boldsymbol{\theta}) \equiv \mathbf{g}(\mathbf{I}_{t-1},\boldsymbol{\theta}) = \frac{\partial}{\partial\boldsymbol{\theta}} f(\mathbf{I}_{t-1},\boldsymbol{\theta}), \qquad w_{t-j}(\mathbf{x})\equiv w(\mathbf Z_{t-j},\mathbf x). \] Recall that \[ \varepsilon_t=Y_t-E(Y_t\mid\mathbf I_{t-1}), \qquad e_t(\boldsymbol{\theta})=Y_t-f(\mathbf I_{t-1},\boldsymbol{\theta}), \] so that under $H_0$, $e_t(\boldsymbol{\theta}_0)=\varepsilon_t$ a.s. Put $\Pi=[0,1]\times\Upsilon$ and write $\boldsymbol{\eta}=(\lambda,\mathbf x')'\in\Pi$.
The process $S_{n,w}(\boldsymbol{\eta},\widehat{\boldsymbol{\theta}}_{f_n})$ is viewed as a random element of the Hilbert space $L_2(\Pi,\nu)$, where $\nu$ is the product of the integrating measure $W$ on $\Upsilon$ and Lebesgue measure on $[0,1]$. For complex-valued functions $f,g \in L_2(\Pi,\nu)$, the inner product is \[ \langle f,g\rangle = \int_{\Pi}f(\boldsymbol{\eta})g^c(\boldsymbol{\eta})\,d\nu(\boldsymbol{\eta}) = \int_{\Pi}f(\lambda,\mathbf x)g^c(\lambda,\mathbf x)\,W(d\mathbf x)d\lambda, \] and $\|f\|=\langle f,f\rangle^{1/2}$. We equip $L_2(\Pi,\nu)$ with the Borel $\sigma$-field generated by this norm and write $\Longrightarrow$ for weak convergence in this space. For a mean-zero $L_2(\Pi,\nu)$-valued random element $Z$ with $E\|Z\|^2<\infty$, its covariance operator is $C_Z(h)=E[\langle Z,h\rangle Z]$, $h\in L_2(\Pi,\nu)$.
Let \[ \Psi_j(\lambda)=\frac{\sqrt{2}\sin(j\pi\lambda)}{j\pi}, \qquad \mathbf b_j(\mathbf x,\boldsymbol{\theta}_0) =E[w_{t-j}(\mathbf x)\mathbf g_t(\boldsymbol{\theta}_0)], \] and define the score-loading function \[ \mathbf G_w(\boldsymbol{\eta}) \equiv \mathbf G_w(\boldsymbol{\eta},\boldsymbol{\theta}_0) = \sum_{j=1}^{\infty} \mathbf b_j(\mathbf x,\boldsymbol{\theta}_0)\Psi_j(\lambda). \] The covariance of the oracle generalized spectral limit is described through the quadratic form
where $h \in L_{2}(\Pi, \nu)$, $\boldsymbol{\eta}_{1}=(\lambda, \mathbf{x}^{\prime})^{\prime}$, and $\boldsymbol{\eta}_{2}= (\varpi, \mathbf{y}^{\prime})^{\prime}$.
The following assumptions collect the necessary regularity conditions.
Assumptions (ref)--(ref) are inherited from escanciano2006goodness, and are deliberately stated at a high level. Assumption (ref)-(ref) are standard but mild conditions on the data-generating process. Assumption (ref) (a) assumes a pseudo-true value under the alternative, which is standard in the theory of misspecified likelihood and other extremum estimation; see, e.g. newey1994large, and white1982maximum,white1994estimation. Assumption (ref) (b) assumes a Bahadur representation of the estimator on the fitting sample, and is satisfied by a broad class of estimators, including commonly used estimators, such as (nonlinear) least squares, maximum likelihood, and the (generalized) method of moments, under standard regularity conditions. Assumption (ref) is a high-level condition on the weight function. Such functions are often called “generically comprehensively revealing” functions; see, for example, bierens1997asymptotic and stinchcombe1998consistent. The boundedness of $w(\cdot)$ is mild, and the required uniform ergodic theorem for $\zeta_t w(\boldsymbol{\xi}_t,\mathbf{x})$ holds under standard uniform integrability conditions; see, e.g., Theorem 3.1 of ling2003asymptotic. Typical examples include $w(\boldsymbol{\xi}_t,\mathbf{x})=\sin(\mathbf{x}'\boldsymbol{\xi}_t)$, $w(\boldsymbol{\xi}_t,\mathbf{x})=[1+\exp(\mathbf{x}'\boldsymbol{\xi}_t)]^{-1}$, $w(\bf{\xi}_t,\mathbf{x} )=\exp( i \mathbf{x}' \boldsymbol{\xi}_t)$, and $w(\boldsymbol{\xi}_t,\mathbf{x})=\mathbb{I}(\boldsymbol{\xi}_t\leq \mathbf{x})$, see Lemma 1 of escanciano2006goodness for further details. In the numerical studies, we focus on the latter two choices. Assumption (ref) allows the observed information set to approximate the infinite past, see similar assumptions in HongLee2005 and WangZhuShao2022.
In addition to the above regularity conditions, we assume that the sample-splitting sequences are arbitrary non-decreasing sequences $\{(f_n, l_n )\}_{n \geq 1}$ which go to infinity with $n$, and $f_n/n \rightarrow \kappa_f$ and $l_n/n \rightarrow \kappa_l$ for some constants $0 < \kappa_f, \kappa_l \leq 1$. Further, let $\kappa_{ra}\equiv \lim_{n \rightarrow \infty} l_n/f_n$ be the limiting ratio of sample-splitting, and $\kappa_{ov}=\lim_{n\rightarrow \infty} \max(0, (f_n+l_n-n)/f_n)$ be the limiting overlap coefficient.
Theorem (ref) states that the split-sample statistic has the oracle null limit, despite using estimated residuals. This is the main theoretical difference from the full-sample generalized spectral test of escanciano2006goodness, where the same observations are used both to estimate the unknown parameter and to construct the residual marked process. In that full-sample construction, the first-order estimation effect is part of the limiting null distribution and must be accounted for in the approximation of critical values. By contrast, the present split-sample construction separates parameter fitting from model checking in a way that makes the residual-based process have the same limiting distribution as the oracle process based on the true errors.
Although the limiting null distribution in Theorem (ref) is still nonstandard and requires approximation by bootstrap, the oracle-equivalence result has an important practical implication. The bootstrap procedure in Section (ref) can apply multipliers directly to the split-sample residuals while keeping $\widehat{\boldsymbol{\theta}}_{f_n}$ fixed. Hence, unlike full-sample residual-based procedures that require re-generating data and re-estimating the model in each bootstrap replication, the proposed method avoids repeated refitting and can be substantially cheaper for nonlinear or numerically intensive models.
We proceed to prove the consistency of the proposed test under the fixed and local alternatives, respectively. First, consider the fixed alternative:
where $\{a_{t}\}$ is strictly stationary and ergodic, with $Ea_{1}^2<\infty$, and importantly, for each $t \in \mathbb{Z}$, $a_{t}$ is $\mathcal{F}_{t-1}$-measurable and $P(a_t=0)<1$. Theorem (ref) shows the asymptotic behavior of $S_{n, w}$ under the global alternative $H_{a}$.
Corollary (ref) indicates that the test is consistent against alternatives that violate at least one of the pairwise conditional mean restrictions, namely $P(E[a_t|\boldsymbol{Z}_{t-j}]=0)<1$ for some $j\geq 1$. As in generalized spectral tests based on pairwise conditional moments, this is a large but not exhaustive class. It is indeed possible that $P(E[a_t| \mathbf I_{t-1}]=0)<1$ while all pairwise projections satisfy $E[a_t | \boldsymbol{Z}_{t-j}]=0$ a.s. for every fixed $j\geq 1$. Such alternatives involve higher-order or genuinely joint dependence on the past and are therefore not captured by a statistic built only from pairwise restrictions.
To further study the consistency properties of the proposed tests, we consider the local alternatives with a sequence of alternative hypotheses tending to the null at the parametric rate $l_n^{-1 / 2}$ as in escanciano2006goodness:
where $\{a_{t}\}$ is the same as in $H_{a}$ (ref). Note that the rate $l_n$ is of the same order as $n$ in the half-splitting scheme.
An additional assumption is needed regarding the behavior of the estimator under these local alternatives.
Theorem (ref) decomposes the local limit into three parts: the oracle Gaussian limit $S_w^0$, the direct local-alternative drift $L_w$, and the indirect drift caused by estimating $\boldsymbol{\theta}_0$ on the fitting sample. If $\boldsymbol{\xi}_{a} \neq \mathbf{0}$, local power may be reduced in directions that are aligned with the score $\mathbf{g}_{t}\left(\boldsymbol{\theta}_{0}\right)$. If $\boldsymbol{\xi}_{a}=\mathbf{0}$, the estimator has no first-order local drift, and the test has nontrivial power against all local alternatives in $\Xi$.
We will show in Section (ref) that, in finite samples, the proposed tests exhibit empirical power comparable to that of the full-sample generalized spectral tests in most linear and nonlinear examples.
The limiting null distribution in Theorem (ref) is nonstandard because the covariance operator of the limiting Gaussian process depends on the underlying data-generating process (DGP) in a very complicated way. Hence, direct tabulation of the critical value is generally infeasible. However, compared with escanciano2006goodness, a key observation is that the asymptotic distribution of $S_{n, w}(\boldsymbol{\eta}, \widehat{\boldsymbol{\theta}}_{f_n})$ is the same as $S_w^0(\boldsymbol{\eta})$, which does not involve any estimation effect of $\widehat{\boldsymbol{\theta}}_{f_n}$. We elaborate more on this point.
In the full-sample generalized spectral test of escanciano2006goodness, because the asymptotic distribution of the test therein involves the estimation effect, the bootstrap must reproduce this effect. Hence, escanciano2006goodness applies the fixed-design wild bootstrap: one first generates bootstrap residuals, then constructs bootstrap responses, then { re-estimates} the model on the bootstrap sample to obtain $\boldsymbol{\theta}_n^*$, and finally recomputes bootstrap residuals using $\boldsymbol{\theta}_n^*$. This is why the procedure requires additional high-level assumptions on the bootstrap estimator.
The split-sample construction in this paper changes the bootstrap problem. By Theorem (ref), under the split-overlap and score-alignment conditions (ref), the residual-based process $S_{n,w}(\cdot,\widehat{\boldsymbol{\theta}}_{f_n})$ has the same limiting null distribution as the oracle process based on the true errors. Hence, the bootstrap only needs to mimic the oracle martingale fluctuation on the checking sample. This means that the fitting-sample estimator $\widehat{\boldsymbol{\theta}}_{f_n}$ can be kept fixed throughout the bootstrap, and no bootstrap responses are generated. This is the main computational advantage of the proposed procedure, especially for nonlinear models or models whose estimation requires cumbersome numerical optimization.
Specifically, let $\{V_t\}$ be a sequence of i.i.d. random variables, independent of the data, with zero mean, unit variance, and bounded support. Conditional on the observed sample, define
where
This is a multiplier version of the checking-sample moment process. It is related to the wild bootstrap of Wu1986bootstrap, liu1988bootstrap, and mammen1993bootstrap, but unlike the full-sample fixed-design wild bootstrap used in escanciano2006goodness, it does not require constructing $Y_t^*$ or re-estimating the parameter.
Common choices of $\{V_t\}$ include Mammen's two-point distribution,
and the Rademacher distribution $P(V_t=1)=P(V_t=-1)=1/2$; see also stute1998bootstrap, de1996bierens, and mammen1993bootstrap.
The complete bootstrap algorithm is summarized in Algorithm (ref).
The algorithm keeps $\widehat{\boldsymbol{\theta}}_{f_n}$ fixed in every bootstrap replication. Therefore, the computational cost of the bootstrap is dominated by recomputing the weighted residual moments, rather than by repeated model estimation. This distinction is minor for very simple linear models, but it can be substantial for nonlinear volatility models, and threshold models, see Section (ref) for numerical evidence.
We then provide the validity of the bootstrap procedure.
To examine the finite-sample performance of the proposed tests, we carry out simulation studies with different DGPs under the nulls and under the alternatives. Here we consider the univariate case, that is, $m=1$ with $Z_{t}=Y_{t}, t \in \mathbb{Z}$. Denote $D_{n, I}^{2}$ and $D_{n, C}^{2}$ to be the proposed CvM tests corresponding to $w(Y_{t}, x)=\mathbbm{1}(Y_{t} \leq x)$ and $w (Y_{t}, x)=\exp (i Y_{t} x)$, where $W(\cdot)$ is chosen to be the empirical CDF $F_{n}(\cdot)$ based on $\{Y_{t-1}\}_{t=n-l_n+1}^{n}$ and the CDF $\Phi(\cdot)$ of a standard normal random variable, respectively. Then,
with $\widehat{\gamma}_{j, I}(x, \widehat{\boldsymbol{\theta}}_{f_n})= (\widehat{\sigma}_{e} n_{j})^{-1} \sum_{k=n-l_n+j}^{n} \widehat{e}_{k} (\widehat{\boldsymbol{\theta}}_{f_n}) \mathbbm{1}(Y_{k-j} \leq x )$, $\widehat{\sigma}_{e}^{2}=l_n^{-1} \sum_{s=n-l_n+1}^{n} \widehat{e}_{s}^{2}(\widehat{\boldsymbol{\theta}}_{f_n})$, and
Other choices of integrating functions $W(\cdot)$ in $D_{n, C}^{2}$ can also be applied. Besides, we also compare the finite-sample performance with that of CvM tests in escanciano2006goodness without sample splitting, which are denoted as $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ using FDWB approximation for the bootstrapped statistics.
For the simulation setup, we consider sample sizes $n=200$ and $n=500$, and the nominal level of $5\%$ for each test. The Monte Carlo experiments are repeated $1,000$ times, and in each replication, the initial $200$ burn-in observations of the generated processes are discarded to ensure stationarity. The number of bootstrap replications is $B=500$, and the random weighting $\{V_{t}\}$ is generated by i.i.d. Mammen's distribution (ref).
First we consider the linear case, where the null model is the AR$(1)$ model as follows:
We examine the goodness-of-fit of this model under the following DGPs:
DGPs 1--4 are introduced to examine the empirical size of the tests under the nulls, and DGPs 5--12 are used to examine the power under the alternatives.
For the computation of all test statistics, we apply the usual least squares residuals under $H_0$, that is, $\widehat{e}_{t} (\widehat{\boldsymbol{\theta}}_{f_n})=Y_{t}-\widehat{b}_{f_n} Y_{t-1}$, for $t=n-l_n+1,\ldots,n$. Table (ref) reports the empirical rejection probabilities (RPs) associated with DGPs 1--4 to examine the empirical size of the tests, and Table (ref) reports the empirical RP associated with DGPs 5-12 to examine the empirical power of the tests in the cases of $n=200$ and $n=500$.
Table (ref) shows that the proposed split-sample generalized spectral testing framework attains empirical sizes close to the nominal 5% level in finite samples under the linear AR(1) null models, and that these sizes are comparable to those of the generalized spectral tests in escanciano2006goodness. This finding supports Theorem (ref), which states that the proposed test has the same null limiting distribution as the oracle process. Table (ref) indicates that the proposed tests exhibit empirical power comparable to that of the full-sample generalized spectral tests, and that the power increases as the sample size grows from $n=200$ to $n=500$, which validates Theorem (ref) for the consistency of the proposed tests.
In the nonlinear case, we are interested in analyzing the size and power against misspecifications in the conditional variance. We consider two common conditional heteroscedasticity models:
1. ARCH-type null model, e.g., ARCH$(1)$ model:
and ARCH$(4)$ model:
2.
GARCH-type null model, e.g., GARCH$(2,2)$ model:
Recall that under the null, $Y_t$ corresponds to $y_t^2$, $E[Y_t | \mathbf{I}_{t-1}] = \sigma_t^2$, and the MDS $\varepsilon_t = Y_t - E[Y_t | \mathbf{I}_{t-1}] = \sigma_t^2 (\eta_t^2 -1)$. We examine the goodness-of-fit of this model under the following DGPs, where for all cases $\eta_{t} \overset{\text{i.i.d.}}{\sim} N(0,1)$:
The DGPs for comparisons are the same as those in Escanciano2008.
To compute the test statistics, we apply the quasi-MLE for estimation of unknown parameters $\boldsymbol{\theta}= (\omega, \phi_1, \ldots, \phi_p, \psi_1, \ldots, \psi_q)' \in \mathbb{R}^{p+q+1}$, and define the residuals to be $\widehat{e}_{t} (\widehat{\boldsymbol{\theta}}_{f_n})=Y_{t}-\sigma_t^2(\widehat{\boldsymbol{\theta}}_{f_n})$, for $t=n-l_n+1,\ldots,n$, where $\widehat{\boldsymbol{\theta}}_{f_n}$ is the quasi-MLE for $\boldsymbol{\theta}_0$ defined in Section (ref) using the first $f_n$ samples.
For the ARCH$(1)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) associated with DGPs 1--2, 5, and 7--11 to examine the empirical size and power of tests at 5% with sample size $n=200$ and $n=500$. It shows that the proposed split-sample tests achieve empirical sizes close to the nominal $5\%$ level, with size performance comparable to that of the original full-sample generalized spectral tests. Their empirical power is also broadly comparable across the nonlinear alternatives considered, and it increases as the sample size rises from $n=200$ to $n=500$.
For the ARCH$(4)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) under the DGPs 3--5 and 7--11. It indicates that the proposed split-sample tests again maintain empirical sizes close to the nominal level and have power comparable to that of the corresponding full-sample tests. Moreover, in several cases, the split-sample procedures outperform the original tests in finite samples, both in terms of size accuracy and empirical power. The power improvement with larger sample sizes is also evident when changing the sample size from $n=200$ to $n=500$.
For the GARCH$(2,2)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) under the DGPs 4 and 6--11. It shows a similar pattern that the proposed split-sample tests deliver empirical sizes close to the nominal $5\%$ level and exhibit empirical power broadly comparable to that of the original full-sample tests. In several scenarios, they even improve upon the original procedures in terms of finite-sample size control or rejection power. In the other settings, the power generally increases with the sample size.
We proceed to study more complex nonlinear cases, for example, the threshold autoregressive (TAR) model. Simulation studies show that our proposed test still has consistently good size and power against model misspecifications.
We consider the following three-regime self-exciting TAR(1,1,1; $d=1$) process as null model, and set the threshold variable to be of lag 1:
We examine the goodness-of-fit of this model under the following DGPs:
To compute the test statistics, for each fixed pair $(r_1,r_2)'$, we first estimate $(\phi_1,\phi_2,\phi_3)'$ by least squares. We then select the optimal threshold pair $(r_1,r_2)'$ over a grid by minimizing the least squares criterion. The residuals $\widehat e_t(\widehat{\boldsymbol\theta}_{f_n})$ are then obtained by plugging in the estimator $\widehat{\boldsymbol{\theta}}_{f_n} = (\widehat\phi_1,\widehat\phi_2,\widehat\phi_3,\widehat r_1,\widehat r_2)'$ using the first $f_n$ samples, for $t=n-l_n+1,\ldots,n$.
Table (ref) reports the empirical sizes and powers (%) of the tests at the nominal $5\%$ level for the null hypothesis $H_0$ of the three-regime TAR$(1,1,1;d=1)$ model in (ref), under DGPs 1--8 and sample sizes $n=200$. We omit the case $n=500$ for now because the full-sample generalized spectral tests are quite time-consuming for the three-regime TAR model, and $n=200$ is sufficient to illustrate the validity of the proposed tests. The results show that the proposed split-sample tests again maintain empirical sizes close to the nominal level and exhibit power comparable to that of the corresponding full-sample tests.
We now collect the computational time comparisons across the simulation designs. The reported times are averaged over five replications for one Monte Carlo experiment. Table (ref) reports the corresponding results for the null models in Sections (ref)-(ref), namely the AR(1), ARCH(1), ARCH(4), GARCH(2,2) and three-regime TAR(1,1,1) models.
The timing results highlight where sample splitting is most useful computationally. In the linear AR case, the proposed and full-sample generalized spectral tests have similar computational costs, because least squares estimation is inexpensive even when repeated inside the bootstrap. A similar pattern appears for the ARCH$(1)$ null, where repeated estimation is still relatively light. The advantage becomes much clearer for higher-order volatility models. Under the ARCH$(4)$ and GARCH$(2,2)$ nulls, the full-sample procedures require repeated bootstrap re-estimation of a larger nonlinear model, whereas the proposed tests estimate the model only once and then apply multipliers directly to the split-sample residuals.
The computational gain is most pronounced for the three-regime TAR model. There, each full-sample bootstrap replication requires another threshold search and model refit, which is costly even for moderate sample sizes. By contrast, the split-sample procedure keeps the fitted parameter fixed throughout the bootstrap. As a result, the proposed tests can be much faster while maintaining accurate size and comparable power performance.
The dynamic behavior of stock returns has been a central topic in financial econometrics, since the specification of the conditional mean and conditional variance directly affects risk measurement, portfolio allocation, and asset pricing. In particular, it is important to distinguish whether the serial dependence in financial returns is mainly driven by the conditional mean, the conditional variance, or both.
We use the same dataset as in the empirical study of Escanciano2008, namely the log differences of the S&P 500 daily stock index in a sample period from January 1, 1988 to May 28, 1993. Following the literature, we delete the last 10% of the observations, leaving 1138 observations. We examine the following three model specifications:
The bootstrap $p$-values are reported in Table (ref). The first four columns $D_{n, I}^{2}$ and $D_{n, C}^{2}$ give the results of the proposed sample-splitting tests in (ref), while the last four columns $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ report the corresponding full-sample tests in escanciano2006goodness, Escanciano2008. For each fitted model, we consider both marginal and joint tests for the conditional mean and conditional variance, denoted by the subscripts $m$ and $v$, respectively. The number of bootstrap replications is $B=500$.
Table (ref) shows that the sample-splitting tests yield conclusions very similar to those from the full-sample tests. For the AR(1) model with conditional homoskedastic errors, the $p$-values are generally large for both the conditional mean and conditional variance tests. This indicates that the linear AR(1) specification captures the conditional mean adequately, and there is no strong evidence against conditional homoskedasticity.
When the AR(1) component of the mean is neglected and only a GARCH(1,1) model is fitted, the conditional mean tests strongly reject the null hypothesis. This suggests that ignoring the linear dependence in the conditional mean leads to misspecification. In contrast, for the AR(1)-GARCH(1,1) model, the conditional mean appears to be reasonably specified, but the conditional variance tests produce zero $p$-values. Hence, the GARCH specification for the conditional variance is rejected. These findings are consistent with the conclusions in the existing study of hong2003diagnostic and Escanciano2008: the S&P 500 returns over this period are better described by a linear AR(1) mean structure with conditional homoskedastic errors, rather than by an AR(1)-GARCH(1,1) model.
Overall, this empirical application confirms the usefulness of the sample-splitting tests. Although the sample-splitting procedure uses only part of the sample for estimation and the remaining part for testing, it delivers conclusions comparable to those from the full-sample tests.
To further illustrate the robustness of the proposed test and its computational advantage, we consider a classical dataset in the TAR literature, that is, the annual Wolf's sunspot data from 1700 to 1979. This dataset has been studied in tong1978threshold, tsay1989testing and the references therein. The series consists of 280 observations and is well known for its asymmetric cyclical behavior. Let $\{Y_t\}_{t=1}^{280}$ denote the annual sunspot series.
Following the model specification in tsay1989testing, we fit $\{Y_t\}_{t=1}^{280}$ using the following three-regime TAR model:
where $Y_{t-2}$ is selected as the threshold variable, and the AR orders are 11, 10, and 10, respectively. The unknown autoregressive coefficients are estimated by least squares for each fixed pair of threshold values, and the thresholds are selected by minimizing the least squares criterion. In tsay1989testing, the martingale difference assumption that $E(\varepsilon_t|\mathbf{I}_{t-1}) = 0$ is imposed but not formally tested. It motivates us to apply the generalized spectral test to examine the MDS assumption.
Table (ref) reports the bootstrap $p$-values for different models. The first four columns $D_{n, I}^{2}$ and $D_{n, C}^{2}$ give the results of the proposed sample-splitting tests in (ref), while the last four columns $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ report the corresponding full-sample tests of escanciano2006goodness, Escanciano2008. The number of bootstrap replications is $B=500$.
The results show strong evidence in favor of the three-regime TAR(11,10,10) specification for the annual sunspot data. The large bootstrap $p$-values indicate that the residuals from the fitted TAR model do not exhibit significant departures from the martingale difference property. Moreover, the sample-splitting tests lead to conclusions very similar to those obtained from the full-sample tests. Meanwhile, the computational gain is substantial. The proposed sample-splitting tests are more than 20 times faster than the full-sample tests. This gain comes from avoiding repeated model re-estimation and threshold search in the bootstrap procedure, which is computationally expensive for TAR models with multiple regimes and threshold parameters.
For comparison, we also fit the dataset $\{Y_t\}_{t=1}^{280}$ using a constant mean model:
or a linear AR(p) model:
with $p=5,10$, and 11. The testing results are summarized in Table (ref). From this table, we can find that the constant model and AR(5) model are strongly rejected by all tests. Meanwhile, the AR(10) model is rejected by most tests at the 10% significance level, and the bootstrap $p$-values for the AR(11) model remain relatively small compared with those for the three-regime TAR model. Overall, the proposed sample-splitting tests produce conclusions consistent with the full-sample tests and suggest that the constant and linear AR specifications are inadequate for the annual sunspot series.
This paper develops a sample-splitting generalized spectral test for diagnostic checking of parametric time-series models. The proposed procedure combines the omnibus, bandwidth-free structure of generalized spectral tests with a split-sample estimation strategy that removes the first-order effect of parameter estimation under suitable split-overlap and score-alignment conditions. The resulting statistic targets pairwise conditional mean dependence in residuals and therefore provides a diagnostic checking for violations of the martingale-difference implication of a correctly specified conditional mean model.
A practical advantage of the method is that critical values can be obtained from a multiplier bootstrap applied directly to the split-sample residuals. The bootstrap avoids generating artificial time series and avoids re-estimating the model in each replication, which can lead to substantial computational savings for nonlinear or numerically intensive models. The simulation results suggest that the test has a reliable size and competitive power across linear, heteroskedastic, and nonlinear dynamic specifications.
Several extensions are worth pursuing. First, the same sample-splitting principle may be useful for broader conditional moment restriction tests, including recent Gaussian-process-based model checks Escanciano2024. Second, the choice of weight function and integrating measure may affect finite-sample power, and a systematic comparison of indicator, characteristic-function, and data-adaptive weights would be valuable. Third, while the present paper focuses on pairwise conditional mean restrictions, extensions toward diagnostics for conditional quantile restrictions, higher-order or joint conditional mean and variance dependence remain important directions for future research.