EconBase
← Back to paper

Generalized Spectral Testing with Sample Splitting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,498 characters · 15 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalized Spectral Testing with Sample Splitting

abstractResidual-based goodness-of-fit tests for parametric time-series models are often complicated by parameter-estimation effects, which can alter the limiting behavior of diagnostic statistics. We propose a sample-splitting generalized spectral test (in the spirit of escanciano2006goodness) for assessing conditional mean specification in linear and nonlinear time-series models. The procedure estimates the model parameter on a fitting subsample and constructs a generalized spectral Cramér–von Mises statistic from residuals computed on a checking/testing subsample. The statistic aggregates pairwise conditional mean restrictions over all lags and is therefore bandwidth-free and free of truncation-lag selection. Under mild regularity conditions and a score-alignment condition, the residual-based process has the same limiting null distribution as the infeasible oracle process based on the true errors. Although the resulting limiting law is still non-pivotal, it can be consistently approximated by a simple multiplier bootstrap that does not require generating bootstrap time series or re-estimating parameters. Such an oracle-equivalence property is in sharp contrast to the original full-sample test, for which parameter estimation contributes an additional first-order term to the limiting process, and requires re-estimating parameters in each bootstrapped sample. We further establish consistency of the proposed test against fixed alternatives and nontrivial power against local alternatives. Extensive simulations and real data analyses show that the proposed test controls size well, has comparable power, and delivers substantial computational savings in models where repeated estimation is costly.

{\it Keywords:} Goodness-of-fit; Sample-splitting; Generalized spectral test; Bootstrap.

Introduction

Diagnostic checking is a central step in time series analysis. Once a parametric dynamic model has been fitted, any subsequent inference, forecasting exercise, or structural interpretation rests on the maintained assumption that the conditional mean has been adequately specified. If the fitted model leaves systematic predictability in the residuals, parameter estimates may still be numerically stable, but the resulting inferential and forecasting conclusions can be misleading. This motivates formal goodness-of-fit tests that examine whether the residuals satisfy the martingale difference property implied by a correctly specified conditional mean.

We consider a strictly stationary and ergodic time series $\{(Y_{t}, \mathbf{Z}_{t-1}^{\prime})\}_{t \in \mathbb{Z}}$ defined on the probability space $(\Omega, \mathcal{F}, P)$. Here $Y_t\in\mathbb R$ denotes the response variable, and $\mathbf Z_{t-1}=(Y_{t-1},\mathbf X_{t-1}')'\in\mathbb R^m$ collects the lagged response and other predetermined explanatory covariates. Let $\mathbf I_{t-1}=(\mathbf Z_{t-1}',\mathbf Z_{t-2}',\ldots)'$ denote the full information set available at time $t-1$, and let $\mathcal F_{t-1}=\sigma(\mathbf I_{t-1})=\sigma(\mathbf Z_s:s\le t-1)$ be the $\sigma$-field generated by $\mathbf{I}_{t-1}$. Under the integrability of $Y_{t}$, the model can be written as

align*[align* omitted — 63 chars of source]

where $m(\mathbf{I}_{t-1})=E[Y_{t} \mid \mathbf{I}_{t-1}]$ is the conditional mean almost surely (a.s.) given the conditioning set $\mathbf{I}_{t-1}$, and $\varepsilon_{t}=Y_{t}-E[Y_{t} \mid \mathbf{I}_{t-1}]$ is a martingale difference sequence (MDS) with respect to $\mathbf{I}_{t-1}$ as the error term.

In this paper, we ask whether there exists a parametric family of functions $\mathcal{M}=\{f(\cdot, \boldsymbol{\theta}): \boldsymbol{\theta} \in \Theta \subset \mathbb{R}^{p} \}$, such that

align[align omitted — 162 chars of source]

or, equivalently, \[ H_0:\quad E\{e_t(\boldsymbol{\theta}_0)\mid \mathbf{I}_{t-1}\}=0\ \text{a.s.}, \] where $e_{t}(\boldsymbol{\theta})=Y_{t}-f(\mathbf{I}_{t-1},\boldsymbol{\theta})$ denotes the parameter induced model error.

Thus, goodness-of-fit can be formulated as a martingale difference hypothesis for the model errors. This formulation is especially natural in dynamic regression settings, where the relevant null is not merely the absence of linear autocorrelation, but the absence of any remaining conditional mean predictability.

Classical diagnostic tools are often based on residual autocorrelations and portmanteau statistics. These procedures remain attractive because they are simple, familiar, and easy to implement; see, among others, BoxPierce1970, Hosking1980, and LiMcLeod1981. Their scope, however, is limited. Portmanteau-type checks are primarily geared toward linear serial dependence, and their standard reference distributions are often justified under innovation assumptions that are stronger than the null one would actually like to test. In particular, when the residuals are uncorrelated but still dependent, as may happen under conditional heteroskedasticity, the usual chi-squared calibration can fail; see Romano1996. This mismatch is important in many nonlinear and heteroskedastic time series models, where the martingale difference null is the appropriate benchmark.

These concerns have led to a rich literature on omnibus specification testing for time series models. Hong1999 introduced generalized spectral methods for detecting serial dependence through empirical characteristic functions, and HongLee2005 extended this line of work to conditional mean models with conditional heteroskedasticity of unknown form. Particularly relevant for the present paper is the bandwidth-free generalized spectral test of escanciano2006goodness, which provides a Cram\'er-von Mises type diagnostic for linear and nonlinear conditional mean models by integrating a continuum of pairwise moment restrictions. Subsequent work has broadened the scope of these ideas to joint and marginal checks for conditional mean and variance models Escanciano2008, as well as to specification testing for conditional distribution models ChenHong2014. More recently, Escanciano2024 has recast model checking in a broader conditional moment restriction framework. In a related direction, WangZhuShao2022 developed MDDM (Martingale Difference Divergence Matrix) based tests for the martingale difference hypothesis in multivariate time series models, and also provided a detailed discussion about the similarity and distinction between the MDDM-based approach and the one by escanciano2006goodness.

A recurring technical difficulty in this literature is the estimation effect. Because the residuals used for diagnostic checking are computed from the same data that are used to estimate the unknown parameter, residual-based empirical processes do not, in general, behave as if the true innovations were observed. For many non-smoothing-based tests, this first-order estimation effect is asymptotically non-negligible, which leads to nonpivotal null distributions and complicates implementation. Existing remedies typically rely either on bootstrap procedures that repeatedly generate pseudo-data and refit the model, or on martingale-transform arguments designed to remove nuisance-parameter effects; see, for example, KoulStute1999, KhmaladzeKoul2004, and Bai2003. While these approaches are theoretically powerful, they may be computationally demanding and can make routine model checking less convenient in practice.

Recent work by davis2025sample shows that sample splitting offers an appealing alternative in time series goodness-of-fit problems. They generalize the half-sample splitting device proposed by durbin1976kolmogorov from an independent data setting to the time series setting. The idea is to estimate the model parameters on one subsample and to assess goodness-of-fit on another, while allowing for a carefully chosen asymptotic overlap between the two parts of the sample. For diagnostics based on the autocorrelation function and the auto-distance correlation function, they show that the resulting residual-based statistics can have the same asymptotic behavior as in the infeasible oracle case, where the innovations are observed if the split-ratio between the fitting sample length and testing sample length is 1/2. This insight is important, but their framework is tailored to tests of residual serial independence and, in the portmanteau setting, it works under an independently and identically distributed (i.i.d.) innovation paradigm that is stronger than the martingale difference null. From the standpoint of conditional mean adequacy, it is therefore natural to ask whether the same sample-splitting idea can be combined with a diagnostic whose null hypothesis is aligned with the martingale difference property itself.

The goal of this paper is to answer that question. We develop a sample-splitting-based generalized spectral test for the goodness-of-fit of parametric conditional mean models. More specifically, we combine the integrated generalized spectral/Cram\'er-von Mises framework of escanciano2006goodness with the sample-splitting strategy of davis2025sample. One part of the sample is used for estimation, while a checking/testing subsample is used to form residuals based on the fitted parameter and then to evaluate pairwise conditional mean restrictions of the form \[ E\{e_t(\boldsymbol{\theta}_0)w(\mathbf Z_{t-j},\mathbf x)\}=0, \qquad j\geq 1, \] using appropriately chosen weighting functions $w(\mathbf Z_{t-j},\mathbf x)$ indexed by lagged explanatory variables $\mathbf{Z}_{t-j}$ and a continuum of values $\mathbf{x}$. In this way, the proposed procedure preserves the omnibus and bandwidth-free nature of generalized spectral diagnostics, while using sample splitting to separate model fitting from model assessment. As in escanciano2006goodness, the resulting test targets pairwise conditional mean dependence rather than full joint conditional mean dependence on the entire past.

The paper makes several contributions. First, it develops a general split-sample generalized spectral testing framework for linear and nonlinear time series models with parametric conditional means. Second, under suitable regularity conditions and an overlap condition linking the fitting and checking subsamples, we show that the split-sample generalized spectral process has the same asymptotic null distribution as the corresponding infeasible oracle process based on the true errors. Thus, sample splitting eliminates the first-order parameter-estimation effect from the limiting null law. Third, although the limiting null distribution remains nonpivotal, it can be consistently approximated by a simple multiplier bootstrap applied directly to the split-sample residuals. This avoids generating bootstrap time series and re-estimating the model in each bootstrap replication, and therefore leads to a markedly lighter computational procedure. This is a major difference from escanciano2006goodness, and the computational gain can be substantial. Similar to the latter paper, our procedure is also tuning parameter-free. Fourth, we establish consistency against fixed alternatives, derive nontrivial power against local alternatives, and verify the required conditions for important model classes, including ARMA and GARCH models.

The proposed methodology complements several existing strands of literature. Relative to davis2025sample, our focus is not residual serial independence per se, but the martingale difference hypothesis implied by a correctly specified conditional mean. Relative to escanciano2006goodness, the contribution is not a new generalized spectral metric, but rather a new inferential device that uses sample splitting to simplify the effect of parameter estimation and to facilitate bootstrap implementation. In this sense, the paper should be viewed as complementary to both earlier MDS testing work and the recent sample-splitting literature.

The rest of the paper is organized as follows. Section (ref) proposes the split-sample generalized spectral process and the associated Cram\'er-von Mises statistic. Section (ref) develops the asymptotic theory under the null hypothesis, fixed alternatives, and local alternatives. Section (ref) presents the multiplier bootstrap approximation algorithm and its theoretical justification. Section (ref) reports extensive simulation studies on the finite-sample performance of the proposed tests, including linear ARMA models, nonlinear volatility models, and threshold autoregressive models. Section (ref) presents empirical illustrations based on S&P 500 dynamics and annual Wolf's sunspot data, demonstrating the validity and computational advantage of the proposed tests. Section (ref) concludes.

Generalized Spectral Tests with Sample Splitting

This section constructs the sample-splitting generalized spectral test for (ref). The population restrictions and their integrated spectral aggregation follow the generalized spectral approach of escanciano2006goodness: the null is represented by a continuum of pairwise conditional-mean restrictions, and the test measures whether the corresponding integrated spectral distribution is flat across the frequency index. The implementation, however, is different in an important way. In escanciano2006goodness, the parameter is estimated from the full sample, and the same full sample is then used to construct the residual-based test statistic. Here, the parameter is estimated on a fitting subsample, whereas the generalized spectral process is formed from residuals on a checking subsample.

We first introduce the generalized spectral test in escanciano2006goodness. Under $H_0$, the model error satisfies $e_t(\boldsymbol{\theta}_0)=\varepsilon_t$ and hence has zero conditional mean given the full past. In particular, for every lag $j\ge 1$,

align[align omitted — 250 chars of source]

Thus, a correctly specified conditional mean implies an infinite collection of pairwise restrictions, one for each lagged vector $\mathbf Z_{t-j}$. To make these restrictions operational, let $\{w(\mathbf Z_{t-j},\mathbf x):\mathbf x\in\Upsilon\subset\mathbb R^s\}$ be a family of bounded weighting functions such that the conditional restriction in (ref) is equivalently characterized by

align[align omitted — 226 chars of source]

As discussed later, usual choices of weight functions $w$ include $w (\mathbf{Z}_{t-j}, \mathbf{x})=\mathbbm{1} (\mathbf{Z}_{t-j} \leq \mathbf{x} )$ with $\mathbf{x} \in[-\infty, \infty]^{m}$, where $\mathbbm{1}(A)$ denotes the indicator function of the event $A$, and $w (\mathbf{Z}_{t-j}, \mathbf{x} )=\exp (i \mathbf{x}^{\prime} \mathbf{Z}_{t-j} )$ with $\mathbf{x} \in \mathbb{R}^{m}$, where $i=\sqrt{-1}$ is the imaginary unit.

The sequence $\{\gamma_{j,w}(\cdot,\boldsymbol{\theta}_0):j\ge1\}$ measures departures from the pairwise martingale-difference implications across all lags. To aggregate these restrictions, define $\gamma_{-j,w}(\cdot,\boldsymbol{\theta}_0)=\gamma_{j,w}(\cdot,\boldsymbol{\theta}_0)$ for $j\ge1$ and consider the generalized spectral transformation

align[align omitted — 234 chars of source]

Therefore, the tests are based on an integrated Fourier transform of the measures $ \{\gamma_{j, w} (\cdot, \boldsymbol{\theta}_{0} ) \}_{j=1}^{\infty}$. Under $H_0$, all nonzero-lag coefficients vanish, so $$f_w(u,\mathbf x,\boldsymbol{\theta}_0) \equiv f_{0, w} (\mathbf{x}, \boldsymbol{\theta}_{0} )=(2 \pi)^{-1} \gamma_{0, w} (\mathbf{x}, \boldsymbol{\theta}_{0} ).$$ Equivalently, the integrated generalized spectral distribution

align[align omitted — 367 chars of source]

reduces to $\gamma_{0,w}(\mathbf x,\boldsymbol{\theta}_0)\lambda$ under $H_0$. Therefore, any departure from this constant function implies a violation of at least one of the pairwise restrictions.

We now describe the proposed sample-splitting implementation. Suppose the observed sample is $\{Y_t,\widehat{\mathbf I}_{t-1}\}_{t=1}^n$, where $\widehat{\mathbf I}_{t-1}=(\mathbf Z_{t-1}',\mathbf Z_{t-2}',\ldots,\mathbf Z_0')'$ is the observed information set and may contain initial values. Let $f_n$ and $l_n$ be two non-decreasing sequences diverging with $n$. The first $f_n$ observations are used for estimation, and the last $l_n$ observations are used for model checking:

itemize• the first $f_n$ data points, $\{Y_{t}, \widehat{\mathbf{I}}_{t-1} \}_{t=1}^{f_n}$, are the fitting sample to estimate the parameter $\boldsymbol{\theta}_0$ of the model; • the last $l_n$ data points, $\{Y_{t}, \widehat{\mathbf{I}}_{t-1} \}_{t=n-l_n+1}^{n}$, are the checking sample used to calculate the residuals based on the estimated parameter.

Let $\widehat{\boldsymbol{\theta}}_{f_n}$ denote the estimator computed from the fitting sample. On the checking sample, define the split-sample residuals \[ \widehat e_t\equiv \widehat e_t(\widehat{\boldsymbol{\theta}}_{f_n}) =Y_t-f(\widehat{\mathbf I}_{t-1},\widehat{\boldsymbol{\theta}}_{f_n}), \qquad t=n-l_n+1,\ldots,n. \] Note that, for lag $j$, the effective sample size is $n_j=l_n-j+1.$ The empirical lag-$j$ weighted moment is

align*[align* omitted — 226 chars of source]

Hence, the sample analogue of the generalized spectral distribution function (ref) is given by

align*[align* omitted — 344 chars of source]

where $ (n_{j} / l_n )^{1 / 2}$ is a finite-sample correction factor that does not affect the asymptotic theory and delivers a better finite-sample performance of the test procedure.

Under $H_0$, the target function is $H_w(\lambda,\mathbf x,\boldsymbol{\theta}_0) =\gamma_{0,w}(\mathbf x,\boldsymbol{\theta}_0)\lambda$ and thus tests can be based on the discrepancy between $\widehat{H}_{w} (\lambda, \mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} )$ and $\widehat{H}_{0, w} (\lambda, \mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} ) \equiv \widehat{\gamma}_{0, w} (\mathbf{x}, \widehat{\boldsymbol{\theta}}_{f_n} ) \lambda$. Define the centered process

align*[align* omitted — 400 chars of source]

Finally, we aggregate the process over $\Pi=[0,1]\times\Upsilon$ using a Cram\'er-von Mises norm:

align[align omitted — 321 chars of source]

where $W$ is an integrating distribution for the index $\mathbf x$. The test rejects $H_0$ for large values of $D_{n,w}^2(\widehat{\boldsymbol{\theta}}_{f_n})$. Because all lags in the checking sample are included with deterministic spectral weights, the procedure does not require choosing a truncation lag or bandwidth.

Asymptotic Theory

This section studies the large-sample behavior of the split-sample generalized spectral process introduced in Section (ref). The main message is that, under a simple score-alignment condition (ref) for the data generating process and a half split-ratio condition on the length of the fitting and checking samples, the residual-based process has the same null limit as the infeasible oracle process that would be computed from the true innovations. We then show that the corresponding Cram\'er-von Mises statistic is consistent against fixed pairwise alternatives and has nontrivial power against local alternatives.

To begin with, we introduce the notation used throughout the asymptotic theory. Let \[ \mathbf{g}_{t}(\boldsymbol{\theta}) \equiv \mathbf{g}(\mathbf{I}_{t-1},\boldsymbol{\theta}) = \frac{\partial}{\partial\boldsymbol{\theta}} f(\mathbf{I}_{t-1},\boldsymbol{\theta}), \qquad w_{t-j}(\mathbf{x})\equiv w(\mathbf Z_{t-j},\mathbf x). \] Recall that \[ \varepsilon_t=Y_t-E(Y_t\mid\mathbf I_{t-1}), \qquad e_t(\boldsymbol{\theta})=Y_t-f(\mathbf I_{t-1},\boldsymbol{\theta}), \] so that under $H_0$, $e_t(\boldsymbol{\theta}_0)=\varepsilon_t$ a.s. Put $\Pi=[0,1]\times\Upsilon$ and write $\boldsymbol{\eta}=(\lambda,\mathbf x')'\in\Pi$.

The process $S_{n,w}(\boldsymbol{\eta},\widehat{\boldsymbol{\theta}}_{f_n})$ is viewed as a random element of the Hilbert space $L_2(\Pi,\nu)$, where $\nu$ is the product of the integrating measure $W$ on $\Upsilon$ and Lebesgue measure on $[0,1]$. For complex-valued functions $f,g \in L_2(\Pi,\nu)$, the inner product is \[ \langle f,g\rangle = \int_{\Pi}f(\boldsymbol{\eta})g^c(\boldsymbol{\eta})\,d\nu(\boldsymbol{\eta}) = \int_{\Pi}f(\lambda,\mathbf x)g^c(\lambda,\mathbf x)\,W(d\mathbf x)d\lambda, \] and $\|f\|=\langle f,f\rangle^{1/2}$. We equip $L_2(\Pi,\nu)$ with the Borel $\sigma$-field generated by this norm and write $\Longrightarrow$ for weak convergence in this space. For a mean-zero $L_2(\Pi,\nu)$-valued random element $Z$ with $E\|Z\|^2<\infty$, its covariance operator is $C_Z(h)=E[\langle Z,h\rangle Z]$, $h\in L_2(\Pi,\nu)$.

Let \[ \Psi_j(\lambda)=\frac{\sqrt{2}\sin(j\pi\lambda)}{j\pi}, \qquad \mathbf b_j(\mathbf x,\boldsymbol{\theta}_0) =E[w_{t-j}(\mathbf x)\mathbf g_t(\boldsymbol{\theta}_0)], \] and define the score-loading function \[ \mathbf G_w(\boldsymbol{\eta}) \equiv \mathbf G_w(\boldsymbol{\eta},\boldsymbol{\theta}_0) = \sum_{j=1}^{\infty} \mathbf b_j(\mathbf x,\boldsymbol{\theta}_0)\Psi_j(\lambda). \] The covariance of the oracle generalized spectral limit is described through the quadratic form

align[align omitted — 339 chars of source]

where $h \in L_{2}(\Pi, \nu)$, $\boldsymbol{\eta}_{1}=(\lambda, \mathbf{x}^{\prime})^{\prime}$, and $\boldsymbol{\eta}_{2}= (\varpi, \mathbf{y}^{\prime})^{\prime}$.

Asymptotic Null Distribution

The following assumptions collect the necessary regularity conditions.

assumption(a). $\{Y_{t}, \mathbf{Z}_{t-1}\}_{t \in \mathbb{Z}}$ is a strictly stationary and ergodic process.\\ (b). $E[\varepsilon_{1}^{2}]<\infty$.
assumptionThe response function $f(\mathbf{I}_{t-1}, \cdot)$ is twice continuously differentiable on $\Theta$. The score $\mathbf{g}_{t}(\boldsymbol{\theta})$ is stationary, ergodic, and $\mathcal{F}_{t-1}$-measurable. There exists an integrable function $\mathbf{M}_t\in \mathcal{F}_{t-1}$ with $|\mathbf{g}(\mathbf{I}_{t-1}, \boldsymbol{\theta})| \leq \mathbf{M}_t$ for all $\theta\in\Theta$, and that $E\mathbf{M}_t^2<\infty.$
assumption(a) The parameter space $\Theta$ is compact in $\mathbb{R}^{p}$, and the true parameter $\boldsymbol{\theta}_{0}$ belongs to the interior of $\Theta$. There exists a unique pseudo-true value $\boldsymbol{\theta}_{*} \in \Theta$ such that the estimator is consistent for $\boldsymbol{\theta}_*$ under both $H_0$ and fixed alternatives.\\ (b) Under $H_{0}$, $\boldsymbol{\theta}_*=\boldsymbol{\theta}_0$, for estimators based on $\{Y_t,\mathbf{Z}_{t-1}\}_{t=1}^{f_n}$, as $f_n \to \infty $, the estimator ${\widehat{\boldsymbol{\theta}}}_{f_n}$ satisfies the expansion \begin{align*} \sqrt{f_n}(\widehat{\boldsymbol{\theta}}_{f_n}-\boldsymbol{\theta}_{0})=\frac{1}{\sqrt{f_n}} \sum_{t=1}^{f_n} \mathbf{h}(Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0})+o_{p}(1), \end{align*} where $\mathbf{h}(\cdot)$ satisfies $E[\mathbf{h}(Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0}) | \mathcal{F}_{t-1}]=\mathbf{0}$, $\mathbf{L}(\boldsymbol{\theta}_{0})= E[\mathbf{h}(Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0}) \mathbf{h}^{\prime}(Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0})]$ exists and is positive definite. By the martingale central limit theorem, the above implies \begin{align*} \sqrt{f_n}(\widehat{\boldsymbol{\theta}}_{f_n}-\boldsymbol{\theta}_{0}) \stackrel{d}{\longrightarrow} \mathcal{N}(\mathbf{0}, \mathbf{L}(\boldsymbol{\theta}_{0})). \end{align*}
assumptionThe integrating function $W(\cdot)$ is a probability distribution function absolutely continuous with respect to Lebesgue measure. The weight function $w(\cdot)$ is such that the equivalence between (ref) and (ref) holds and is uniformly bounded. Also, $w(\cdot)$ satisfies the uniform law of large numbers, \begin{align*} \sup _{\mathbf{x} \in \Upsilon_{c}} \bigg|n^{-1} \sum_{t=1}^{n} \zeta_{t} w(\boldsymbol{\xi}_{t}, \mathbf{x})-E \big[\zeta_{t} w(\boldsymbol{\xi}_{t}, \mathbf{x})\big]\bigg| \rightarrow 0, \quad a.s., \end{align*} whenever $\{(\zeta_{t}, \boldsymbol{\xi}_{t}^{\prime}), t=0, \pm 1, \ldots\}$ is a strictly stationary and ergodic process with $\zeta_{t} \in \mathbb{R}, \boldsymbol{\xi}_{t} \in \mathbb{R}^{m}, E|\zeta_{1}|<\infty$, and $\Upsilon_{c}$ is any compact subset of $\Upsilon \subset \mathbb{R}^{s}$.
assumptionThe observed information set at period $t$, $\widehat{\mathbf{I}}_{t}$, may contain some assumed initial values and satisfies $\lim _{n \rightarrow \infty} \sum_{t=1}^{n}(E \sup _{\boldsymbol{\theta} \in \Theta}(f(\mathbf{I}_{t-1}, \boldsymbol{\theta})-f(\widehat{\mathbf{I}}_{t-1}, \boldsymbol{\theta}))^{2})^{1 / 2} <\infty$.

Assumptions (ref)--(ref) are inherited from escanciano2006goodness, and are deliberately stated at a high level. Assumption (ref)-(ref) are standard but mild conditions on the data-generating process. Assumption (ref) (a) assumes a pseudo-true value under the alternative, which is standard in the theory of misspecified likelihood and other extremum estimation; see, e.g. newey1994large, and white1982maximum,white1994estimation. Assumption (ref) (b) assumes a Bahadur representation of the estimator on the fitting sample, and is satisfied by a broad class of estimators, including commonly used estimators, such as (nonlinear) least squares, maximum likelihood, and the (generalized) method of moments, under standard regularity conditions. Assumption (ref) is a high-level condition on the weight function. Such functions are often called “generically comprehensively revealing” functions; see, for example, bierens1997asymptotic and stinchcombe1998consistent. The boundedness of $w(\cdot)$ is mild, and the required uniform ergodic theorem for $\zeta_t w(\boldsymbol{\xi}_t,\mathbf{x})$ holds under standard uniform integrability conditions; see, e.g., Theorem 3.1 of ling2003asymptotic. Typical examples include $w(\boldsymbol{\xi}_t,\mathbf{x})=\sin(\mathbf{x}'\boldsymbol{\xi}_t)$, $w(\boldsymbol{\xi}_t,\mathbf{x})=[1+\exp(\mathbf{x}'\boldsymbol{\xi}_t)]^{-1}$, $w(\bf{\xi}_t,\mathbf{x} )=\exp( i \mathbf{x}' \boldsymbol{\xi}_t)$, and $w(\boldsymbol{\xi}_t,\mathbf{x})=\mathbb{I}(\boldsymbol{\xi}_t\leq \mathbf{x})$, see Lemma 1 of escanciano2006goodness for further details. In the numerical studies, we focus on the latter two choices. Assumption (ref) allows the observed information set to approximate the infinite past, see similar assumptions in HongLee2005 and WangZhuShao2022.

In addition to the above regularity conditions, we assume that the sample-splitting sequences are arbitrary non-decreasing sequences $\{(f_n, l_n )\}_{n \geq 1}$ which go to infinity with $n$, and $f_n/n \rightarrow \kappa_f$ and $l_n/n \rightarrow \kappa_l$ for some constants $0 < \kappa_f, \kappa_l \leq 1$. Further, let $\kappa_{ra}\equiv \lim_{n \rightarrow \infty} l_n/f_n$ be the limiting ratio of sample-splitting, and $\kappa_{ov}=\lim_{n\rightarrow \infty} \max(0, (f_n+l_n-n)/f_n)$ be the limiting overlap coefficient.

theoremUnder Assumptions (ref)--(ref) and $H_{0}$, if $\kappa_{ra} = 2 \kappa_{ov}$, and the following score-alignment condition holds, that is, \begin{equation} E\big[w_{t-j}(\mathbf{x}) \mathbf{g}_{t}(\boldsymbol{\theta}_0)\big]^{\prime} \mathbf{L}(\boldsymbol{\theta}_{0}) = E \Big[ \varepsilon_t w_{t-j}(\mathbf{x}) \mathbf{h}(Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0}) \Big]^{\prime}, \quad \forall j \geq 1, \end{equation} then the process $S_{n, w}$ satisfies \begin{equation*} S_{n, w} \Longrightarrow S_w^0, \quad in L_{2}(\Pi, v), \end{equation*} where $S_{w}^{0}(\cdot)$ is a Gaussian process in $L_{2}(\Pi, v)$ with mean $\mathbf{0}$ and covariance operator $C_{S_{w}^{0}}$ satisfying $\sigma_{h}^{2}=\langle C_{S_{w}^{0}}(h), h \rangle, \forall h \in L_{2}(\Pi, v)$, and $\sigma_{h}^{2}$ is defined in (ref). Moreover, by the continuous mapping theorem, we have \begin{align*} D_{n, w}^{2}(\widehat{\boldsymbol{\theta}}_{f_n}) \stackrel{d}{\longrightarrow} D_{\infty, w}^{2}(\boldsymbol{\theta}_{0}) :=\int_{\Pi} |S_{w}^0 (\lambda, \mathbf{x}, \boldsymbol{\theta}_{0})|^{2} W(d \mathbf{x}) d \lambda. \end{align*}

Theorem (ref) states that the split-sample statistic has the oracle null limit, despite using estimated residuals. This is the main theoretical difference from the full-sample generalized spectral test of escanciano2006goodness, where the same observations are used both to estimate the unknown parameter and to construct the residual marked process. In that full-sample construction, the first-order estimation effect is part of the limiting null distribution and must be accounted for in the approximation of critical values. By contrast, the present split-sample construction separates parameter fitting from model checking in a way that makes the residual-based process have the same limiting distribution as the oracle process based on the true errors.

Although the limiting null distribution in Theorem (ref) is still nonstandard and requires approximation by bootstrap, the oracle-equivalence result has an important practical implication. The bootstrap procedure in Section (ref) can apply multipliers directly to the split-sample residuals while keeping $\widehat{\boldsymbol{\theta}}_{f_n}$ fixed. Hence, unlike full-sample residual-based procedures that require re-generating data and re-estimating the model in each bootstrap replication, the proposed method avoids repeated refitting and can be substantially cheaper for nonlinear or numerically intensive models.

remarkThe equality $\kappa_{ra}=2\kappa_{ov}$ balances the amount of checking-sample information against the overlap between fitting and checking samples. A simple choice is $f_n=n/2$ and $l_n=n$. Condition (ref) is a score-alignment condition: it matches the covariance between the residual moment process and the estimator influence function with the covariance generated by the model score. A similar requirement appears in davis2025sample for portmanteau statistics under the i.i.d. condition for the error process. This paper extends it to the martingale difference setting. More importantly, in Appendix (ref) and (ref), we verify condition (ref) for the ARMA Gaussian MLE and GARCH Quasi-MLE examples, respectively.

Consistency and Local Alternatives

We proceed to prove the consistency of the proposed test under the fixed and local alternatives, respectively. First, consider the fixed alternative:

align[align omitted — 122 chars of source]

where $\{a_{t}\}$ is strictly stationary and ergodic, with $Ea_{1}^2<\infty$, and importantly, for each $t \in \mathbb{Z}$, $a_{t}$ is $\mathcal{F}_{t-1}$-measurable and $P(a_t=0)<1$. Theorem (ref) shows the asymptotic behavior of $S_{n, w}$ under the global alternative $H_{a}$.

theoremUnder Assumptions (ref)--(ref) and $H_{a}$ (ref), \begin{align*} l_n^{-1/2}S_{n,w}(\cdot,\widehat{\boldsymbol{\theta}}_{f_n}) \stackrel{p}{\longrightarrow} L_w(\cdot), \qquad in L_2(\Pi,\nu), \end{align*} where \begin{align*} L_w(\boldsymbol{\eta}):=\sum_{j=1}^\infty \varsigma_j(\mathbf{x})\Psi_j(\lambda), \quad \varsigma_j(\mathbf{x})=E[a_t w_{t-j}(\mathbf{x})]. \end{align*}
corollaryLet $\Xi$ denote the class of alternatives $\{a_t\}$ such that, under $H_a$ (ref), there exists some $j\ge 1$ for which $\varsigma_j(\mathbf{x})\neq 0$ on a subset of $\Upsilon$ with positive W-measure. Then we have \begin{align*} D_{n,w}^2(\widehat{\boldsymbol{\theta}}_{f_n}) \stackrel{p}{\longrightarrow}\infty, \end{align*} which means $D_{n,w}^2$ is consistent against all alternatives in $\Xi$.

Corollary (ref) indicates that the test is consistent against alternatives that violate at least one of the pairwise conditional mean restrictions, namely $P(E[a_t|\boldsymbol{Z}_{t-j}]=0)<1$ for some $j\geq 1$. As in generalized spectral tests based on pairwise conditional moments, this is a large but not exhaustive class. It is indeed possible that $P(E[a_t| \mathbf I_{t-1}]=0)<1$ while all pairwise projections satisfy $E[a_t | \boldsymbol{Z}_{t-j}]=0$ a.s. for every fixed $j\geq 1$. Such alternatives involve higher-order or genuinely joint dependence on the past and are therefore not captured by a statistic built only from pairwise restrictions.

commentIt is important to mention that although the set $\Xi$ forms a large class of alternatives, it does not cover all possible alternatives. The test statistic $D_{n, w}^{2}$ will not be able to detect those alternatives in the complement of $\Xi$.

To further study the consistency properties of the proposed tests, we consider the local alternatives with a sequence of alternative hypotheses tending to the null at the parametric rate $l_n^{-1 / 2}$ as in escanciano2006goodness:

equation[equation omitted — 144 chars of source]

where $\{a_{t}\}$ is the same as in $H_{a}$ (ref). Note that the rate $l_n$ is of the same order as $n$ in the half-splitting scheme.

An additional assumption is needed regarding the behavior of the estimator under these local alternatives.

assumptionThe estimator $\widehat{\boldsymbol{\theta}}_{n}$ satisfies the asymptotic expansion under $H_{a, n}$ (ref), $$ \sqrt{f_n} (\widehat{\boldsymbol{\theta}}_{f_n}-\boldsymbol{\theta}_{0})= \boldsymbol{\xi}_{a}+\frac{1}{\sqrt{f_n}} \sum_{t=1}^{f_n} \mathbf{h} (Y_{t}, \mathbf{I}_{t-1}, \boldsymbol{\theta}_{0} )+o_{p}(1) $$ where the function $\mathbf{h}(\cdot)$ is as in Assumption (ref) and $\boldsymbol{\xi}_{a} \in \mathbb{R}^{p}$.
theoremUnder the sequence of alternative hypotheses (ref) and Assumptions (ref)--(ref), if $\kappa_{ra} = 2 \kappa_{ov}$ and condition (ref) holds, then \begin{align*} S_{n, w} \Longrightarrow S_{w}^0 + L_{w}(\cdot)-\sqrt{\kappa_{ra}} \mathbf{G}_{w}^{\prime}(\cdot) \boldsymbol{\xi}_{a}, \end{align*} where $S_{w}^0$ and $L_{w}$ are the processes defined in Theorems (ref) and (ref), and $\kappa_{ra}\equiv \lim_{n \rightarrow \infty} l_n/f_n$. Moreover, \begin{align*} D_{n, w}^{2}(\widehat{\boldsymbol{\theta}}_{f_n}) \stackrel{d}{\longrightarrow} \int_{\Pi} |S_{w}^0(\boldsymbol{\eta}, \boldsymbol{\theta}_{0})+L_{w}(\boldsymbol{\eta})-\sqrt{\kappa_{ra}} \mathbf{G}_{w}^{\prime}(\boldsymbol{\eta}) \boldsymbol{\xi}_{a}| ^{2} d \nu(\boldsymbol{\eta}). \end{align*}

Theorem (ref) decomposes the local limit into three parts: the oracle Gaussian limit $S_w^0$, the direct local-alternative drift $L_w$, and the indirect drift caused by estimating $\boldsymbol{\theta}_0$ on the fitting sample. If $\boldsymbol{\xi}_{a} \neq \mathbf{0}$, local power may be reduced in directions that are aligned with the score $\mathbf{g}_{t}\left(\boldsymbol{\theta}_{0}\right)$. If $\boldsymbol{\xi}_{a}=\mathbf{0}$, the estimator has no first-order local drift, and the test has nontrivial power against all local alternatives in $\Xi$.

remarkCompared with escanciano2006goodness, the local power has two changes in the limiting process. Escanciano's statistic has the local limit $S_w+L_w(\cdot) - \mathbf{G}_{w}^{\prime}(\cdot) \boldsymbol{\xi}_{a}$, where the Gaussian component $S_w$ contains the first-order estimation effect. In contrast, the split-sample statistic has the limit $S_{w}^0 + L_{w}(\cdot)-\sqrt{\kappa_{ra}} \mathbf{G}_{w}^{\prime}(\cdot) \boldsymbol{\xi}_{a}$. Thus, sample splitting removes the estimation effect from the stochastic part of the limit, and changes the deterministic score-induced drift from $\mathbf{G}_{w}^{\prime} \boldsymbol{\xi}_{a}$ to $\sqrt{\kappa_{ra}} \mathbf{G}_{w}^{\prime} \boldsymbol{\xi}_{a}$. Consequently, there is no uniform local power dominance. If the local alternative is nearly orthogonal to the score direction, so that $\mathbf{G}_{w}^{\prime} \boldsymbol{\xi}_{a}$ is small, then the two tests have essentially the same local drift $L_w$, while the split-sample test has the oracle stochastic component. In this case, sample splitting can yield higher local power. If $\mathbf{G}_{w}^{\prime} \boldsymbol{\xi}_{a}$ is not negligible, the power gain or loss depends on the sign and magnitude of $L_w(\cdot) - \sqrt{\kappa_{ra}} \mathbf{G}_{w}^{\prime}(\cdot) \boldsymbol{\xi}_{a}$ relative to $L_w(\cdot) - \mathbf{G}_{w}^{\prime}(\cdot) \boldsymbol{\xi}_{a}$.

We will show in Section (ref) that, in finite samples, the proposed tests exhibit empirical power comparable to that of the full-sample generalized spectral tests in most linear and nonlinear examples.

Bootstrap Approximation

The limiting null distribution in Theorem (ref) is nonstandard because the covariance operator of the limiting Gaussian process depends on the underlying data-generating process (DGP) in a very complicated way. Hence, direct tabulation of the critical value is generally infeasible. However, compared with escanciano2006goodness, a key observation is that the asymptotic distribution of $S_{n, w}(\boldsymbol{\eta}, \widehat{\boldsymbol{\theta}}_{f_n})$ is the same as $S_w^0(\boldsymbol{\eta})$, which does not involve any estimation effect of $\widehat{\boldsymbol{\theta}}_{f_n}$. We elaborate more on this point.

In the full-sample generalized spectral test of escanciano2006goodness, because the asymptotic distribution of the test therein involves the estimation effect, the bootstrap must reproduce this effect. Hence, escanciano2006goodness applies the fixed-design wild bootstrap: one first generates bootstrap residuals, then constructs bootstrap responses, then { re-estimates} the model on the bootstrap sample to obtain $\boldsymbol{\theta}_n^*$, and finally recomputes bootstrap residuals using $\boldsymbol{\theta}_n^*$. This is why the procedure requires additional high-level assumptions on the bootstrap estimator.

The split-sample construction in this paper changes the bootstrap problem. By Theorem (ref), under the split-overlap and score-alignment conditions (ref), the residual-based process $S_{n,w}(\cdot,\widehat{\boldsymbol{\theta}}_{f_n})$ has the same limiting null distribution as the oracle process based on the true errors. Hence, the bootstrap only needs to mimic the oracle martingale fluctuation on the checking sample. This means that the fitting-sample estimator $\widehat{\boldsymbol{\theta}}_{f_n}$ can be kept fixed throughout the bootstrap, and no bootstrap responses are generated. This is the main computational advantage of the proposed procedure, especially for nonlinear models or models whose estimation requires cumbersome numerical optimization.

Specifically, let $\{V_t\}$ be a sequence of i.i.d. random variables, independent of the data, with zero mean, unit variance, and bounded support. Conditional on the observed sample, define

align[align omitted — 269 chars of source]

where

align[align omitted — 262 chars of source]

This is a multiplier version of the checking-sample moment process. It is related to the wild bootstrap of Wu1986bootstrap, liu1988bootstrap, and mammen1993bootstrap, but unlike the full-sample fixed-design wild bootstrap used in escanciano2006goodness, it does not require constructing $Y_t^*$ or re-estimating the parameter.

Common choices of $\{V_t\}$ include Mammen's two-point distribution,

align[align omitted — 151 chars of source]

and the Rademacher distribution $P(V_t=1)=P(V_t=-1)=1/2$; see also stute1998bootstrap, de1996bierens, and mammen1993bootstrap.

The complete bootstrap algorithm is summarized in Algorithm (ref).

algorithm[algorithm omitted — 1,387 chars of source]

The algorithm keeps $\widehat{\boldsymbol{\theta}}_{f_n}$ fixed in every bootstrap replication. Therefore, the computational cost of the bootstrap is dominated by recomputing the weighted residual moments, rather than by repeated model estimation. This distinction is minor for very simple linear models, but it can be substantial for nonlinear volatility models, and threshold models, see Section (ref) for numerical evidence.

We then provide the validity of the bootstrap procedure.

theoremAssume that Assumptions (ref)--(ref) and the conditions in Theorem (ref) hold. Under the null hypothesis $H_0$, under any fixed alternative hypothesis or under the local alternatives, \begin{align*} S_{n,w}^* \underset{*}{\Longrightarrow} \widetilde{S}_w \quad a.s., \end{align*} where $\widetilde{S}_w$ is the same Gaussian process of Theorem (ref) but with $\theta_*$ replacing $\boldsymbol{\theta}_0$ and $\underset{*}{\Longrightarrow}$ denoting weak convergence a.s. under the bootstrap law (see gine1990bootstrapping).

Simulation Studies

To examine the finite-sample performance of the proposed tests, we carry out simulation studies with different DGPs under the nulls and under the alternatives. Here we consider the univariate case, that is, $m=1$ with $Z_{t}=Y_{t}, t \in \mathbb{Z}$. Denote $D_{n, I}^{2}$ and $D_{n, C}^{2}$ to be the proposed CvM tests corresponding to $w(Y_{t}, x)=\mathbbm{1}(Y_{t} \leq x)$ and $w (Y_{t}, x)=\exp (i Y_{t} x)$, where $W(\cdot)$ is chosen to be the empirical CDF $F_{n}(\cdot)$ based on $\{Y_{t-1}\}_{t=n-l_n+1}^{n}$ and the CDF $\Phi(\cdot)$ of a standard normal random variable, respectively. Then,

align*[align* omitted — 173 chars of source]

with $\widehat{\gamma}_{j, I}(x, \widehat{\boldsymbol{\theta}}_{f_n})= (\widehat{\sigma}_{e} n_{j})^{-1} \sum_{k=n-l_n+j}^{n} \widehat{e}_{k} (\widehat{\boldsymbol{\theta}}_{f_n}) \mathbbm{1}(Y_{k-j} \leq x )$, $\widehat{\sigma}_{e}^{2}=l_n^{-1} \sum_{s=n-l_n+1}^{n} \widehat{e}_{s}^{2}(\widehat{\boldsymbol{\theta}}_{f_n})$, and

align*[align* omitted — 297 chars of source]

Other choices of integrating functions $W(\cdot)$ in $D_{n, C}^{2}$ can also be applied. Besides, we also compare the finite-sample performance with that of CvM tests in escanciano2006goodness without sample splitting, which are denoted as $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ using FDWB approximation for the bootstrapped statistics.

commentFor ease of comparison, we adopt exactly the same simulation setup and test statistics as those in escanciano2006goodness. The chosen test statistics are as follows: \begin{itemize} • The classical Portmanteau test of ljung1978measure, designed for ARMA$(p, q)$ models, with statistics \begin{align*} LB_{m} = n(n+2) \sum_{j=1}^{m}(n-j)^{-1} \widehat{\rho}_{e, j}^{2}, \end{align*} where $\widehat{\rho}_{e, j}$ is the autocorrelation coefficient of residuals at lag $j$ from the ARMA $(p, q)$ fitted model. Under i.i.d. errors and $H_{0}$, the asymptotic distribution of $LB_{m}$ can be approximated by a $\chi_{m-p-q}^{2}$ distribution ($m>p+q$). • The generalized spectral test proposed by hong2005generalized for a parametric conditional mean under possible conditional heteroscedasticity. The test statistic is \begin{align} HL_{n}(p)=[L_{2}^{2}(p)-\widehat{C}_{1}(p)] /[\sqrt{\widehat{D}_{1}(p)}], \end{align} where \begin{align*} L_{2}^{2}(p)=\sum_{j=1}^{n} n_{j} k^{2}\Big(\frac{j}{p}\Big) \int_{\mathbb{R}} |\widehat{\gamma}_{j}^{e} (x, \widehat{\boldsymbol{\theta}}_{n})|^{2} W(dx), \end{align*} and $\widehat{\gamma}_{j}^{e}(x, \boldsymbol{\theta}_{n})=n_{j}^{-1} \sum_{t=j}^{n} e_{t}(\widehat{\boldsymbol{\theta}}_{n}) \widehat{\psi}_{t-j}(x)$, $\widehat{\psi}_{t}(x)=\exp (i x e_{t}(\widehat{\boldsymbol{\theta}}_{n})) -\widehat{\varphi}(x)$, $\widehat{\varphi}(x)=n^{-1} \sum_{t=1}^{n} \exp (i x e_{t}(\widehat{\boldsymbol{\theta}}_{n}))$, $k(\cdot)$ is a symmetric kernel, $p$ is a bandwidth parameter, and $W(\cdot)$ is an integrating function. In (ref), $\widehat{C}_{1}(p)$ and $\widehat{D}_{1}(p)$ are the centering and scaling factors to obtain an asymptotic standard normal null distribution, which depend on the higher dependence structure between the errors and the regressors. In hong2005generalized, the centering and scaling factors are \begin{align*} \widehat{C}_{1}(p)=&\sum_{j=1}^{n} n_{j}^{-1} k^{2}\Big(\frac{j}{p}\Big) \sum_{t=j}^{n} \widehat{e}_{t}^{2}(\widehat{\boldsymbol{\theta}}_{n}) \int_{\mathbb{R}} |\widehat{\psi}_{t-j}(x)|^{2} W(dx),\\ \widehat{D}_{1}(p) =& 2 \sum_{j=1}^{n-1} \sum_{l=1}^{n-1} k^{2}\Big(\frac{j}{p}\Big) k^{2}\Big(\frac{l}{p}\Big) \\ & \times \int_{\mathbb{R}^{2}} \bigg| \frac{1}{n-\max (j, l)+1} \sum_{t=\max (j,l)}^{n} \widehat{e}_{t}^{2}(\widehat{\boldsymbol{\theta}}_{n}) \widehat{\psi}_{t-j}(x) \widehat{\psi}_{t-l}(y)\bigg|^{2} W(dx) W(dy). \end{align*} Under $H_{0}$ and certain conditions, $HL_{n}(p)$ converges to a standard normal random variable. As in the simulations of hong2005generalized, we choose the CDF of $N(0,1)$ as the integrating function $W$ and the Daniell kernel $k(z)=\sin (\pi z) /(\pi z)$. • The CvM test of bierens1982consistent, which considers a fixed number of lags in the conditioning set. The Bierens test is based on the weight function $w(\mathbf{I}_{t-1}^{d}, \mathbf{x})=\exp (i \mathbf{x}^{\prime} \mathbf{I}_{t-1}^{d})$ and the CDF of a multivariate standard normal random vector as the integrating measure, where $\mathbf{I}_{t-1}^{d}=(Y_{t-1}, \ldots, Y_{t-d})^{\prime}$ is the $d$-lagged values of the series. The Bierens test statistics is defined as \begin{align*} CvM_{\exp,d} = \frac{1}{\widehat{\sigma}_{e}^2 n} \sum_{t=1}^{n} \sum_{s=1}^{n} e_{t} (\widehat{\boldsymbol{\theta}}_{n}) e_{s}(\widehat{\boldsymbol{\theta}}_{n}) \exp \Big(-\frac{1}{2}|\mathbf{I}_{t-1}^{d}-\mathbf{I}_{s-1}^{d}|^{2}\Big). \end{align*} • The CvM and Kolmogorov-Smirnov (KS) tests of a multivariate version of Koul1999 proposed by escanciano2007model. It is based on residual marked empirical processes and uses the FDWB approximation for a large class of residual marked tests, including the Bierens test and the multivariate extension of Koul1999 as special cases. Denote $CvM_{d}$ and $KS_{d}$ to be the CvM and KS statistics proposed by escanciano2007model, which are given by \begin{align*} & CvM_{d}=\frac{1}{\widehat{\sigma}_{e}^{2} n^{2}} \sum_{j=1}^{n} \bigg[\sum_{t=1}^{n} \widehat{e}_{t}(\widehat{\boldsymbol{\theta}}_{n}) \mathbbm{1} (\mathbf{I}_{t-1}^{d} \leq \mathbf{I}_{j-1}^{d})\bigg]^{2},\\ & KS_{d}=\max_{1 \leq i \leq n} \bigg|\frac{1}{\widehat{\sigma}_{e} \sqrt{n}} \sum_{t=1}^{n} \widehat{e}_{t} (\widehat{\boldsymbol{\theta}}_{n}) \mathbbm{1}(\mathbf{I}_{t-1}^{d} \leq \mathbf{I}_{i-1}^{d})\bigg|. \end{align*} \end{itemize}

For the simulation setup, we consider sample sizes $n=200$ and $n=500$, and the nominal level of $5\%$ for each test. The Monte Carlo experiments are repeated $1,000$ times, and in each replication, the initial $200$ burn-in observations of the generated processes are discarded to ensure stationarity. The number of bootstrap replications is $B=500$, and the random weighting $\{V_{t}\}$ is generated by i.i.d. Mammen's distribution (ref).

Linear ARMA Model

First we consider the linear case, where the null model is the AR$(1)$ model as follows:

align[align omitted — 151 chars of source]

We examine the goodness-of-fit of this model under the following DGPs:

enumerate• AR(1) model: $Y_{t}= 0.6 Y_{t-1}+\varepsilon_{t}$; • AR(1) model with exponential centered noise (AR-EXP): $Y_{t}=0.6 Y_{t-1}+\bar{\epsilon}_{t}$, $\bar{\epsilon}_{t}=\epsilon_{t}-1, \epsilon_{t} \sim \exp (1)$; • AR(1) model with heteroscedasticity (AR-HET): $Y_{t}=0.6 Y_{t-1}+h_{t} \varepsilon_{t}$, $ h_{t}^{2}=0.1+ 0.1 Y_{t-1}^{2}$; • AR(1) model plus a bilinear term (AR-BIL): $Y_{t}=0.6 Y_{t-1}+ 0.1 Y_{t-1} \varepsilon_{t}+\varepsilon_{t}$; • AR(2) model: $Y_{t} = 0.6 Y_{t-1}- 0.5 Y_{t-2}+\varepsilon_{t}$; • ARMA(1,1) model: $Y_{t}= 0.6 Y_{t-1}+0.9 \varepsilon_{t-1}+\varepsilon_{t}$; • Bilinear model (BIL): $Y_{t}=0.6 Y_{t-1} + 0.7 \varepsilon_{t-1} Y_{t-2}+\varepsilon_{t}$; • Nonlinear moving average model (NLMA): $Y_{t}=0.6 Y_{t-1}+0.7 \varepsilon_{t-1} \varepsilon_{t-2}+\varepsilon_{t}$; • Threshold autoregressive model (TAR): $Y_{t}=0.6 Y_{t-1} + \varepsilon_{t}$ if $Y_{t-1}<1$ and $Y_{t}=-0.5 Y_{t-1} + \varepsilon_{t}$ if $Y_{t-1} \geq 1$; • Sign autoregressive model (SIGN): $Y_{t}=\operatorname{sign} (Y_{t-1})+ 0.43 \varepsilon_{t}$, where $\operatorname{sign}(x)=\mathbbm{1}(x>0)-\mathbbm{1}(x<0)$; • Temp map model (TEM MAP): $Y_{t}=\alpha^{-1} Y_{t-1}$ if $0 \leq Y_{t-1}<\alpha$ and $Y_{t}=(1-\alpha)^{-1} (1-Y_{t-1})$ if $\alpha \leq Y_{t-1} \leq 1$, where $\alpha=0.49999$ and $Y_{0}$ is generated from the uniform distribution on $[0,1]$; • Nonlinear autoregressive model (NAR): $Y_{t}=0.6 Y_{t-1} + 0.7 \sin (0.3 \pi Y_{t-2})+\varepsilon_{t}$.

DGPs 1--4 are introduced to examine the empirical size of the tests under the nulls, and DGPs 5--12 are used to examine the power under the alternatives.

table[table omitted — 830 chars of source]
table[table omitted — 1,119 chars of source]

For the computation of all test statistics, we apply the usual least squares residuals under $H_0$, that is, $\widehat{e}_{t} (\widehat{\boldsymbol{\theta}}_{f_n})=Y_{t}-\widehat{b}_{f_n} Y_{t-1}$, for $t=n-l_n+1,\ldots,n$. Table (ref) reports the empirical rejection probabilities (RPs) associated with DGPs 1--4 to examine the empirical size of the tests, and Table (ref) reports the empirical RP associated with DGPs 5-12 to examine the empirical power of the tests in the cases of $n=200$ and $n=500$.

Table (ref) shows that the proposed split-sample generalized spectral testing framework attains empirical sizes close to the nominal 5% level in finite samples under the linear AR(1) null models, and that these sizes are comparable to those of the generalized spectral tests in escanciano2006goodness. This finding supports Theorem (ref), which states that the proposed test has the same null limiting distribution as the oracle process. Table (ref) indicates that the proposed tests exhibit empirical power comparable to that of the full-sample generalized spectral tests, and that the power increases as the sample size grows from $n=200$ to $n=500$, which validates Theorem (ref) for the consistency of the proposed tests.

comment\begin{table}[t] \begin{center} \caption{\textcolor{red}{Empirical Size (%) of Tests at $5\%$ with sample size $n=100$, with intercept in linear $H_0$ (ref).}} \begin{tabular}{l*{4}{S}} \hline \hline $n=100$ & {AR(1)} & {AR(1)-EXP} & {AR(1)-HET} & {AR(1)-BIL} \\ \hline ${D}_{n, I}^{2}$ & 8.4 & 7.1 & 10.1 & 13.0 \\ $D_{n, C}^{2}$ & 10.2 & 7.7 & 17.0 & 11.8 \\ $\widetilde{D}_{n, I}^{2}$ & 3.5 & 3.2 & 5.9 & 4.1 \\ $\widetilde{D}_{n, C}^{2}$ & 5.2 & 4.9 & 7.8 & 10.4 \\ $H L_{n}(2)$ & 0.8 & 0.3 & 1.1 & 1.3 \\ $H L_{n}(6)$ & 1.1 & 0.3 & 0.9 & 0.6 \\ $H L_{n}(10)$ & 1.0 & 0.4 & 0.9 & 0.9 \\ $L B_{2}$ & 5.5 & 4.5 & 27.5 & 4.8 \\ $L B_{6}$ & 4.5 & 3.3 & 31.9 & 5.3 \\ $L B_{10}$ & 5.4 & 4.5 & 27.1 & 4.4 \\ $CvM_{\exp,1}$ & 4.9 & 4.8 & 5.8 & 10.0 \\ $CvM_{\exp,3}$ & 4.8 & 3.3 & 8.4 & 7.3 \\ $CvM_{\exp,5}$ & 4.3 & 3.2 & 7.7 & 6.2 \\ $K S_{1}$ & 3.2 & 4.5 & 3.8 & 5.8 \\ $K S_{3}$ & 5.2 & 4.9 & 4.5 & 4.8 \\ $K S_{5}$ & 6.0 & 5.0 & 6.6 & 5.1 \\ $CvM_{1}$ & 4.1 & 3.8 & 4.2 & 5.7 \\ $CvM_{3}$ & 4.2 & 3.2 & 7.6 & 0.6 \\ $CvM_{5}$ & 4.9 & 5.5 & 10.2 & 0.8 \\ \hline \end{tabular} \end{center} \end{table} \begin{table}[htbp] \caption{\textcolor{red}{Empirical Power (%) of Tests at $5 \%$ with sample size $n=100$, with intercept in linear $H_0$ (ref).}} \begin{tabular}{l *{8}{S}} \hline \hline $n=100$ & {AR$(2)$} & {ARMA$(1,1)$} & {BIL} & {NLMA} & {TAR} & {SIGN} & {TEM MAP} & {NAR} \\ \hline ${D}_{n, I}^{2}$ & 1.9 & 6.1 & 4.6 & 5.0 & 5.2 & 43.3 & 87.6 & 24.9 \\ $D_{n, C}^{2}$ & 2.3 & 4.7 & 19.6 & 11.3 & 4.9 & 48.3 & 5.7 & 35.9 \\ $\widetilde{D}_{n, I}^{2}$ & 97.3 & 35.0 & 13.5 & 11.4 & 4.6 & 58.6 & 100.0 & 39.2 \\ $\widetilde{D}_{n, C}^{2}$ & 84.1 & 7.5 & 24.7 & 16.2 & 7.5 & 59.3 & 100.0 & 46.7 \\ $H L_{n}(2)$ & 37.6 & 60.0 & 1.1 & 2.0 & 0.8 & 44.7 & 46.2 & 10.1 \\ $H L_{n}(6)$ & 93.4 & 60.8 & 1.1 & 0.9 & 0.7 & 38.8 & 18.0 & 15.7 \\ $H L_{n}(10)$ & 91.5 & 49.0 & 0.7 & 0.8 & 0.5 & 34.1 & 12.9 & 11.4 \\ $L B_{2}$ & 99.6 & 97.2 & 19.6 & 7.9 & 6.2 & 49.6 & 7.0 & 32.7 \\ $L B_{6}$ & 98.7 & 80.6 & 22.3 & 6.4 & 5.7 & 41.4 & 6.4 & 22.0 \\ $L B_{10}$ & 97.4 & 71.7 & 16.9 & 7.2 & 4.8 & 36.5 & 6.0 & 18.7 \\ $CvM_{\exp,1}$ & 0.7 & 0.5 & 24.1 & 18.2 & 6.9 & 58.5 & 100.0 & 19.1 \\ $CvM_{\exp,3}$ & 93.1 & 33.0 & 38.1 & 23.3 & 5.2 & 60.4 & 68.0 & 61.4 \\ $CvM_{\exp,5}$ & 67.4 & 15.6 & 15.9 & 10.4 & 3.8 & 58.6 & 41.7 & 33.2 \\ $K S_{1}$ & 1.4 & 1.3 & 15.2 & 10.6 & 5.0 & 58.3 & 100.0 & 13.3 \\ $K S_{3}$ & 88.0 & 32.0 & 13.4 & 7.8 & 6.5 & 35.6 & 6.6 & 36.9 \\ $K S_{5}$ & 21.5 & 22.1 & 7.9 & 5.4 & 5.0 & 24.2 & 0.1 & 21.7 \\ $CvM_{1}$ & 0.4 & 0.8 & 16.3 & 14.1 & 5.8 & 58.4 & 100.0 & 12.5 \\ $CvM_{3}$ & 85.4 & 31.3 & 16.5 & 6.4 & 6.1 & 34.0 & 73.0 & 37.3 \\ $CvM_{5}$ & 2.5 & 17.8 & 2.1 & 2.2 & 5.4 & 31.9 & 20.5 & 25.9 \\ \hline \end{tabular} \end{table} \begin{table}[t] \begin{center} \caption{\textcolor{red}{Empirical Size (%) of Tests at $5\%$ with sample size $n=500$, with intercept in linear $H_0$ (ref).}} \begin{tabular}{l*{4}{S}} \hline \hline $n=500$ & {AR(1)} & {AR(1)-EXP} & {AR(1)-HET} & {AR(1)-BIL} \\ \hline ${D}_{n, I}^{2}$ & 7.0 & 7.2 & 10.8 & 11.7 \\ $D_{n, C}^{2}$ & 7.1 & 6.9 & 26.3 & 13.5 \\ $\widetilde{D}_{n, I}^{2}$ & 2.8 & 3.7 & 6.1 & 8.1 \\ $\widetilde{D}_{n, C}^{2}$ & 4.9 & 3.9 & 9.2 & 11.1 \\ \hline \end{tabular} \end{center} \end{table} \begin{table}[t] \caption{\textcolor{red}{Empirical Power (%) of Tests at $5 \%$ with sample size $n=500$, with intercept in linear $H_0$ (ref).}} \begin{tabular}{l *{8}{S}} \hline \hline $n=500$ & {AR$(2)$} & {ARMA$(1,1)$} & {BIL} & {NLMA} & {TAR} & {SIGN} & {TEM MAP} & {NAR} \\ \hline ${D}_{n, I}^{2}$ & 85.5 & 6.6 & 7.1 & 2.8 & 5.6 & 95.8 & 100.0 & 25.0 \\ $D_{n, C}^{2}$ & 99.4 & 10.1 & 72.0 & 36.3 & 7.1 & 98.8 & 11.8 & 87.0 \\ $\widetilde{D}_{n, I}^{2}$ & 100.0 & 99.8 & 62.3 & 63.7 & 10.3 & 98.8 & 100.0 & 100.0 \\ $\widetilde{D}_{n, C}^{2}$ & 100.0 & 68.0 & 81.8 & 80.7 & 14.3 & 98.8 & 100.0 & 99.8 \\ \hline \end{tabular} \end{table} {\color{red} \\ 9.27 Discussion: 1. try simulations of sample-spliting version of (shao&zhu, 2022), multivariate response, see the performance. only compare the selected test statistics. 2. n=2000/5000, see the first two tests, whether there is improvement. 3. computational time, list a table. n=100, 500, 2000, 5000, five replications, take the mean, use own computer. 4. try rademacher distribution, P(w=1)=0.5, P(w=-1)=0.5, n=100 select other DGPs in the literature (shao&zhu, 2022) Or try other Null hypothesis, GARCH(1,1) simulation, may increase improved computational time. 5. Look into GARCH model goodness-of-fit literature, compare. Point: 1. comparable performance; 2. improved computational time \\ 12.6 Discussion: Improve computational time: try more bootstrap times B, choose other complicated DGP, multivariate case,... } \begin{table}[t] \begin{center} \caption{\textcolor{red}{Empirical Size (%) of Tests at $5\%$ with sample size $n=100$ using Rademacher weight, without intercept in linear $H_0$ (ref).}} \begin{tabular}{l*{4}{S}} \hline \hline $n=100$ & {AR(1)} & {AR(1)-EXP} & {AR(1)-HET} & {AR(1)-BIL} \\ \hline ${D}_{n, I}^{2}$ & 5.8 & 5.9 & 6.1 & 14.6 \\ $D_{n, C}^{2}$ & 6.0 & 6.5 & 6.5 & 10.6 \\ $\widetilde{D}_{n, I}^{2}$ & 5.2 & 5.1 & 4.3 & 6.2 \\ $\widetilde{D}_{n, C}^{2}$ & 5.2 & 5.6 & 4.3 & 8.7 \\ \hline \end{tabular} \end{center} \end{table} \begin{table}[t] \begin{center} \caption{\textcolor{red}{Empirical Size (%) of Tests at $5\%$ with sample size $n=100$ using Normal weight $N(0,1)$, without intercept in linear $H_0$ (ref).}} \begin{tabular}{l*{4}{S}} \hline \hline $n=100$ & {AR(1)} & {AR(1)-EXP} & {AR(1)-HET} & {AR(1)-BIL} \\ \hline ${D}_{n, I}^{2}$ & 5.8 & 6.7 & 6.1 & 13.4 \\ $D_{n, C}^{2}$ & 5.1 & 6.6 & 9.0 & 9.9 \\ $\widetilde{D}_{n, I}^{2}$ & 4.5 & 4.8 & 4.3 & 5.0 \\ $\widetilde{D}_{n, C}^{2}$ & 4.1 & 4.4 & 4.4 & 6.3 \\ \hline \end{tabular} \end{center} \end{table}

Nonlinear Volatility Model

In the nonlinear case, we are interested in analyzing the size and power against misspecifications in the conditional variance. We consider two common conditional heteroscedasticity models:

1. ARCH-type null model, e.g., ARCH$(1)$ model:

align[align omitted — 187 chars of source]

and ARCH$(4)$ model:

align[align omitted — 203 chars of source]

2.

commentOur null model is a GARCH$(1,1)$ model: \begin{align} &H_0: y_t=\sigma_t \eta_t, \sigma_{t}^{2}= \omega + \phi y_{t-1}^{2} + \psi \sigma_{t-1}^{2}, \quad where \quad \eta_{t} \overset{i.i.d.}{\sim} N(0,1). \end{align}

GARCH-type null model, e.g., GARCH$(2,2)$ model:

align[align omitted — 252 chars of source]

Recall that under the null, $Y_t$ corresponds to $y_t^2$, $E[Y_t | \mathbf{I}_{t-1}] = \sigma_t^2$, and the MDS $\varepsilon_t = Y_t - E[Y_t | \mathbf{I}_{t-1}] = \sigma_t^2 (\eta_t^2 -1)$. We examine the goodness-of-fit of this model under the following DGPs, where for all cases $\eta_{t} \overset{\text{i.i.d.}}{\sim} N(0,1)$:

enumerate• ARCH(1) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.9 + 0.1 y_{t-1}^{2}$; • ARCH(2) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.9 + 0.1 y_{t-1}^{2} + 0.8 y_{t-2}^{2}$; • ARCH(4) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.9 + 0.1 y_{t-1}^{2} + 0.2 y_{t-2}^{2} + 0.2 y_{t-3}^{2} + 0.1 y_{t-4}^{2}$; • ARCH(5) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.9 + 0.1 y_{t-1}^{2} + 0.2 y_{t-2}^{2} + 0.2 y_{t-3}^{2} + 0.1 y_{t-4}^{2} + 0.3 y_{t-5}^{2}$; • GARCH(1,1) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.01 + 0.29 y_{t-1}^{2} + 0.7 \sigma_{t-1}^{2}$; • GARCH(2,2) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}=0.1 + 0.2 y_{t-1}^{2} + 0.2 y_{t-2}^{2} + 0.3 \sigma_{t-1}^{2} + 0.1 \sigma_{t-2}^{2}$; • EGARCH(1,1) model: $y_t=\sigma_t \eta_t, ~ \log \sigma_{t}^{2}=0.01 + 0.9 \log \sigma_{t-1}^{2} + 0.3(|\eta_{t-1}| - (2/\pi)^{1/2}) - 0.8 \eta_{t-1}$; • Stochastic volatility (SV) model: $y_t=\sigma_t \eta_t, ~ \sigma_{t}^{2}= 0.1 y_{t-1}^2 + \exp (0.98 \log \sigma_{t-1}^2 + v_t)$; • Bilinear model (BIL): $y_{t}=0.8 \eta_{t-1} y_{t-1}+\eta_{t}$; • Logistic map (LM): $y_{t}= 4 y_{t-1} (1 - y_{t-1})$, where $y_{0}$ is generated from the uniform distribution on $[0,1]$; • Nonlinear moving average model (NLMA): $y_{t}=0.8 \eta_{t-1}^2+\eta_{t}$.

The DGPs for comparisons are the same as those in Escanciano2008.

To compute the test statistics, we apply the quasi-MLE for estimation of unknown parameters $\boldsymbol{\theta}= (\omega, \phi_1, \ldots, \phi_p, \psi_1, \ldots, \psi_q)' \in \mathbb{R}^{p+q+1}$, and define the residuals to be $\widehat{e}_{t} (\widehat{\boldsymbol{\theta}}_{f_n})=Y_{t}-\sigma_t^2(\widehat{\boldsymbol{\theta}}_{f_n})$, for $t=n-l_n+1,\ldots,n$, where $\widehat{\boldsymbol{\theta}}_{f_n}$ is the quasi-MLE for $\boldsymbol{\theta}_0$ defined in Section (ref) using the first $f_n$ samples.

table[table omitted — 1,147 chars of source]

For the ARCH$(1)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) associated with DGPs 1--2, 5, and 7--11 to examine the empirical size and power of tests at 5% with sample size $n=200$ and $n=500$. It shows that the proposed split-sample tests achieve empirical sizes close to the nominal $5\%$ level, with size performance comparable to that of the original full-sample generalized spectral tests. Their empirical power is also broadly comparable across the nonlinear alternatives considered, and it increases as the sample size rises from $n=200$ to $n=500$.

comment\begin{table}[t] \begin{center} \caption{Computational time (s) for one Monte Carlo experiment for each test with five replications, for nonlinear $H_0$ (ref) under ARCH$(1)$ model with DGP 1.} \begin{tabular}{l*{5}{S}} \hline \hline & {$n=100$} & {$n=200$} & {$n=500$} & {$n=1000$} & {$n=2000$} \\ \hline ${D}_{n, I}^{2}$ & 0.47 & 4.49 & 66.04 & 589.89 & 5333.19 \\ $D_{n, C}^{2}$ & 0.60 & 6.71 & 97.96 & 422.46 & 7872.65 \\ $\widetilde{D}_{n, I}^{2}$ & 1.61 & 7.23 & 71.83 & 597.11 & 5360.70 \\ $\widetilde{D}_{n, C}^{2}$ & 1.72 & 9.69 & 103.86 & 432.85 & 8075.34 \\ \hline \end{tabular} \end{center} \end{table}

For the ARCH$(4)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) under the DGPs 3--5 and 7--11. It indicates that the proposed split-sample tests again maintain empirical sizes close to the nominal level and have power comparable to that of the corresponding full-sample tests. Moreover, in several cases, the split-sample procedures outperform the original tests in finite samples, both in terms of size accuracy and empirical power. The power improvement with larger sample sizes is also evident when changing the sample size from $n=200$ to $n=500$.

table[table omitted — 1,138 chars of source]
comment\begin{table}[t] \begin{center} \caption{Computational time (s) for one Monte Carlo experiment for each test with five replications, for nonlinear $H_0$ (ref) under ARCH$(4)$ model with DGP 3.} \begin{tabular}{l*{4}{S}} \hline \hline & {$n=100$} & {$n=200$} & {$n=500$} & {$n=1000$} \\ \hline ${D}_{n, I}^{2}$ & 0.45 & 3.19 & 59.89 & 502.58 \\ $D_{n, C}^{2}$ & 0.50 & 3.14 & 44.54 & 303.35 \\ $\widetilde{D}_{n, I}^{2}$ & 22.60 & 40.18 & 127.86 & 633.49 \\ $\widetilde{D}_{n, C}^{2}$ & 23.24 & 40.12 & 112.50 & 432.64 \\ \hline \end{tabular} \end{center} \end{table}

For the GARCH$(2,2)$ null (ref), Table (ref) reports the empirical rejection probabilities (RPs) under the DGPs 4 and 6--11. It shows a similar pattern that the proposed split-sample tests deliver empirical sizes close to the nominal $5\%$ level and exhibit empirical power broadly comparable to that of the original full-sample tests. In several scenarios, they even improve upon the original procedures in terms of finite-sample size control or rejection power. In the other settings, the power generally increases with the sample size.

table[table omitted — 1,092 chars of source]
comment\begin{table}[t] \begin{center} \caption{Computational time (s) for one Monte Carlo experiment for each test with five replications, under nonlinear $H_0$ (ref) under GARCH$(2,2)$ model with DGP 6.} \begin{tabular}{l*{4}{S}} \hline \hline & {$n=100$} & {$n=200$} & {$n=500$} & {$n=1000$} \\ \hline ${D}_{n, I}^{2}$ & 0.56 & 3.76 & 70.40 & 590.62 \\ $D_{n, C}^{2}$ & 0.80 & 4.50 & 65.55 & 454.28 \\ $\widetilde{D}_{n, I}^{2}$ & 48.32 & 34.16 & 123.32 & 697.42 \\ $\widetilde{D}_{n, C}^{2}$ & 46.60 & 34.97 & 118.38 & 556.55 \\ \hline \end{tabular} \end{center} \end{table} \begin{table}[t] \begin{center} \caption{\textcolor{red}{Computational time (s) for one Monte Carlo experiment for each test with five replications, under nonlinear $H_0$ (ref) under GARCH$(1,1)$ model with DGP 1.}} \begin{tabular}{l*{4}{S}} \hline \hline & {$n=100$} & {$n=500$} & {$n=1000$} & {$n=2000$} \\ \hline ${D}_{n, I}^{2}$ & 1.02 & 73.77 & & 5193.80 \\ $D_{n, C}^{2}$ & 1.48 & 101.47 & & 6864.54 \\ $\widetilde{D}_{n, I}^{2}$ & 4.12 & 85.76 & & \\ $\widetilde{D}_{n, C}^{2}$ & 5.78 & 132.53 & & \\ \hline \end{tabular} \end{center} \end{table}

Threshold Autoregressive Model

We proceed to study more complex nonlinear cases, for example, the threshold autoregressive (TAR) model. Simulation studies show that our proposed test still has consistently good size and power against model misspecifications.

commenttwo-regime self-exciting TAR(1,1; $d=1$) process as null model: \begin{align} H_0: Y_{t}=\begin{cases} \phi_1 Y_{t-1}+\varepsilon_t, & if Y_{t-1} < r, \\ \psi_1 Y_{t-1}+\varepsilon_t, & if Y_{t-1} \geq r, \end{cases} \quad where \quad \varepsilon_{t} \overset{i.i.d.}{\sim} N(0,1). \end{align}

We consider the following three-regime self-exciting TAR(1,1,1; $d=1$) process as null model, and set the threshold variable to be of lag 1:

align[align omitted — 343 chars of source]

We examine the goodness-of-fit of this model under the following DGPs:

enumerate• Self-exciting TAR(1,1,1; $d=1$) model: $Y_{t}=0.6 Y_{t-1} + \varepsilon_{t}$ if $Y_{t-1}<-1$; $Y_{t}= -0.5 Y_{t-1} + \varepsilon_{t}$ if $-1 \leq Y_{t-1}<1$; and $Y_{t}=0.3 Y_{t-1} + \varepsilon_{t}$ if $Y_{t-1} \geq 1$; • AR(2) model: $Y_{t} = 0.6 Y_{t-1}- 0.5 Y_{t-2}+\varepsilon_{t}$; • ARMA(1,1) model: $Y_{t}= 0.6 Y_{t-1}+0.9 \varepsilon_{t-1}+\varepsilon_{t}$; • Bilinear model (BIL): $Y_{t}=0.6 Y_{t-1} + 0.7 \varepsilon_{t-1} Y_{t-2}+\varepsilon_{t}$; • Nonlinear moving average model (NLMA): $Y_{t}=0.6 Y_{t-1}+0.7 \varepsilon_{t-1} \varepsilon_{t-2}+\varepsilon_{t}$; • Sign autoregressive model (SIGN): $Y_{t}=\operatorname{sign} (Y_{t-1})+ 0.43 \varepsilon_{t}$, where $\operatorname{sign}(x)=\mathbbm{1}(x>0)-\mathbbm{1}(x<0)$; • Temp map model (TEM MAP): $Y_{t}=\alpha^{-1} Y_{t-1}$ if $0 \leq Y_{t-1}<\alpha$ and $Y_{t}=(1-\alpha)^{-1} (1-Y_{t-1})$ if $\alpha \leq Y_{t-1} \leq 1$, where $\alpha=0.49999$ and $Y_{0}$ is generated from the uniform distribution on $[0,1]$; • Nonlinear autoregressive model (NAR): $Y_{t}=0.6 Y_{t-1} + 0.7 \sin (0.3 \pi Y_{t-2})+\varepsilon_{t}$.

To compute the test statistics, for each fixed pair $(r_1,r_2)'$, we first estimate $(\phi_1,\phi_2,\phi_3)'$ by least squares. We then select the optimal threshold pair $(r_1,r_2)'$ over a grid by minimizing the least squares criterion. The residuals $\widehat e_t(\widehat{\boldsymbol\theta}_{f_n})$ are then obtained by plugging in the estimator $\widehat{\boldsymbol{\theta}}_{f_n} = (\widehat\phi_1,\widehat\phi_2,\widehat\phi_3,\widehat r_1,\widehat r_2)'$ using the first $f_n$ samples, for $t=n-l_n+1,\ldots,n$.

Table (ref) reports the empirical sizes and powers (%) of the tests at the nominal $5\%$ level for the null hypothesis $H_0$ of the three-regime TAR$(1,1,1;d=1)$ model in (ref), under DGPs 1--8 and sample sizes $n=200$. We omit the case $n=500$ for now because the full-sample generalized spectral tests are quite time-consuming for the three-regime TAR model, and $n=200$ is sufficient to illustrate the validity of the proposed tests. The results show that the proposed split-sample tests again maintain empirical sizes close to the nominal level and exhibit power comparable to that of the corresponding full-sample tests.

table[table omitted — 764 chars of source]
comment\begin{table}[t] \begin{center} \caption{Computational time (s) for one Monte Carlo experiment of each test with five replications, under the nonlinear $H_0$ (ref) for the three-regime TAR$(1,1; d=1)$ model with DGP 1.} \begin{tabular}{l*{4}{S}} \hline \hline & {$n=100$} & {$n=200$} & {$n=300$} & {$n=400$} \\ \hline ${D}_{n, I}^{2}$ & 1.03 & 6.16 & 22.73 & 54.72 \\ $D_{n, C}^{2}$ & 1.51 & 8.13 & 28.67 & 69.85 \\ $\widetilde{D}_{n, I}^{2}$ & 285.13 & 3649.35 & 22201.49 & 85787.14 \\ $\widetilde{D}_{n, C}^{2}$ & 279.50 & 3609.34 & 22279.66 & 95049.60 \\ \hline \end{tabular} \end{center} \end{table}

Computational Time

We now collect the computational time comparisons across the simulation designs. The reported times are averaged over five replications for one Monte Carlo experiment. Table (ref) reports the corresponding results for the null models in Sections (ref)-(ref), namely the AR(1), ARCH(1), ARCH(4), GARCH(2,2) and three-regime TAR(1,1,1) models.

table[table omitted — 1,038 chars of source]

The timing results highlight where sample splitting is most useful computationally. In the linear AR case, the proposed and full-sample generalized spectral tests have similar computational costs, because least squares estimation is inexpensive even when repeated inside the bootstrap. A similar pattern appears for the ARCH$(1)$ null, where repeated estimation is still relatively light. The advantage becomes much clearer for higher-order volatility models. Under the ARCH$(4)$ and GARCH$(2,2)$ nulls, the full-sample procedures require repeated bootstrap re-estimation of a larger nonlinear model, whereas the proposed tests estimate the model only once and then apply multipliers directly to the split-sample residuals.

The computational gain is most pronounced for the three-regime TAR model. There, each full-sample bootstrap replication requires another threshold search and model refit, which is costly even for moderate sample sizes. By contrast, the split-sample procedure keeps the fitted parameter fixed throughout the bootstrap. As a result, the proposed tests can be much faster while maintaining accurate size and comparable power performance.

Empirical Application

S&P 500 Dynamics

The dynamic behavior of stock returns has been a central topic in financial econometrics, since the specification of the conditional mean and conditional variance directly affects risk measurement, portfolio allocation, and asset pricing. In particular, it is important to distinguish whether the serial dependence in financial returns is mainly driven by the conditional mean, the conditional variance, or both.

We use the same dataset as in the empirical study of Escanciano2008, namely the log differences of the S&P 500 daily stock index in a sample period from January 1, 1988 to May 28, 1993. Following the literature, we delete the last 10% of the observations, leaving 1138 observations. We examine the following three model specifications:

enumerate• an AR(1) model with conditional homoskedastic errors; • a GARCH(1,1) model without the AR(1) component in the conditional mean; • an AR(1)-GARCH(1,1) model, where both the conditional mean and conditional variance are modeled.

The bootstrap $p$-values are reported in Table (ref). The first four columns $D_{n, I}^{2}$ and $D_{n, C}^{2}$ give the results of the proposed sample-splitting tests in (ref), while the last four columns $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ report the corresponding full-sample tests in escanciano2006goodness, Escanciano2008. For each fitted model, we consider both marginal and joint tests for the conditional mean and conditional variance, denoted by the subscripts $m$ and $v$, respectively. The number of bootstrap replications is $B=500$.

Table (ref) shows that the sample-splitting tests yield conclusions very similar to those from the full-sample tests. For the AR(1) model with conditional homoskedastic errors, the $p$-values are generally large for both the conditional mean and conditional variance tests. This indicates that the linear AR(1) specification captures the conditional mean adequately, and there is no strong evidence against conditional homoskedasticity.

When the AR(1) component of the mean is neglected and only a GARCH(1,1) model is fitted, the conditional mean tests strongly reject the null hypothesis. This suggests that ignoring the linear dependence in the conditional mean leads to misspecification. In contrast, for the AR(1)-GARCH(1,1) model, the conditional mean appears to be reasonably specified, but the conditional variance tests produce zero $p$-values. Hence, the GARCH specification for the conditional variance is rejected. These findings are consistent with the conclusions in the existing study of hong2003diagnostic and Escanciano2008: the S&P 500 returns over this period are better described by a linear AR(1) mean structure with conditional homoskedastic errors, rather than by an AR(1)-GARCH(1,1) model.

table[table omitted — 930 chars of source]

Overall, this empirical application confirms the usefulness of the sample-splitting tests. Although the sample-splitting procedure uses only part of the sample for estimation and the remaining part for testing, it delivers conclusions comparable to those from the full-sample tests.

Sunspot Data for the TAR Model

To further illustrate the robustness of the proposed test and its computational advantage, we consider a classical dataset in the TAR literature, that is, the annual Wolf's sunspot data from 1700 to 1979. This dataset has been studied in tong1978threshold, tsay1989testing and the references therein. The series consists of 280 observations and is well known for its asymmetric cyclical behavior. Let $\{Y_t\}_{t=1}^{280}$ denote the annual sunspot series.

Following the model specification in tsay1989testing, we fit $\{Y_t\}_{t=1}^{280}$ using the following three-regime TAR model:

align[align omitted — 349 chars of source]

where $Y_{t-2}$ is selected as the threshold variable, and the AR orders are 11, 10, and 10, respectively. The unknown autoregressive coefficients are estimated by least squares for each fixed pair of threshold values, and the thresholds are selected by minimizing the least squares criterion. In tsay1989testing, the martingale difference assumption that $E(\varepsilon_t|\mathbf{I}_{t-1}) = 0$ is imposed but not formally tested. It motivates us to apply the generalized spectral test to examine the MDS assumption.

Table (ref) reports the bootstrap $p$-values for different models. The first four columns $D_{n, I}^{2}$ and $D_{n, C}^{2}$ give the results of the proposed sample-splitting tests in (ref), while the last four columns $\widetilde{D}_{n,I}^2$ and $\widetilde{D}_{n,C}^2$ report the corresponding full-sample tests of escanciano2006goodness, Escanciano2008. The number of bootstrap replications is $B=500$.

The results show strong evidence in favor of the three-regime TAR(11,10,10) specification for the annual sunspot data. The large bootstrap $p$-values indicate that the residuals from the fitted TAR model do not exhibit significant departures from the martingale difference property. Moreover, the sample-splitting tests lead to conclusions very similar to those obtained from the full-sample tests. Meanwhile, the computational gain is substantial. The proposed sample-splitting tests are more than 20 times faster than the full-sample tests. This gain comes from avoiding repeated model re-estimation and threshold search in the bootstrap procedure, which is computationally expensive for TAR models with multiple regimes and threshold parameters.

For comparison, we also fit the dataset $\{Y_t\}_{t=1}^{280}$ using a constant mean model:

align*[align* omitted — 47 chars of source]

or a linear AR(p) model:

align*[align* omitted — 79 chars of source]

with $p=5,10$, and 11. The testing results are summarized in Table (ref). From this table, we can find that the constant model and AR(5) model are strongly rejected by all tests. Meanwhile, the AR(10) model is rejected by most tests at the 10% significance level, and the bootstrap $p$-values for the AR(11) model remain relatively small compared with those for the three-regime TAR model. Overall, the proposed sample-splitting tests produce conclusions consistent with the full-sample tests and suggest that the constant and linear AR specifications are inadequate for the annual sunspot series.

table[table omitted — 981 chars of source]

Conclusion and Discussion

This paper develops a sample-splitting generalized spectral test for diagnostic checking of parametric time-series models. The proposed procedure combines the omnibus, bandwidth-free structure of generalized spectral tests with a split-sample estimation strategy that removes the first-order effect of parameter estimation under suitable split-overlap and score-alignment conditions. The resulting statistic targets pairwise conditional mean dependence in residuals and therefore provides a diagnostic checking for violations of the martingale-difference implication of a correctly specified conditional mean model.

A practical advantage of the method is that critical values can be obtained from a multiplier bootstrap applied directly to the split-sample residuals. The bootstrap avoids generating artificial time series and avoids re-estimating the model in each replication, which can lead to substantial computational savings for nonlinear or numerically intensive models. The simulation results suggest that the test has a reliable size and competitive power across linear, heteroskedastic, and nonlinear dynamic specifications.

Several extensions are worth pursuing. First, the same sample-splitting principle may be useful for broader conditional moment restriction tests, including recent Gaussian-process-based model checks Escanciano2024. Second, the choice of weight function and integrating measure may affect finite-sample power, and a systematic comparison of indicator, characteristic-function, and data-adaptive weights would be valuable. Third, while the present paper focuses on pairwise conditional mean restrictions, extensions toward diagnostics for conditional quantile restrictions, higher-order or joint conditional mean and variance dependence remain important directions for future research.