EconBase
← Back to paper

Self-normalized tests for multistep conditional predictive ability

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,106 characters · 12 sections · 31 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Self-Normalized Tests for Multistep Conditional Predictive Ability

abstractThis paper proposes self-normalized tests for multistep conditional predictive ability in forecast comparison. By normalizing the sample mean of the transformed loss differential using functionals of its cumulative sum (CUSUM) process, specifically an adjusted-range normalizer for scalars and a matrix normalizer for vectors, our approach avoids direct estimation of the long-run covariance matrix. Consequently, it eliminates the need for the ad hoc bandwidth, kernel, and lag-truncation choices required by traditional methods. We establish the asymptotic theory for these statistics, deriving pivotal null limiting distributions and proving test consistency. Monte Carlo simulations show that the proposed tests effectively mitigate the finite-sample size distortions associated with traditional heteroskedasticity and autocorrelation consistent (HAC) methods, while retaining strong empirical power against conditional predictability alternatives.

\onehalfspacing

Introduction

Forecast comparison remains a central issue in time series econometrics. The traditional literature on forecast evaluation, such as DM95 and West96, primarily focuses on testing the validity of specific forecasting models at the population level. In practice, however, it is often more relevant to compare forecasting methods rather than the underlying specifications of particular models. A forecasting method may encompass model-based, semiparametric, nonparametric, or algorithmic approaches; the primary object of interest is the sequence of forecasts generated from the information available at each forecast origin. The conditional predictive ability (CPA) framework proposed by GW06 is well suited to this perspective, as it tests whether relative forecast performance is predictable based on the information set available at the time the forecasts are formed.

While the CPA framework is versatile, evaluating multistep forecasts within this setup can introduce nontrivial statistical complications in finite samples. Analytically, overlapping forecast errors tend to give rise to an $\mathrm{MA}(\tau-1)$ serial correlation structure within the transformed loss differential. Consequently, the asymptotic null distribution of the test statistic depends on a complex nuisance parameter: a long-run variance in the scalar case, or a long-run covariance matrix in the vector-valued case. To address this nuisance parameter, standard CPA tests typically employ heteroskedasticity and autocorrelation consistent (HAC) estimation (NW87,Andrews91). However, HAC estimators necessitate the selection of bandwidths and kernel functions, and their convergence rates can be relatively slow. In finite samples, particularly when the forecast horizon $\tau$ is large or the underlying data exhibit high persistence, direct HAC estimation may yield ill-conditioned covariance matrices. This instability can lead to pronounced size distortions (i.e., over-rejection of the null hypothesis), which may constrain the reliability of traditional multistep CPA tests in macroeconomic and financial applications where evaluation samples are limited.

To circumvent the limitations associated with direct HAC estimation, we propose a self-normalized (SN) multistep CPA test. The SN approach, systematically developed by KV02, KV05 and Shao10, replaces the explicit estimation of nuisance covariance matrices with a normalizer constructed from the recursive fluctuations (e.g., partial-sum processes) of the underlying data. While this paradigm has recently been extended to various complex settings, such as structural break testing (Shao10, Hong2024), high-dimensional time series inference (WangShao2020, Wangetal2022), and unconditional forecast evaluation (Li2025), accommodating the information-set-measurable test functions required for conditional predictive ability presents an unresolved methodological challenge. Building upon the principle of self-normalization, we formulate our test by interacting the $\tau$-step-ahead loss differential, $\Delta L_{m,t+\tau}$, with an $\mathscr{F}_t$-measurable test function $h_t$, yielding a combined sequence $Z_{m,t+\tau} = h_t \Delta L_{m,t+\tau}$. The construction of the appropriate normalizer is governed by the dimensionality of the test function. For a scalar test function, we employ the adjusted range of the centered cumulative sum (CUSUM) process of $Z_{m,t+\tau}$. Conversely, for a vector-valued test function, for which extending the adjusted range is mathematically nontrivial due to cross-sectional dependence, we adopt a matrix CUSUM self-normalizer. A primary theoretical advantage of this methodology over traditional HAC procedures is that the limiting behavior of both the sample mean and the self-normalizer involves the same unknown long-run variance in the scalar case, or the same unknown long-run covariance matrix in the vector-valued case. Consequently, when the test statistic is constructed from these two objects, the nuisance variance or covariance parameter cancels out asymptotically. This mechanism yields a pivotal limiting distribution, thereby avoiding subjective bandwidth, kernel, and lag-truncation choices and improving finite-sample reliability.

This paper makes three main contributions to the literature on forecast evaluation. First, we develop a self-normalized conditional predictive ability (SN-CPA) testing framework applicable to both scalar and vector-valued $\mathscr{F}_t$-measurable test functions. By using functionals of the cumulative sum (CUSUM) process of the transformed loss differential, this method circumvents direct long-run covariance estimation and the subjective tuning parameters inherent in conventional HAC procedures. Second, we establish the asymptotic properties of the proposed test statistics. We show that the self-normalizers asymptotically account for the unknown long-run covariance matrix, yielding pivotal limiting null distributions. Furthermore, we prove test consistency against fixed alternatives and derive noncentral limits under local alternatives. Third, Monte Carlo simulations demonstrate that the SN-CPA tests systematically control the finite-sample size distortions commonly associated with HAC-based methods, especially over extended forecast horizons.

The rest of the paper is organized as follows. Section (ref) introduces the multistep CPA framework and the proposed statistics. Section (ref) presents the asymptotic theory. Section (ref) reports the Monte Carlo evidence. Section (ref) concludes. Proofs are available upon request.

Conditional predictive ability framework

This section introduces the multistep CPA framework and explains the choice of normalizers used in the main text.

Forecast comparison setup

Let $W_t=(Y_t,X_t^T)^T$ be the observed process, where $Y_t$ is the variable of interest and $X_t$ is a vector of predictor variables. Let $\mathscr F_t=\sigma(W_1,\ldots,W_t)$ denote the information set available at time $t$. At forecast origin $t$, consider two competing forecasting methods for the future value $Y_{t+\tau}$, where $\tau > 1$ is the forecast horizon. The two forecasts are denoted by $\widehat f_t$ and $\widehat g_t$. They may be generated by parametric, semiparametric, or nonparametric methods, provided that both forecasts are $\mathscr F_t$-measurable.

We use a finite rolling-window estimation scheme for forecast evaluation. Let $m_{f,t}$ and $m_{g,t}$ denote the estimation-window lengths used by the two forecasting methods at forecast origin $t$. These window lengths may be method-specific constants or $\mathscr F_t$-measurable random integers determined by the forecasting methods. Throughout the paper, we assume that they are uniformly bounded by a finite constant $\bar m$, namely, $\sup_t \max \{m_{f,t}, m_{g,t} \} \leq \bar{m} <\infty$. Let $m$ denote the first forecast origin at which both methods can produce a $\tau$-step-ahead forecast. The out-of-sample evaluation period is $t=m,\ldots,T-\tau$, and hence $ n=T-\tau-m+1$. The asymptotics are taken with $m$ fixed and $n\to\infty$, so they are driven by the evaluation sample rather than by an increasing estimation window. To avoid excessive notation, we suppress the explicit dependence of the forecasts on $m_{f,t}$ and $m_{g,t}$.

Let $L(\cdot,\cdot)$ be a prespecified loss function, and define the $\tau$-step-ahead loss differential as $\Delta L_{m,t+\tau} = L(Y_{t+\tau},\widehat f_t) - L(Y_{t+\tau},\widehat g_t)$. The null hypothesis of equal conditional predictive ability (CPA) is formulated as

equation*[equation* omitted — 117 chars of source]

Testing this conditional moment restriction directly against all $\mathscr{F}_t$-measurable alternatives is practically infeasible. Following GW06, we introduce a $q\times 1$ vector of $\mathscr{F}_t$-measurable test functions, $h_t$, to project the conditional hypothesis into a tractable finite-dimensional space. Then, $H_0$ implies the necessary condition $\mathbb{E}[Z_{m,t+\tau}] = \mathbf{0}$, where $Z_{m,t+\tau} = h_t \Delta L_{m,t+\tau} \in \mathbb{R}^q$ is the transformed loss differential. Thus, the CPA framework operationalizes forecast conditional evaluation by testing whether the unconditional mean of $Z_{m,t+\tau}$ is zero.

The specification of $h_t$ is crucial, as it dictates the specific directions in which the test has power. When $h_t \equiv 1$, the projected restriction reduces to a conventional unconditional forecast comparison. Conversely, incorporating lagged loss differences or relevant macroeconomic state variables into $h_t$ directs the test's power toward detecting persistent or state-dependent predictability in relative forecast performance. In this paper, we focus on fixed and low-dimensional test functions. \footnote{While a high-dimensional $h_t$ encompasses a broader information set, it inevitably inflates the dimensionality of the normalizer, introducing estimation noise that erodes finite-sample power. Exploring consistent conditional moment testing via infinite-dimensional spaces (e.g., Bie90; SW98) presents a different statistical trade-off and is beyond the scope of this paper.}

Multistep dependence and choice of normalizers

table[table omitted — 461 chars of source]

For $\tau>1$, overlapping forecast errors generally make $Z_{m,t+\tau}$ serially dependent under the null. Hence the nuisance object is a long-run variance when $q=1$, and a long-run covariance matrix when $q>1$. Standard CPA procedures handle this dependence through HAC estimation, which requires bandwidth, kernel, and lag-truncation choices. The purpose of the proposed normalizers is to avoid directly estimating this long-run variance or covariance matrix.

When $q=1$, $Z_{m,t+\tau}$ is scalar. A conventional studentized statistic would require an estimate of the long-run variance. We instead use an adjusted-range normalizer based on the CUSUM process. This normalizer cancels out the unknown long-run variance without requiring HAC estimation. The use of the adjusted range is also in line with recent range-based self-normalized procedures for time-series testing, such as Hong2024.

When $q>1$, $Z_{m,t+\tau}$ is both vector-valued and serially dependent. The nuisance object is then a full long-run covariance matrix. HAC-based Wald procedures require estimating the precision matrix and are therefore sensitive to tuning choices and finite-sample estimation error. A componentwise range normalizer is not sufficient in this setting, because it does not account for the joint long-run dependence across coordinates. Moreover, the range-based multivariate extensions typically require ordering or decorrelation choices; see Hong2024. We therefore use a matrix CUSUM self-normalizer, following the self-normalization principle of Shao10, without introducing additional tuning-parameter choices. This construction absorbs the unknown long-run covariance matrix and avoids HAC bandwidth, kernel, lag-truncation, and separate scale-estimation choices.

These considerations lead to the two multistep cases summarized in Table (ref).

Self-normalized statistics and asymptotic theory

This section presents the self-normalized statistics and their asymptotic theory for multistep CPA testing. Throughout this section, $B$ denotes a standard one-dimensional Brownian motion and $\mathbb B(r)=B(r)-rB(1)$ denotes the associated Brownian bridge. In the multivariate case, $B^q$ and $\mathbb{B}^q(r)=B^q(r)-rB^q(1)$ denote their $q$-dimensional counterparts.

We consider fixed and local alternatives. The fixed alternative is

equation*[equation* omitted — 180 chars of source]

and we write $V_{m,t+\tau}^{F}=Z_{m,t+\tau}-\mu_F$. The local alternative is

equation*[equation* omitted — 190 chars of source]

and we write $V_{m,t+\tau}^{L}=Z_{m,t+\tau}-n^{-1/2}\mu_L$. When $q=1$, $\mu_F$ and $\mu_L$ are understood as nonzero real constants.

The fixed and local alternatives describe the power behavior of the proposed statistics. They are stronger than the global alternatives in GW06, but have a direct CPA interpretation: $h_t$ contains information that predicts relative forecast performance. A fixed drift makes the statistic diverge, whereas an $n^{-1/2}$-local drift enters the limiting self-normalized Brownian functional as a noncentrality parameter. In both cases, the CUSUM normalizer removes these drifts from the denominator but not from the numerator.

Remark (State-dependent alternatives). The fixed and local alternatives above are stated with constant conditional drifts for notational simplicity. The same conclusions continue to hold under state-dependent conditional drifts, provided that the average drift is nonzero under fixed alternatives and the average local drift coefficient is nonzero under local alternatives, while the centered drift components do not affect the centered CUSUM terms used in the normalizers below.

Let $g_{m,t+\tau}:=\mathbb{E}[Z_{m,t+\tau}\mid \mathscr F_t]$ and $\overline g_{m,n,\tau}:=n^{-1}\sum_{t=m}^{m+n-1}g_{m,t+\tau}$. Consider a fixed alternative under which $\overline g_{m,n,\tau} \overset{p}{\longrightarrow} \mu_F$ for some $\mu_F\in\mathbb R^q$ with $\mu_F\neq \mathbf{0}$. Define $V_{m,t+\tau}^{F}:= Z_{m,t+\tau}-g_{m,t+\tau}$. Assume that the FCLT continues to hold when $Z_{m,t+\tau}$ is replaced by $V_{m,t+\tau}^{F}$. Suppose further that the centered drift component satisfies $\sup_{0\le r\le 1}\left\|n^{-1/2}\sum_{t=m}^{m+\lfloor nr\rfloor-1}\left(g_{m,t+\tau}-\overline g_{m,n,\tau}\right)\right\|=O_p(1)$. Then the centered CUSUM term based on $Z_{m,t+\tau}$ remains of stochastic order $O_p(1)$. Indeed, it can be decomposed into the centered CUSUM term based on $V_{m,t+\tau}^{F}$, which is $O_p(1)$ by the FCLT, and the centered drift component above. On the other hand, $\overline Z_{m,n,\tau}=\overline g_{m,n,\tau}+n^{-1}\sum_{t=m}^{m+n-1}V_{m,t+\tau}^{F}=\mu_F+o_p(1)$. Thus, the fixed-alternative consistency conclusion stated below for constant drifts continues to hold under this state-dependent fixed alternative.

For local alternatives, let $g_{m,t+\tau}=n^{-1/2}a_{m,t+\tau}$ and $\overline a_{m,n,\tau}:=n^{-1}\sum_{t=m}^{m+n-1}a_{m,t+\tau}\overset{p}{\longrightarrow}\mu_L$, for some $\mu_L\in\mathbb R^q$ with $\mu_L\neq\mathbf{0}$. Define $V_{m,t+\tau}^{L}:=Z_{m,t+\tau}-n^{-1/2}a_{m,t+\tau}$. Assume analogously that the FCLT continues to hold with $Z_{m,t+\tau}$ replaced by $V_{m,t+\tau}^{L}$, and that $\sup_{0\le r\le 1}\left\|n^{-1/2}\sum_{t=m}^{m+\lfloor nr\rfloor-1}\left(a_{m,t+\tau}-\overline a_{m,n,\tau}\right)\right\|=O_p(1)$. Then the same decomposition argument shows that the local drift is removed from the centered CUSUM term but remains in the numerator. Hence, the local noncentral limiting conclusions stated below for constant local drifts continue to hold under this state-dependent local alternative, with the noncentrality parameter determined by the limit $\mu_L$.

Scalar multistep case: $q=1$

Define the CUSUM process

equation*[equation* omitted — 170 chars of source]

where $ \overline Z_{m,n,\tau} = n^{-1}\sum_{t=m}^{T-\tau} Z_{m,t+\tau}$ and $\lfloor x \rfloor =\max \{z\leq x \colon z\in\mathbb{Z} \}$. The adjusted range is $R^h_{m,n,\tau}= \sup_{0\leq r\leq 1}T^h_{m,n,\tau}(r)- \inf_{0\leq r\leq 1}T^h_{m,n,\tau}(r)$, and the scalar multistep CPA test statistic is

equation*[equation* omitted — 140 chars of source]

To show the limiting distribution under $H_0$, we impose the following regularity conditions.

assumption(1) The sequence $\{Z_{m, t+\tau} \colon t\geq m \}$ is strongly $\alpha$-mixing with coefficients $\alpha(k)$, and there exists $\eta>0$ such that $\sum_{k=1}^{\infty} \alpha(k)^{\eta/(2+\eta)}<\infty$. (2) For the same $\eta>0$, $ \sup_{t\geq m}\mathbb{E}|Z_{m, t+\tau}|^{2+\eta}<\infty$. (3) There exists $\sigma_{\tau}^2 \in (0,\infty)$ such that, for every $r,s \in[0,1]$, $n^{-1}\mathrm{Cov}\left( \sum_{t=m}^{m+\lfloor nr\rfloor-1} Z_{m, t+\tau}, \sum_{t=m}^{m+\lfloor ns\rfloor-1} Z_{m, t+\tau}\right) \longrightarrow (r\wedge s)\sigma_{\tau}^2$.

Assumption (ref) is a high-level regularity condition imposed directly on the sequence $\{Z_{m,t+\tau}=h_t\Delta L_{m,t+\tau},\ t\geq m\}$, where $m$ is the starting point of the forecast evaluation period. These conditions can be verified from primitive assumptions on the underlying process $W_t$, the forecasting methods, the loss function, and the test function $h_t$, as in GW06; for details, see the Supplementary Material. Moreover, Assumption (ref) is consistent with standard strong-mixing invariance principles for the dependent sequence; see Herrndorf84,Her85 and deJ00. Theorem (ref) below establishes pivotal null limits.

theoremSuppose Assumption (ref) holds. Then, under \(H_0\), \begin{equation*} Q^h_{m,n,\tau} \Rightarrow \frac{B(1)^2} {\left( \sup_{0\leq r\leq 1}\mathbb B(r)-\inf_{0\leq r\leq 1}\mathbb B(r) \right)^2}. \end{equation*}

Under Assumption (ref) and \(H_0\), the partial sum process satisfies

equation[equation omitted — 139 chars of source]

Hence, by the continuous mapping theorem (CMT), $T^h_{m,n,\tau}(\cdot)\Rightarrow \sigma_\tau \mathbb B(\cdot)$ and $R^h_{m,n,\tau}\Rightarrow\sigma_\tau\left(\sup_{0\le r\le 1}\mathbb B(r)-\inf_{0\le r\le 1}\mathbb B(r)\right)$. Since the limiting range is strictly positive almost surely, $\mathbb{P}(R^h_{m,n,\tau}>0)\to 1$. Hence \(Q^h_{m,n,\tau}\) is well defined with probability approaching one.

As demonstrated in Theorem (ref), the unknown long-run standard deviation $\sigma_\tau$ cancels out asymptotically, yielding a pivotal null limit. Unlike the standard HAC-based procedures in GW06, this self-normalized approach bypasses explicit variance estimation and its associated tuning sensitivities. The adjusted range is motivated by recent range-based self-normalized procedures, such as Hong2024, and we adapt it to multistep CPA testing with persistent loss differentials.

To evaluate the power properties of the proposed test, Theorem (ref) below establishes its consistency against fixed alternatives and derives the noncentral limit under local alternatives.

theoremSuppose the statistic $Q^h_{m,n,\tau}$ is defined as above. \begin{enumerate} • Under $H^{(1,\tau),F}_{A,h}$, if Assumption (ref) holds with $Z_{m,t+\tau}$ replaced by $V_{m,t+\tau}^{F}$, then $ Q^h_{m,n,\tau}\xrightarrow{p}\infty$. • Under \(H^{(1,\tau),L}_{A,h}\), if Assumption (ref) holds with $Z_{m,t+\tau}$ replaced by $V_{m,t+\tau}^{L}$, then, with $J_\tau=\mu_L/\sigma_\tau$, \begin{equation*} Q^h_{m,n,\tau}\Rightarrow \frac{(B(1)+J_\tau)^2}{\left(\sup_{0\leq r\leq 1}\mathbb B(r) -\inf_{0\leq r\leq 1}\mathbb B(r) \right)^2}. \end{equation*} \end{enumerate}

Unlike GW06, which focuses mainly on fixed alternatives, Theorem (ref) also derives the noncentral limit under local alternatives and thereby characterizes local asymptotic power.

Multivariate multistep case: $q>1$

In this subsection, we use a matrix CUSUM self-normalizer, following the self-normalization principle of Shao10, rather than range normalization because of the joint long-run dependence across coordinates.

Define

equation*[equation* omitted — 166 chars of source]

where $\overline Z_{m,n,\tau}= n^{-1} \sum_{t=m}^{T-\tau}Z_{m, t+\tau}$. The matrix self-normalizer is

equation*[equation* omitted — 131 chars of source]

The multivariate multistep CPA test statistic is

equation*[equation* omitted — 192 chars of source]
assumption(1) The sequence $\{Z_{m, t+\tau} \colon t\geq m \}$ is strongly $\alpha$-mixing with coefficients $\alpha(k)$, and there exists $\eta>0$ such that $\sum_{k=1}^{\infty} \alpha(k)^{\eta/(2+\eta)}<\infty$. (2) For the same $\eta>0$, $\sup_{t\geq m} \mathbb{E}\|Z_{m,t+\tau}\|^{2+\eta}<\infty$. (3) There exists a positive definite $q\times q$ matrix $\Sigma_\tau$ such that, for every \(r,s\in[0,1]\), $n^{-1} \mathrm{Cov} \left( \sum_{t=m}^{m+\lfloor nr\rfloor-1}Z_{m,t+\tau}, \sum_{t=m}^{m+\lfloor ns\rfloor-1}Z_{m,t+\tau} \right) \longrightarrow (r\wedge s)\Sigma_\tau $.

Assumption (ref) is a high-level regularity condition imposed directly on $Z_{m,t+\tau}$. Since $m$ and $\tau$ are fixed, this process is treated as an ordinary $\mathbb{R}^q$-valued strongly mixing sequence indexed by $t\geq m$. The mixing, moment, and covariance conditions are sufficient to apply the strong-mixing FCLT; see Herrndorf84,Her85 and Davidson94. Theorem (ref) establishes the null limit of $Q^{h,q}_{m,n,\tau}$.

theoremSuppose Assumption (ref) holds. Then, under $H_0$, \begin{equation*} Q^{h,q}_{m,n,\tau}\Rightarrow B^q(1)^T \left( \int_0^1 \mathbb B^q(r)\mathbb B^q(r)^T dr \right)^{-1} B^q(1). \end{equation*}

The multivariate case is analogous to the scalar case, with the range functional replaced by the matrix CUSUM self-normalizer. Under Assumption (ref) and \(H_0\),

equation[equation omitted — 148 chars of source]

Hence, by the CMT and the Riemann sum representation of $U_{m,n,\tau}$, $T^{h,q}_{m,n,\tau}(\cdot)\Rightarrow\Sigma_\tau^{1/2}\mathbb B^q(\cdot)$ and $U_{m,n,\tau}\Rightarrow\Sigma_\tau^{1/2}\left(\int_0^1\mathbb B^q(r)\mathbb B^q(r)^\top\,dr\right)\Sigma_\tau^{1/2}$. Since the limiting matrix is nonsingular almost surely, \(U_{m,n,\tau}\) is nonsingular with probability approaching one, and hence \(Q^{h,q}_{m,n,\tau}\) is well defined with probability approaching one.

Unlike GW06, this pivotal limit in Theorem (ref) is achieved without the explicit estimation of the long-run precision matrix. The matrix CUSUM self-normalizer asymptotically absorbs the nuisance covariance matrix $\Sigma_\tau$, thereby extending the tuning-free principle of Shao10 to multistep forecast comparisons.

Having established the pivotal null limit for the multivariate test, Theorem (ref) characterizes its asymptotic power against both fixed and local alternatives.

theoremSuppose the statistic $Q^{h,q}_{m,n,\tau}$ is defined as above. \begin{enumerate} • Under $H^{(q,\tau),F}_{A,h}$, if Assumption (ref) holds with $Z_{m,t+\tau}$ replaced by $V_{m,t+\tau}^{F}$, then $Q^{h,q}_{m,n,\tau}\xrightarrow{p}\infty $. • Under $H^{(q,\tau),L}_{A,h}$, if Assumption (ref) holds with $Z_{m,t+\tau}$ replaced by $V_{m,t+\tau}^{L}$, then, with $ J_\tau=\Sigma_\tau^{-1/2}\mu_L$, we have \begin{equation*} Q^{h,q}_{m,n,\tau} \Rightarrow (B^q(1)+J_\tau)^T\left(\int_0^1 \mathbb B^q(r)\mathbb B^q(r)^T dr \right)^{-1}(B^q(1)+J_\tau). \end{equation*} \end{enumerate}

The matrix self-normalizer asymptotically absorbs the nuisance covariance $\Sigma_\tau$ under the null hypothesis and implicitly standardizes the local drift under the alternative. Consequently, the noncentrality vector $J_\tau$ represents the local drift standardized by the long-run covariance, while preserving local asymptotic power without requiring explicit HAC estimation.

Monte Carlo simulation

This section reports Monte Carlo simulation results for the finite-sample performance of the proposed self-normalized CPA tests. For notational simplicity in the simulation and empirical sections, we denote the scalar statistic $Q^h_{m,n,\tau}$ as $Q_1$, and the multivariate statistic $Q^{h,q}_{m,n,\tau}$ as $Q_q$ (specifically, $Q_2$ when $q=2$).

Data-generating processes (DGPs)

We consider two DGPs. For each DGP, size and power are evaluated under the same structural framework by varying a single control parameter. Empirical size is computed when this parameter makes the conditional mean of the loss differential zero, and power is computed when it makes the conditional mean nonzero.

DGP 1: Conditional relative performance with overlapping errors. In the spirit of GW06, we generate the predictor \(x_t\), used in the test function \(h_t=(1,x_t)^T\) for the multivariate test and \(h_t=x_t\) for the scalar test, as a persistent AR(1) process: $x_t=\rho x_{t-1}+u_t$, where $u_t\sim \text{i.i.d.}\ N(0,1)$ and $\rho \in \{0.2, 0.5, 0.8\}$ controls the persistence level. The sequence \(x_t\) is standardized to maintain a unit unconditional variance.

The transformed loss differential is generated by $\Delta L_{m,t+\tau}=\delta x_t+\varepsilon_{t+\tau}$, where the forecast error $\varepsilon_{t+\tau}$ is modeled as an $\mathrm{MA}(\tau-1)$ process to explicitly capture the overlap in multistep forecasts: $ \varepsilon_{t+\tau} = c \sum_{j=0}^{\tau-1}\theta_j v_{t+\tau-j}, $ with \(\theta_0=1\) and \(\theta_j=0.5\) for \(j>0\). The scaling constant \(c\) ensures that the unconditional variance of \(\varepsilon_{t+\tau}\) is normalized to one.

When \(\delta=0\), we have $\mathbb E[\Delta L_{m,t+\tau}\mid \mathscr F_t]=0$, so the design evaluates empirical size under the exact null hypothesis. When \(\delta>0\), $ \mathbb E[\Delta L_{m,t+\tau}\mid \mathscr F_t]=\delta x_t$. Importantly, because $\mathbb E[x_t]=0$, the unconditional mean $\mathbb E[\Delta L_{m,t+\tau}]$ remains identically zero regardless of \(\delta\). In this paper, we set $\delta=0.2$. Thus, relative performance is conditionally predictable but not unconditionally biased, thereby isolating the power against conditional alternatives while demonstrating the failure of unconditional benchmark tests.

DGP 2: State-dependent performance. To evaluate the tests' sensitivity to regime shifts and business cycle fluctuations, we consider a state-dependent framework. The state variable $S_t \in \{0,1\}$ is assumed to be in the information set $\mathscr{F}_t$ and is drawn from a Bernoulli distribution with $P(S_t=1)=p$, where the state probability $p \in \{0.2, 0.5, 0.8\}$. The test functions are defined as $h_t=(1,S_t)^T$ for the multivariate test and $h_t=S_t$ for the scalar test.

The transformed loss differential is generated by: $\Delta L_{m,t+\tau} = d(S_t-p) + \varepsilon_{t+\tau}$, where the control parameter $d$ dictates the magnitude of state-dependent predictability, and $\varepsilon_{t+\tau}$ follows the same overlapping $\mathrm{MA}(\tau-1)$ structure as defined above.

When $d=0$, the conditional mean satisfies $\mathbb E[\Delta L_{m,t+\tau}\mid \mathscr F_t]=0$, establishing the exact null hypothesis for empirical size evaluation. When $d>0$, the relative forecast performance varies with the realized state $S_t$, so that the conditional expectation becomes $\mathbb E[\Delta L_{m,t+\tau}\mid \mathscr F_t]=d(S_t-p)$.

Crucially, by construction, the unconditional expectation of the state indicator is $\mathbb E[S_t]=p$. Consequently, the unconditional mean of the loss differential remains identically zero ($\mathbb E[\Delta L_{m,t+\tau}] = 0$) for all values of $d$. This deliberately asymmetric design ensures that while the competing models exhibit strong conditional predictability across different states, they appear equally accurate unconditionally.

Implementation and test statistics

For both DGPs, we consider evaluation sample sizes \(n \in \{50, 100, 150, 200, 250, 300, 350, 400\}\) and perform \(5{,}000\) Monte Carlo replications. Empirical rejection frequencies are evaluated at the 1%, 5%, and 10% nominal significance levels. To capture the conditional information, we employ the scalar test function \(h_t = x_t\) (for DGP 1) and \(h_t = S_t\) (for DGP 2) in the one-dimensional tests. For the multidimensional tests, we naturally expand the information set to include an intercept, utilizing the vector-valued test functions \(h_t = (1, x_t)^T\) and \(h_t = (1, S_t)^T\), respectively.

We systematically compare five test statistics to benchmark the finite-sample performance: The proposed vector-valued SN-CPA statistic (\(Q_{2}\)); the proposed scalar SN-CPA statistic (\(Q_{1}\)); the unconditional self-normalized DM statistic of Li2025 (\(T_{\mathrm{SN}}\)); the traditional HAC-based conditional predictive ability Wald statistic of GW06 (\(T_{\mathrm{GW}}\)); and the traditional HAC-based unconditional Diebold-Mariano statistic of DM95 (\(T_{\mathrm{DM}}\)).

Note that all HAC covariance matrices are estimated using the NeweyWest function from the sandwich package in R, maintaining all default settings. Because the asymptotic null distributions of our proposed SN-CPA statistics (\(Q_{1}\) and \(Q_{2}\)) are non-standard functionals of Brownian motions and Brownian bridges, we simulate their critical values by Monte Carlo methods. Following the theoretical representations in Section (ref), we approximate the continuous-time Wiener processes using discrete Gaussian random walks on a fine grid with \(N=200,000\) steps and $M = 10{,}000$ independent replications. The critical values are reported in Table (ref). Moreover, the critical values for the benchmark statistics (\(T_{\mathrm{SN}}\), \(T_{\mathrm{GW}}\), and \(T_{\mathrm{DM}}\)) are drawn from their well-established asymptotic distributions or existing literature.

table[table omitted — 1,340 chars of source]

Empirical size and power

Tables (ref), (ref), and Figure (ref) report the empirical size and power ($\delta=0.2$) for DGP 1. The results for DGP 2 yield highly similar conclusions; to conserve space, we focus our discussion on DGP 1 in the main text and relegate the comprehensive tables and figures for DGP 2 to the Supplementary Material. By comparing the proposed statistics with traditional benchmarks, we uncover the following key empirical findings:

First, traditional HAC-based tests ($T_{\mathrm{GW}}$ and $T_{\mathrm{DM}}$) tend to over-reject the null hypothesis in finite samples, particularly as the forecast horizon $\tau$ and data persistence $\rho$ increase. In contrast, the proposed statistic $Q_{2}$ effectively mitigates these finite-sample size distortions, maintaining accurate size control near the nominal level across all considered sample sizes and horizons. Second, the unconditional tests ($T_{\mathrm{SN}}$ and $T_{\mathrm{DM}}$) display flat power curves hovering around 5%, which naturally aligns with their theoretical design focusing on average performance rather than dynamic, information-set-based predictability. Meanwhile, the proposed $Q_{1}$ and $Q_{2}$ tests exhibit robust, monotonically increasing power that converges to one as the sample size $n$ grows. Third, although the HAC-based $T_{\mathrm{GW}}$ test occasionally reports numerically higher raw power than $Q_{2}$, this apparent advantage should be interpreted with caution, as it is partially driven by its empirical size inflation under the null. When accounting for size control, $Q_{2}$ provides a more balanced and reliable measure of true statistical sensitivity. Finally, while $Q_{1}$ can detect specific directional predictability, it is empirically less stable than its multidimensional counterpart. By augmenting the test function with an intercept ($h_t = (1, x_t)^T$), $Q_{2}$ preserves the full joint conditional moment structure. This prevents information loss, yielding consistently tighter size control and a more comprehensive evaluation of predictive ability.

These empirical findings align with our established asymptotic theory. The finite-sample challenges associated with HAC methods illustrate the inherent difficulty of explicitly estimating the long-run covariance matrix induced by overlapping $MA(\tau-1)$ forecast errors. Conversely, the accurate size control and robust power of $Q_{2}$ validate our primary theoretical claim: the matrix CUSUM self-normalizer absorbs the complex serial and cross-sectional dependence. By entirely bypassing ad hoc bandwidth selections, our proposed methodology offers a reliable and tuning-free alternative for finite-sample forecast evaluation.

table[table omitted — 3,707 chars of source]
table[table omitted — 3,707 chars of source]
center[center omitted — 537 chars of source]

Conclusion

In this paper, we propose a unified self-normalized framework for testing multistep conditional predictive ability. By employing functionals of the CUSUM process of the transformed loss differential, our approach avoids direct long-run covariance estimation and the associated ad hoc bandwidth selections required by conventional HAC methods. We establish asymptotic theory for both scalar and vector-valued test functions, deriving pivotal limiting null distributions, proving consistency against fixed alternatives, and obtaining noncentral limits under local alternatives. Monte Carlo simulations show that the proposed SN-CPA tests substantially alleviate the pronounced, horizon-dependent size distortions that affect traditional HAC-based methods, while maintaining robust empirical power.

\makeatletter \@setaddresses \global\let\addresses\@empty \makeatother