Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,979 characters · 8 sections · 46 citation commands
On the Size Control of the Hybrid Test for Predictive Ability
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} superior predictive ability test; hybrid test; asymptotic validity
\spacingset{1.45}
A test of superior predictive ability (SPA) compares many forecasting methods. More precisely, it tests whether a certain forecasting method outperforms a finite set of alternative forecasting methods. Most notably in the literature of tests of SPA, white2000 developed a framework for a SPA test and proposed a SPA test called the Reality Check for data snooping. hansen2005test proposed a SPA test featuring improved power in the framework of white2000. Finally, song2012testing devised a SPA test, called the hybrid test, which delivers better power against certain local alternative hypotheses under which both of the SPA tests of white2000 and hansen2005test perform poorly.
One of the main challenges of all SPA tests lies in finding a suitable critical value. This is because popular test statistics used in this setting are asymptotically non-pivotal and depend on parameters that cannot be consistently estimated. More concretely, the null hypothesis of a SPA test can be written as $H_0: \mu \leq 0_M$ for a parameter $\mu \in \mathbb{R}^M$ where $0_M$ is a $M$-dimensional vector with zeros, inequality applies elementwise, and $M\geq 2$ in general. It then follows that the limiting distribution of standard test statistics depends on exactly which of the elements in the vector $\mu$ are equal to zero. This property prevents researchers from using any tabulated critical values.
To circumvent the above problem, white2000 proposed to use a critical value from the so-called least favorable distribution in his Reality Check test. The approach exploits the fact that the distribution of its test statistic $T$ under $\mu=0_M$ is stochastically largest over all possible null distributions satisfying $\mu\leq 0_M$. The distribution under $\mu=0_M$ is then called the least favorable case. The Reality Check approximates the least favorable distribution using the bootstrap and takes the $1-\alpha$ quantile of the distribution as the critical value, where $\alpha$ is a significance level. The resulting critical value converges to a value which is always larger than or equal to $1-\alpha$ quantile of the limiting distribution of $T$ under any null distribution and thus the approach yields a test with correct asymptotic size.
song2012testing followed the Reality Check in the construction of the hybrid test and used the same least favorable distribution, i.e., the one associated with $\mu=0_M$. However, in this article, we show that this null distribution is not the least favorable one for the type of test statistic of the hybrid test which, in particular, combines two different test statistics. Whereas one of the test statistics is stochastically largest under $\mu=0_M$, the other one is not. Consequently, the hybrid test which employs bootstrap approximations to the distribution with $\mu=0_M$ fails to control the rejection probability under the null and loses pointwise asymptotic validity.
As the main contribution of this article, we show that the hybrid test may not be pointwise asymptotically of level $\alpha$ under reasonable conditions. This implies that a researcher could reject the null hypothesis $H_0$ with probability higher than the significance level $\alpha$ in the limit even when $H_0$ is true. A simple example in \hyperref[subsec:example]{Section 3.1} shows that the rejection probability under the null distribution could be over $11\%$ when the significance level $\alpha$ is set at $5\%$. Our results illustrate that the cause of the problem lies behind the fact that the bootstrap procedure that the hybrid test uses approximates neither the asymptotic distribution of the test statistics nor their least favorable distribution. As an easy fix, we propose a modified hybrid test which is pointwise asymptotically of level $\alpha$, again under reasonable conditions. Our proposed modification follows the generalized moment selection method by andrews2010inference after accounting for the fact one of the test statistics in the hybrid test does not exhibit certain monotonicity properties that are required for the generalized moment selection method.
Neither devising an alternative test to the hybrid test nor analyzing theoretical properties of the modified hybrid test is the main focus of this article. Yet, it is worth making a remark on asymptotic validity in SPA tests. When it comes to controlling the asymptotic type I error, uniform asymptotic validity, which is a necessary condition of pointwise asymptotic validity, is regarded as the gold-standard as it guarantees uniform control of Type I error over all null distributions at a fixed sample size. For such a reason, the literature of testing moment inequalities has addressed the importance of uniform asymptotic validity. However, the property has not been yet discussed in depth in the literature of SPA tests, of which the reason we attribute to complexity coming from dealing with dependent data.\footnote{Not only the papers discussed in this article, white2000 and hansen2005test, but also a recent paper on conditional SPA test by LiEtAl2021REStud deals with pointwise asymptotic validity, not uniform asymptotic validity. For readers who are interested in uniform asymptotic validity in the context of testing moment inequality, see canayshaikh.} As a first step of engaging uniform asymptotic validity in this literature, we formally define uniform asymptotic validity in the context of SPA tests following andrews2010inference and show that the modified hybrid test is not uniformly asymptotically valid in \hyperref[sec:appendixC]{Appendix C}. Developing a general framework for uniform asymptotic validity for SPA tests could be interesting, yet we leave it as a topic for future research.
This article is organized as follows. \hyperref[sec:setup]{Section 2} lays out notation and describes the hybrid test as originally proposed by Song (2012). \hyperref[sec:body]{Section 3} presents the main result of the properties of the hybrid test and proposes the modified hybrid test with its formal properties. \hyperref[sec:montecarlo]{Section 4} explores the Monte Carlo simulations of the hybrid test and the modified one. Lastly, \hyperref[sec:conclusion]{Section 5} concludes. The proofs of the formal results are included in \hyperref[sec:appendixB]{Appendix B}.
In this section, we introduce the hybrid test proposed by song2012testing in a simple setup as in hansen2005test where the data generating process is stationary. The framework in song2012testing incorporates more general settings, yet a simple setup better serves our purpose by allowing us to solely focus on verification of the theoretical property of the hybrid test. We explain how our setup is different from the original one in more detail in \hyperref[sec:body]{Section 3}.
Consider a situation where we aim to predict a $\tau$-ahead unknown random variable $\xi_{t+\tau}$ at time $t$. Suppose that we have $M+1$ different forecasts: a benchmark $\varphi_{t,0}$ and a finite set of alternative forecasts $\varphi_{t,m}$, $m\in\mathbf{M} \equiv \{1, \ldots, M\}$. The objective of the hybrid test is to test whether the benchmark forecast is superior to all other alternative forecasts in terms of predictive ability. To compare the predictive ability, we assess the risk (of prediction) of the $m$th forecast in terms of expected risk, $E[\Lambda(\xi_{t+\tau}, \varphi_{t,m})]$ for $m=0, 1, \ldots, M$. An example of such risk is the squared error $\Lambda(\xi_{t+\tau}, \varphi_{t,m})= \|\xi_{t+\tau}-\varphi_{t,m}\|^2$ for $ \xi_{t+\tau}, \varphi_{t,m}\in \mathbb{R}$. For more examples, refer to song2012testing or references therein.
If the expected risk of the forecast $\varphi_{t,m}$ is greater than or equal to that of the benchmark $\varphi_{t,0}$, we say the benchmark forecast $\varphi_{t,0}$ dominates the forecast $\varphi_{t,m}$ (in terms of predictive ability). Let us define the relative risk variable between the benchmark forecast $\varphi_{t,0}$ and the $m$th alternative forecast $\varphi_{t,m}$ as
Then the null hypothesis that the benchmark forecast $\varphi_{t,0}$ dominates all alternative forecasts in $\mathbf{M}$ can be formulated as
The relative risk variables are building blocks for our analysis. We regard $\{d_t: t= 1, \ldots, n\}$ where $d_t =(d_{t,1}, \ldots, d_{t,M})^t \in \mathbb{R}^M$ as our observations and abstract away from how they are constructed. We assume $\{d_t: t= 1, \ldots, n\}$ are stationary under distribution $F$ with mean $\mu\equiv E[d_t]$ and thus expectation in (ref) is taken with respect to $F.$
Before proceeding, it is useful to define what we mean by pointwise asymptotically valid tests. A test $\phi_n=\phi_n(d_1, \ldots, d_n)$ for the null hypothesis $H_0$ is said to be pointwise asymptotically of level $\alpha$ if it satisfies
for a data generating process satisfying the null hypothesis $H_0$ and the maintained assumptions. Interchangeably, we say $\phi_n$ attains pointwise asymptotic validity. If a test fails to satisfy condition (ref), then we can always find some distribution $F$ under $H_0$ along which the rejection probability $E[\phi_n]$ exceeds the significance level $\alpha$ infinitely often as the sample size $n$ grows.
An unusual feature of the hybrid test is that the test uses two pairs of a test statistic and a critical value. In \hyperref[sec:body]{Section 3}, we show how this feature complicates theoretical analysis of pointwise asymptotic validity. The first pair $(\hat{T}^r_n, \hat{c}^{r*}_n)$ is adopted from the Reality Check which tests the one-sided null hypothesis (ref). We call $\hat{T}^r_n$ and $\hat{c}^{r*}_n$ one-sided test statistic and one-sided critical value, respectively. The second pair $(\hat{T}^s_n, \hat{c}^{s*}_n)$ is adopted from the symmetrized test by lmw2005. We call $\hat{T}^s_n$ and $\hat{c}^{s*}_n$ two-sided test statistic and two-sided critical value. The name, two-sided test, comes from the fact that the two-sided statistic $\hat{T}^s_n$ is originally proposed to test a two-sided null hypothesis $H^s_0: \mu \leq 0_M \text{ or } \mu \geq 0_M$.
The test statistics and critical values are defined through two functions, $T^r: \mathbb{R}^M \mapsto \mathbb{R}$ and $T^s: \mathbb{R}^M \mapsto \mathbb{R}$ defined as
and through two statistics
where $\hat{\sigma}^2_{n,m}, m=1, \ldots,M$ are some estimators for the asymptotic variance of $\sqrt{n}(\hat{d}_n - \mu).$ For now, we assume that these statistics satisfy
for all $m=1,\ldots,M$ where $\Sigma$ is a $M\times M$ semi-positive matrix and $\Sigma_{i,j}$ is the $(i,j)$ component of $\Sigma$. Note $\hat{D}_n$ is an estimator of $\text{diag}(\Sigma)$. The two test statistics of the hybrid test are
Defining critical values is not as straightforward as defining the test statistics. Define two values $\bar{c}^r(\alpha, \gamma)$ and $\bar{c}^s(\alpha, \gamma)$ satisfying these two conditions:
for a significance level $\alpha \in (0,1)$ and a fixed tuning parameter $\gamma \in (0,1)$. If the critical values $\hat{c}^{r*}_n $ and $\hat{c}^{s*}_n$ were consistent for $\bar{c}^r(\alpha, \gamma)$ and $\bar{c}^s(\alpha, \gamma)$, the hybrid test would be asymptotically pointwise of level $\alpha$. The two values are, however, tricky to estimate. To see this, let us rewrite the test statistics as
Because $T^q$ is a continuous mapping for any $q\{r,s\}$, the weak convergence of $\hat{T}^q_n$ follows from the weak convergence of the argument. The first term in the argument $\sqrt{n}\hat{D}^{-1/2}_{n}(\hat{d}_n-\mu)$ is stochastically bounded by (ref), yet we cannot consistently estimate the second term $\sqrt{n}\hat{D}^{-1/2}_{n}\mu$ unless $\mu=0_M$ because at least one component of $\mu$ diverges.
Instead of directly dealing with the tricky part, the Reality Check takes an indirect approach based on the least favorable case to define a critical value. To explain, it is useful to define a distribution function
for any $x, \eta \in \mathbb{R} $ for $q \in \{r,s\}$. Note $J^q_n(x;\mu)$ is simply the distribution function of the test statistic $\hat{T}^q_n$ for $q \in \{r,s\}$. The Reality Check then takes advantage of the fact that $T^r(\hat{D}^{-1/2}_n \sqrt{n} (\hat{d}_n - \mu))$ can be consistently estimated by conventional data-dependent methods and it stochastically dominates the one-sided test statistic $\hat{T}^r_n$ (at first order) under $H_0$, i.e.,
which immediately follows from monotonicity of $T^r(\cdot).$ Equivalently, any quantile of $T^r(\hat{D}^{-1/2}_n$ $\sqrt{n} (\hat{d}_n - \mu))$ is greater than or equal to that of $\hat{T}^r_n$ under $H_0$, i.e.,
where $ (J^q_n)^{-1} (t; \eta) = \inf\{x \in \mathbb{R}: J^q_n(x;\eta) \geq 1-t\}$ for any $t \in [0,1]$ and for $q\in \{r,s\}$. In this sense, the distribution of $T^r(\hat{D}^{-1/2}_n \sqrt{n} (\hat{d}_n - \mu))$, or the distribution of $\hat{T}^r_n$ at $\mu=0_M$ is called the least favorable case of $\hat{T}^r_n$ under $H_0$. The Reality Check defines its critical value $\underbar{c}^r$ as the $1-\alpha$ quantile of the limit distribution of the least favorable case, i.e., $\lim_{n\to \infty} J^r_n (\cdot; 0_M) $. As intended, the stochastic dominance then guarantees the test $1\{ \hat{T}^r_n >\underbar{c}^r \}$ being asymptotically of level $\alpha$.
Following the Reality Check, the hybrid test defines the bootstrap-based critical value $\hat{c}^{q*}_n$ as a consistent estimator of the $1-\alpha$ quantile of $\lim_{n\to \infty} J^r_n (\cdot; 0_M) $ for $q \in \{r,s\}$. Here we explain the procedure step by step. Consider a bootstrap sample $\{\hat{d}_{n,b}^{*}: 1\leq b \leq B\}$ obtained by the stationary bootstrap by RomanoPolitis1994JASA where we denote the $m$th element of $\hat{d}_{n,b}^{*}$ as $\hat{d}_{n,b,m}^{*}$. Define a centred bootstrap sample as
Define the bootstrap test statistics as $\{(\hat{T}_{n,b}^{r *}, \hat{T}_{n,b}^{s *})\}_{b=1}^{B}$ where $\hat{T}^{q*}_{n,b}$ is
and $\{\hat{\sigma}_{n,m}: m \in \mathbf{M} \}$ are not bootstrapped. Choose $\gamma \in (0,1]$. The hybrid test defines the two-sided critical value $\hat{c}^{s*}_n$ as the $1-\alpha\gamma$ quantile of the bootstrap sample $\{\hat{T}^{s*}_{n,b}\}^B_{b=1}$, i.e.
Given $\hat{c}^{s*}_n $, the one-sided critical value $\hat{c}^{r*}_n $ is defined as the $1-\alpha(1-\gamma)$ quantile of the bootstrap sample $\{\hat{T}^{r*}_{n,b} \cdot 1\{\hat{T}^{s*}_{n,b} \leq \hat{c}^{s*}_{n} \}\}^B_{b=1}$ for $\gamma \in (0,1)$, i.e.
and $\hat{c}^{r*}_n= \infty$ for $\gamma=1.$
Finally, given the two pairs, $(T^r_n, \hat{c}^{r*}_n)$ and $(T^s_n, \hat{c}^{s*}_n)$, the hybrid test rejects the null hypothesis if $T^r_n>\hat{c}^{r*}_n$ or $T^s_n> \hat{c}^{s*}_n$. For brevity, define $\phi^r_n \equiv 1\{T^r_n>\hat{c}^{r*}_n\}$ and $\phi^s_n\equiv 1\{T^s_n> \hat{c}^{s*}_n\}$, say one-sided test and two-sided test. Then the hybrid test is defined as
That is, the rejection of the hybrid test is the union of the two rejection regions by $\phi^r_n$ and $\phi^s_n.$ The tuning parameter $\gamma \in (0,1]$ determines how much the two-sided test $\phi^s_n$ contributes to forming the rejection region as opposed to the one-sided test $\phi^r_n.$ If $\gamma $ is zero, the hybrid test coincides with the one-sided test $\phi^r_n$. If $\gamma $ is 1, then the hybrid test corresponds to the two-sided test $\phi^s_n$. We restrict the tuning parameter $\gamma$ to be in $(0,1]$ as the asymptotic properties of the one-sided test can be found in white2000.
In this section, we investigate asymptotic properties of the hybrid test under the null hypothesis, which are not formally discussed in song2012testing. First, we provide a simple example where the rejection probability of the hybrid test exceeds significance level $\alpha$ in the limit. Then we present the main results generalizing the observation made in the example.
The following set of assumptions strengthens the setup of song2012testing and facilitates our theoretical analysis of the hybrid test.
Assumption (ref) implies that the marginal distribution of $d_t$ does not vary over time and has a finite mean. This ensures that null hypothesis $H_0$ in (ref) is well-defined. Stationarity of data generating processes is often made in the literature so as to invoke asymptotic normality and bootstrap consistency (for example, Assumption 1 in hansen2005test and Assumption A in white2000). The asymptotic normality in Assumption (ref) is necessary in order to enable the inference of the test statistics. song2012testing requires asymptotic normality of a generic statistic $\tilde{d}_n$ in place of $\hat{d}_n$, yet we consider the case where $\tilde{d}_n$ is given as a sample mean $\hat{d}_n$ as in white2000 and hansen2005test. The bootstrap consistency in Assumption (ref) guarantees that the bootstrap approximates the distribution of $\sqrt{n}(\hat{d}_n-\mu)$ for large $n$. This is necessary to justify that the hybrid test uses the bootstrap critical values. Assumption (ref) and (ref) can be attained by imposing $\alpha$-mixing condition onto the process $\{d_t\}^n_{t=1}$ as in hansen2005test, yet we maintain high-level assumptions building on the literature. Assumption (ref) guarantees that no element of $\sqrt{n}(\hat{d}_n-\mu)$ degenerates in the limit, which is made to innocuously simplify the analysis. Finally, another high-level Assumption (ref) assumes consistency of the variance estimator, which is implicitly assumed in song2012testing in his use of studentized statistics.
\color{black}
To gain intuitions on the asymptotic properties of the hybrid test, we consider a simple example where the number of alternative forecasts is two, $M=2$. Consider the distribution function $F$ in Assumption (ref) such that $\mu_{1}=0$ and $\mu_{2}<0$. Namely, the first alternative forecast is as risky as the benchmark, whereas the benchmark dominates the second alternative forecast. We further assume that the covariance matrix in Assumption (ref) is the identify matrix, $\Sigma=I_2$, and it is known. Thus we simply use $\hat{D}_n = I_2$. For simplicity, let $\gamma=0.5$.
First, we derive the asymptotic distribution of the test statistics, $\hat{T}^{r}_{n}$ and $\hat{T}^{s}_{n}$. By Assumption (ref), we have
This and the condition that $\mu_{1}=0, \mu_{2}<0$ together imply that $\sqrt{n}\hat{d}_{n,2}$ diverges to $-\infty$ as $n\to\infty$ while $\sqrt{n}\hat{d}_{n,1}$ is stochastically bounded. Then the two test statistics depend only on $\sqrt{n}\hat{d}_{n,1}$ for large $n$, eventually yielding the following approximation:
Meanwhile, $(\hat{T}^{r*}_{n,b}, \hat{T}^{s*}_{n,b})$ weakly converges to a distribution different from $(Z_1, Z_1)$ in (ref). According to Assumption (ref), we have
with probability approaching 1, and then continuous mapping theorem gives
with probability approaching 1 where $V\equiv (V_1, V_2) \sim N(0_2, I_2)$.
Our salient finding comes from a comparison between the two asymptotic distributions of $\hat{T}^s_n$ and $\hat{T}^{s*}_{n,b}$: the marginal asymptotic distribution of $\hat{T}^{s*}_{n,b}$ does not stochastically dominate that of $\hat{T}^s_n$. In fact, the two distribution functions of $Z_1$ and $T^s(V)$, which are the limit distributions of $\hat{T}^s_n$ and $\hat{T}^{s*}_{n,b}$, cross each other. This results in
where $\Phi$ and $J^s$ are the distribution functions of $Z_1$ and $T^s(V)$. This rather unexpected result suggests that the distribution of $\hat{T}^s_n$ at $\mu=0_2$ that the bootstrap test statistic $\hat{T}^{s*}_{n,b}$ approximates is not the least favorable case of $\hat{T}^s_n$ under the null hypothesis $\mu \leq 0_2$ for large $n$. This contrasts to the fact that the distribution $\hat{T}^r_n$ at $\mu=0_2$ is the least favorable case of $\hat{T}^r_n$ under $\mu \leq 0_2$ for any $n$ in the sense of (ref).
The finding implies that the hybrid test may not attain pointwise asymptotic validity. Simple algebra provides a formula for the distribution function: $J^s(x) =-2\Phi^2(x) + 4\Phi(x) -1$ for $x \in [0, \infty)$. It is easy to show that the two-sided critical value $\hat{c}^{s*}_n$ is consistent for the $1-\alpha/2$ quantile of $J^s$, i.e.
Then (ref) and (ref) combine to give
where $c^r(\alpha/2)$ is the probability limit of $\hat{c}^{r*}_n$ and the inequality holds by the definition of minimum. The last term $1-\Phi(c^s(\alpha/2))$ exceeds the significance level $\alpha$ if $\alpha\in (0, 0.25)$. The gap between $E[\phi_n]$ and $\alpha$, the violation of pointwise asymptotic validity of the hybrid test, could be substantive: $1-\Phi(c^s(\alpha/2))$ is 0.158, 0.112, and 0.05 for different values of $\alpha$, 0.10, 0.05, and 0.01 respectively.
The example in the previous subsection shows that, given all the assumptions are maintained, there exist distributions $F$ under which the distribution of $\hat{T}^s_n$ at $\mu=0_M$ is not the least favorable case of $\hat{T}^s_n$ under the null hypothesis in (ref) for large $n$ in the sense of (ref). This phenomenon, in fact, holds in general and it leads to the hybrid test not being pointwise asymptotically valid. We start with introducing Lemma (ref), of which the proof is delegated to \hyperref[sec:appendixB]{Appendix B}.
Generalizing the result in (ref), this lemma says that under the distribution $F$ and for sufficiently large $n$
where $J^s_n$ is defined in (ref). In other words, the distribution of $\hat{T}^s_n$ at $\mu=0_M$ is not the least favorable case of $\hat{T}^s_n$ under $\mu \leq 0_M$ for large $n$ in the sense of (ref).
The implication that immediately follows from Lemma (ref) is that the two-sided test $\phi^s_n$ alone fails to control the size. If the tuning parameter $\gamma \in (0,1]$ is $1$, then the hybrid test coincides with the two-sided test, i.e $\phi_n= \phi^s_n$, and so the hybrid test is not pointwise asymptotically of level $\alpha.$ Rather unexpectedly, the following theorem tells us that the over-rejection of the null hypothesis driven by the two-sided test is overriding even for small $\gamma$ and consequently the hybrid test is not pointwise asymptotically valid for any $\gamma \in (0,1]$ for some $\alpha$.
Before proceeding, we define the parameter space $\mathcal{F}$ as
and define the subset of $\mathcal{F}$ which satisfies the null hypothesis as $\mathcal{F}_0 \equiv \{(\Sigma, F) \in \mathcal{F}: E[d_t] \leq 0_M\}$.
Theorem (ref) claims that the hybrid test is not pointwise asymptotically of level $\alpha$ for any $\alpha \in (0, \bar{\alpha})$ for some $\bar{\alpha}$ under the parameterization $\mathcal{F}$. More specifically, it provides three sufficient conditions of a data generating process $(\Sigma, F)$ under which the asymptotic rejection probability exceeds the significance level $\alpha \in (0,\bar{\alpha})$. This has a substantial implication in practice: one may reject the null hypothesis at a higher rate than the desired level $\alpha$ even for large $n$ when the null hypothesis holds true. This result further implies that any power gain that the hybrid test is reported to possess over other SPA tests may be due to over-rejection of the hybrid test.
The magnitude of the upper bound $\bar{\alpha}$ could be of practical interest as one can carry out the hybrid test without taking the risk of committing the type I error over the conventional significance levels if $\bar{\alpha}$ is smaller than 0.01. The value of $\bar{\alpha}$ is, however, a priori unknown as it relies on $M$ as well as the number of the alternative forecasts which attain the same risk as the benchmark, i.e. $M_0 \equiv |\{m \in \mathbf{M}: \mu_m=0\}|$. In fact, once $M$, $M_0$ and $\gamma$ are fixed, $\bar{\alpha}$ can be obtained by numerical approximation. We tabulated some values of $\bar{\alpha}$ under $\gamma=0.5$ in Table (ref) to see how large $\bar{\alpha}$ could be. The numbers in Table (ref) reveal that the value of $\bar{\alpha}$ varies systemically as the ratio of $M_0$ to $M$ varies. $\bar{\alpha}$ approaches to $\gamma$ as the ratio increases to 1 and approaches to $0$ as the ratio diminishes to zero. We present the values of $\bar{\alpha}$ under $\gamma=0.25$ and $\gamma=0.75$ in \hyperref[sec:appendixA]{Appendix A}. The result implies that one cannot use conventional significance levels $\{0.01, 0.05, 0.1\}$ when the ratio $M_0/M$ exceeds 0.5.
Among the three conditions postulated in Theorem (ref), the first condition in Theorem (ref) is crucial because it prevents the distribution of the test statistics from degenerating. The condition is satisfied if the set of alternative forecasts $\mathbf{M}$ contains at least one forecast that attains the same risk as the benchmark. This condition is violated if all the alternative forecasts in $\mathbf{M}$ are strictly riskier than the benchmark forecast. In this case, both test statistics diverge to $-\infty$ while the critical values converge to some fixed numbers so the rejection probability converges to zero. Therefore, if the first condition is violated, the conclusion no longer holds.
The second condition states that the set $\mathbf{M}$ contains at least one forecast that is riskier than the benchmark. Recall that in the example from the previous subsection $\mu_{2}<0$ plays the key role drawing the conclusion by having the limiting distribution of $\hat{T}^s_n$ deviate from that of $\hat{T}^{s*}_{n,b}$. In the same manner, the second condition in the theorem causes the asymptotic distribution of the test statistics to differ from that of the bootstrap test statistics. If the second condition is not satisfied, then $\mu$ is zero under the null hypothesis so the bootstrap test statistics correctly approximate the limiting distribution of the test statistics. The probability to reject the null hypothesis, therefore, converges to the significance level $\alpha$, rather than exceeding $\alpha$.
Unlike the first two, the last condition is not a necessary condition. It requires that the covariance of $(Z_i, Z_j)$ is zero for any $i\neq j \in \mathbf{M}$ where $Z $ is the random vector from $N(0_M, \Sigma)$ in Assumption (ref). In a simple case where $M=2$, we can easily show that the result still holds even if $\text{cov}(Z_1, Z_2) > 0.$ The condition is posited to simplify the proof.
As our foremost finding of this article, Theorem (ref) shows that the hybrid test is not pointwise asymptotically of level $\alpha$ for some $\alpha$. Lemma (ref) identifies the root of this problem. Namely, the distribution from which the limit of the two-sided critical value $\hat{c}^{s*}_n$ is obtained does not stochastically dominate the asymptotic distribution of the two-sided test statistic $\hat{T}^s_n$, i.e., $\lim_{n\to\infty}J^s_n(\cdot, \mu)$ where $J^s_n$ is defined in (ref). This problem can be solved if we approximate the asymptotic distribution of $\hat{T}^s_n$ and obtain a critical value from it. The moment selection technique provides a simple way to do so.
The moment selection technique was originally proposed by hansen2005test. The purpose was to improve the power of the Reality Check which exploits the least favorable case, because the Reality Check tends to perform conservatively by picking a rather large critical value in the sense of (ref). andrews2010inference, canay2010, and bugni2010 independently developed similar techniques in the context of testing moment inequality. See canayshaikh for more details.
Our interest does not lie in improving the power of the hybrid test. Nonetheless, the moment selection technique can serve to rectify the problem that we state in Theorem (ref). Below we explain how we can apply the generalized moment selection method by andrews2010inference to the hybrid test.
First, we normalize test statistics so that their values are zero under the null hypothesis. Specifically, define two modified test statistics $\tilde{T}^r_n$ and $\tilde{T}^s_n$ as
where $S^q: \mathbb{R}^M \to \mathbb{R}$ for $q \in \{r,s\}$ are real-valued functions such that $S^r (x) = \max_{m \in\mathbf{M}} (x \vee 0)$ and $S^s (x) = \min(\max_{m \in\mathbf{M}} (x \vee 0), \max_{m \in\mathbf{M}} ((-x) \vee 0))$. The operation $a \vee b$ denotes the maximum between $a$ and $b$.
Second, we define the moment selecting vector $\hat{\psi}_n = (\hat{\psi}_ {n,1}, \ldots, \hat{\psi}_ {n,M} )^t$ where its $m$th element is
$\kappa_n$ is a non-stochastic sequence of non-negative numbers such that $\kappa_n \to \infty$ and $\kappa_n/ \sqrt{n} \rightarrow 0 $ as $n \to \infty$. $\kappa_n$ is a tuning parameter that a researcher has to choose. andrews2010inference recommend $\kappa_n = \sqrt{\log n}.$
We suggest two types of data-dependent critical values. The first type is bootstrap-based. We define the critical values as follows:
where
$P^*_n$ is the bootstrap probability conditional on the sample $\{d_t\}^n_{t=1}$ and $B$ is some large number. The second type is simulation-based. We define critical values $\tilde{c}^{q}_n (1-\alpha)$ for $q \in \{r,s\}$ as
where $\hat{\Omega}_n \equiv \hat{D}^{-\frac{1}{2}}_n \hat{\Sigma}_n \hat{D}^{-\frac{1}{2}}_n$, $\hat{\Sigma}_n$ is an estimator of $\Sigma$ in Assumption (ref), and $\hat{\Omega}_n^{1/2}$ is a symmetric positive semi-definite matrix such that $\hat{\Omega}^{1/2}_n \hat{\Omega}^{1/2}_n=\hat{\Omega}_n$. $P^{\#}$ a probability measure conditioned on $( \hat{\Omega}^{\frac{1}{2}}_n, \hat{\psi}_n)$; $Z^{\#}$ follows the standard normal distribution and is independent from the sample $\{d_t\}^n_{t=1}$. Given $( \hat{\Omega}^{\frac{1}{2}}_n, \hat{\psi}_n)$, we can obtain $\tilde{c}^q_n(1-\alpha)$ by simulating $\{Z^{\#}_1, \ldots, Z^{\#}_R\}$ for some large $R$.
The biggest difference of the modified bootstrap test statistics in (ref) from the original bootstrap test statistics in (ref) is that the moment selecting vector $\hat{\psi}_n$ defined in (ref) is added to the centred bootstrap statistics $\tilde{d}^*_{n,b}$. The bootstrap statistic $\tilde{d}^*_{n,b} $ hinders the bootstrap test statistics from approximating asymptotic distribution of the test statistics when $\mu$ is not zero. To be specific, if $\mu_m <0$ for some $m \in \mathbf{M}$, $\sqrt{n} \hat{d}_{n,m} /\hat{\sigma}_{n,m}$ diverges to $-\infty$ whereas $\sqrt{n}\tilde{d}^*_{n,b,m} /\hat{\sigma}_{n,m}$ remains stochastically bounded. The moment selecting vector aids the approximation by adding a quantity diverging to $-\infty$ to $\sqrt{n}\tilde{d}^*_{n,b,m}/\hat{\sigma}_{n,m}$ when $\sqrt{n}\hat{d}_{n,m}/\hat{\sigma}_{n,m}$ is sufficiently small. By the same logic, the moment selecting vector allows $S^q ( \hat{\Omega}^{\frac{1}{2}}_n Z^{\#} + \hat{\psi}_n)$ in (ref) to approximate the distribution of the test statistic $\tilde{T}^q_n$ for $q\in\{r,s\}$.
Given the test statistics and the critical values, we are ready to define the modified hybrid test.
Because both types of critical values use an estimator for $\Sigma$, we strengthen Assumption (ref) and require consistency of $\hat{\Sigma}$.
With this reinforced assumption, we define $\mathcal{F}^{pt}$ as
Since Assumption (ref) implies Assumption (ref), we have $\mathcal{F}^{pt} \subset \mathcal{F}.$ We define the subset of $\mathcal{F}^{pt}$ that satisfies the null hypothesis as $\mathcal{F}^{pt}_0$, i.e., $\mathcal{F}^{pt}_0 \equiv \{(\Sigma, F)\in \mathcal{F}^{pt}: E[d_{t}]\leq 0\}$, over which we show pointwise asymptotic validity of the modified hybrid test.
Proposition (ref) shows that the modified hybrid test is pointwise asymptotically of level $\alpha$ within the class of data generating processes, $\mathcal{F}^{pt}$. That is, that given a data generating process $(\Sigma, F)\in \mathcal{F}^{pt}_0$, one could carry out the modified hybrid test while keeping the probability of committing the Type I error less than the significance level $\alpha$ for large $n$.
Though the modified hybrid test adopts the general moment selection approach, the result in andrews2010inference does not directly apply to our setup. They show the proposed test enjoys uniform asymptotic validity, which is a necessary condition of pointwise asymptotic validity and we formally define in \hyperref[sec:appendixC]{Appendix C}. Monotonicity of their test statistic plays a crucial role in attaining uniform asymptotic validity as pointed out by canayshaikh. However, the two-sided test statistic $\tilde{T}^s_n$ of the modified hybrid test is not monotone in $\hat{d}_n$ and thus violates Assumption 1(a) and 3 in andrews2010inference. Therefore, we alter their proof and achieve pointwise asymptotic validity. The proof of Proposition (ref) can be found in \hyperref[sec:appendixB]{Appendix B}.
Uniform asymptotic validity is a stronger condition than pointwise asymptotic validity in the sense that the former implies the latter. Due to the complication derived from the fact that SPA tests intrinsically deal with dependent data such as time series, the literature of SPA tests has developed focusing on pointwise asymptotic validity in contrast to the moment inequality tests in which the importance of uniform asymptotic validity has been addressed. The main focus of this paper is neither to propose an alternative to the hybrid test and nor to investigate theoretical properties of the modified hybrid test. Nonetheless, we provide a small example showing that the modified hybrid test does not satisfy uniform asymptotic validity under conditions used by andrews2010inference in \hyperref[sec:appendixC]{Appendix C}.
While Theorem (ref) tells us that the hybrid test could be pointwise asymptotically invalid, it doesn't inform us how pronounced the distortion could be in a finite sample. In this section, we explore how significantly pointwise asymptotic invalidity manifests in a finite sample through Monte Carlo simulation. Furthermore, we study the finite sample rejection probabilities of the modified hybrid test under the null hypothesis.
We use the simulation design similar to those considered in song2012testing and hansen2005test. As in \hyperref[sec:setup]{Section 2}, suppose we have a benchmark forecast and $M$ distinct alternative forecasts. We observe $n$ realized relative risks between the benchmark forecast $\varphi_{t,0}$ and $m$th alternative forecast $\varphi_{t,m}$, $d_{t,m}$ for $t=1, \ldots, n$ and $m=1,\ldots,M$. We are interested in testing the null hypothesis in (ref) meaning that the benchmark is superior to all the alternative forecasts in terms of expected risk.
For simulation, we draw realized relative risks independently from a normal distribution, i.e., $d_{t} \sim \text{i.i.d.}~ N(- \lambda_{M_0}, V)$ where $\lambda_{M_0}$ is an $M$ dimensional vector of which the first $M_0$ elements are zeros and the rest $M-M_0$ elements are ones. $M_0$ refers to the number of the alternative forecasts of which expected risks are the same as that of the benchmark as before. The relative risk $\mu =-c \lambda_k$ is non-positive and hence the design satisfies the null hypothesis. The i.i.d. observations imply that Assumption (ref), (ref) and (ref) are satisfied. The variance-covariance matrix $V$ is designed to satisfy the third condition of Theorem (ref). The off-diagonal elements of the variance matrix $V$ are zeros and the $M$ diagonal elements are determined by a random draw from a uniform distribution over $[1,2]$ at the beginning of the simulation and are fixed during the simulation.
The sample size $n$ is 200 and we draw $M \times n$ random numbers. The number of Monte Carlo repetitions and the bootstrap samples are 5,000 and 500 respectively. For the number of alternative forecasts, we consider $M \in \{50, 100\}$. For the significance level, we consider 0.01, 0.05, and 0.10. We use $\gamma=0.5$ as recommended by song2012testing, and $\kappa_n = \sqrt{\log n}$ for the tuning parameter in the modified hybrid test as recommended by andrews2010inference.
Table \hyperref[tab:MCsimu]{2} reports the simulated rejection probabilities. Hyb. indicates the hybrid test in (ref) while Boot. and Simu. refer to the modified hybrid test in Definition \hyperref[def:modifiedhybridtest]{1} with the bootstrap-based and simulation-based critical values in (ref) and (ref) respectively.
Table \hyperref[tab:MCsimu]{2} provides evidence supporting Theorem (ref). Many simulated rejection probabilities of the hybrid test exceed the significance level $\alpha$ when $M_0$ is strictly less than $M$. This phenomenon is the most pronounced when $M_0$ is slightly less than $M$, and the extent of distortion is not marginal. For example, the rejection probabilities of the hybrid test with $M=50$ and $M_0=45$ are 0.208, 0.149, and 0.070 which are almost twice, three times, and seven times larger than their corresponding significance levels $\alpha=0.10, 0.05$ and $0.01$. We have similar results in the case with $M=100$ and $M_0=95.$
Furthermore, there is a noticeable pattern in the simulated probabilities. First, when all inequalities are binding, that is, when the second condition in Theorem (ref) is not satisfied, the probabilities are close to the nominal level $\alpha$. This is because, under this data generating process, the bootstrap distribution correctly approximates the distribution of the test statistics and hence the rejection probability converges exactly to the nominal level. Second, as $M_0$ decreases, the probabilities abruptly increase, exceeding the nominal level $\alpha$ but decline gradually. This is because both test statistics converge to $\max_{m \in M_0} Z_m$, which decreases in $M_0$, where $\{Z_m: m=1, \ldots, M\}$ are independent random variables from the standard normal distribution. On the contrary, the limiting distributions of the bootstrap test statistics do not depend on $M_0$. This difference leads to diminishing rejection probabilities along decreasing $M_0$. Finally, the probabilities fall below $\alpha$ when the ratio $M_0/M$ is small: less than 0.4 for $\alpha=0.10$, 0.3 for $\alpha=0.05$, and 0.2 for $\alpha=0.01$ for the case $M=50.$ This is consistent with our findings from Table (ref) that $\bar{\alpha}$ decreases as the ratio $M_0/M$ diminishes.
Contrary to the hybrid test and as expected from our result in Proposition (ref), the simulated rejection probabilities of the modified hybrid tests are less than the nominal level except only two cases with $M=M_0=50$ and $\alpha=0.01$. The modified hybrid test appears to be conservative in that the simulated rejection probabilities are close to $\alpha/2$ when $M_0$ is strictly less than $M$. This is because two test statistics $\tilde{T}^r_n$ and $ \tilde{T}^s_n$ converge in distribution to the same distribution to which $\hat{T}^r_n$ and $ \hat{T}^s_n$ converge. Furthermore, comparing the simulated probabilities from `Boot.' and `Simu' tells us that two different critical values of the modified hybrid test, one bootstrap-based and the other simulation-based yield similar results.
song2012testing proposes the hybrid test but does not formally discuss its theoretical properties. In this article, we demonstrate with a simple example that the hybrid test may be not pointwise asymptotically of level $\alpha$ at commonly used significance levels, and provide a formal result generalizing this observation. Pointwise asymptotic invalidity of the hybrid test has a practical implication in that a researcher may commit the Type I error with probability larger than $\alpha$ even with a large sample. As an easy fix, we propose a modified hybrid test by adjusting the generalized moment selection approach to our setup, which is often used in the literature of testing moment inequality. We prove that the modified hybrid test enjoys pointiwse asymptotic validity. Finally we present Monte Carlo results supporting the theoretical findings.
While SPA tests and moment inequality tests share many common features, the literature of the former has pursued pointwise asymptotic validity and that of the latter has addressed the importance of uniform asymptotic validity. We attribute the reason to the complex nature of dependent data generating processes that SPA tests deal with. Modifying the hybrid test so as to gain uniform asymptotic validity or developing a general framework for uniform asymptotic validity in the context of SPA tests could be an interesting topic, yet it is beyond the scope of this paper and we leave it for future research.
This supplementary file consists of three appendices. \hyperref[sec:appendixA]{Appendix A} presents the tables for the values of $\bar{\alpha}$'s in Theorem (ref) under $\gamma=0.25$ and $\gamma=0.75$. \hyperref[sec:appendixB]{Appendix B} provides auxiliary lemmas and proofs for the main results in \hyperref[sec:body]{Section 3}. Finally, we define uniform asymptotic validity and show that the modified hybrid test is not uniformly asymptotically of level $\alpha$ in \hyperref[sec:appendixC]{Appendix C}.