EconBase
← Back to paper

Multi-frequency-band tests for white noise under heteroskedasticity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,754 characters · 12 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Multi-frequency-band tests for white noise under heteroskedasticity

frontmatter\begin{aug} , \and \end{aug} \begin{abstract} This paper proposes a new family of multi-frequency-band (MFB) tests for the white noise hypothesis by using the maximum overlap discrete wavelet packet transform (MODWPT). The MODWPT allows the variance of a process to be decomposed into the variance of its components on different equal-length frequency sub-bands, and the MFB tests then measure the distance between the MODWPT-based variance ratio and its theoretical null value jointly over several frequency sub-bands. The resulting MFB tests have the chi-squared asymptotic null distributions under mild conditions, which allow the data to be heteroskedastic. The MFB tests are shown to have the desirable size and power performance by simulation studies, and their usefulness is further illustrated by two applications. \end{abstract} \begin{keyword} \kwd{Heteroskedasticity; Maximum overlap discrete wavelet packet transform; Testing for white noise; Variance ratio test; Wavelets} \end{keyword}

Introduction

Consider a stochastic sequence $\{y_t\}$ with $E(y_t)=0$ for all $t\in\mathbb{Z}$. A long standing problem in time series analysis is to detect the null hypothesis that $\{y_t\}$ is white noise, i.e.,

equation[equation omitted — 83 chars of source]

In the time domain, Box and Pierce (1970) and later Ljung and Box (1978) proposed portmanteau tests to detect $H_0$ by checking whether $E(y_t y_{t-k})=0$ at some finite lags $k=1,...,K$. Their portmanteau tests require $\{y_t\}$ to be independent and identically distributed (i.i.d.), while the i.i.d. condition is restrictive in many economic and financial applications. To relax this condition, Lobato, Nankervis and Savin (2001) constructed a modified portmanteau test, which is valid when $\{y_t\}$ is a martingale difference sequence (MDS). This method was further studied by Escanciano and Lobato (2009) with a data-driven method to select an optimal lag. For the non-MDS $\{y_t\}$, some robust versions of portmanteau test were proposed in Romano and Thombs (1996) and Horowitz, Lobato, Nankervis and Savin (2006) by implementing the block bootstrap methods, Lobato (2001) by using the self-normalization technique, and Lobato, Nankervis and Savin (2002) and Zhu (2016) by estimating the asymptotic variance matrix of the first $K$ sample autocorrelations of $\{y_t\}$. However, all of the aforementioned tests require $\{y_t\}$ to be stationary, and they are thus not applicable for heteroskedastic $\{y_t\}$ (i.e., $Ey_t^2\not\equiv$ a constant for all $t$).

In the frequency domain, Gen\c{c}ay and Signori (2015) recently introduced a family of multi-scale tests for $H_0$, and their tests work for the heteroskedastic $\{y_t\}$. To illustrate the idea of multi-scale tests, we simply assume that $\{y_t\}$ is a covariance stationary process. The multi-scale tests first apply the maximum overlap discrete wavelet transform (MODWT) to $\{y_t\}$, and then obtain its high frequency component $W_m\equiv\{W_{m,t}\}$ and low frequency component $V_m\equiv\{V_{m,t}\}$ at each scale $m$, where $W_m$ and $V_m$ are related to the frequency sub-bands $[\frac{1}{2^{m+1}}, \frac{1}{2^m}]$ and $[0,\frac{1}{2^{m+1}}]$, respectively, and they are decomposed recursively from $V_{m-1}$; see the left panel in Figure\,(ref) for the decomposition way of MODWT. Next, Gen\c{c}ay and Signori (2015) showed that if $\{y_t\}$ is white noise,

equation[equation omitted — 119 chars of source]

where $\mbox{var}(W_{m,t})$ is the MODWT-based wavelet variance, and so $\mbox{var}(W_{m,t})/\mbox{var}(y_t)$ is the MODWT-based wavelet variance ratio (WVR). Motivated by ((ref)), the multi-scale tests detect $H_0$ by measuring the distance (under certain norm) between the sample version of MODWT-based WVR and $\frac{1}{2^m}$ at each scale $m$ (or jointly over the first $m$ scales). With the aid of wavelet method, the multi-scale tests are particularly suitable in situations where the data $\{y_t\}$ have jumps, kinks, seasonality and non-stationary features. This advantage does not hold for the Fourier-based frequency-domain tests in Hong (1996), Paparoditis (2000), Fan and Zhang (2004), Escanciano and Velasco (2006), and Shao (2011a). Besides the multi-scale tests, some other wavelet-based frequency-domain tests were constructed based on the wavelet spectral density estimator. In this context, Lee and Hong (2001) applied the idea of Hong (1996) to construct an asymptotically pivotal test, but their test requires $\{y_t\}$ to be stationary and homoskedastic, and its result is usually sensitive to the choice of the finest scale especially when the sample size is small; Duchesne, Li and Vandermeerschen (2010) and Li, Yao and Duchesne (2014) further developed some wavelet-based tests by using the idea of Fan (1996), however, their methods are only applicable for the stationary i.i.d. data, with some bootstrap methods to obtain the critical values.

figure[figure omitted — 360 chars of source]

Although the multi-scale tests have the aforementioned advantage over the existing ones, they have a drawback due to the decomposition way of MODWT. To see it clearly, we note that for any covariance stationary process $\{y_t\}$ and $m=1,2,...$,

equation[equation omitted — 153 chars of source]

(see Gen\c{c}ay and Signori (2015)), where $S_y(f)$ is the spectral density function of $\{y_t\}$, and it is flat under $H_0$. The result ((ref)) implies that the MODWT-based WVR at scale $m$ essentially measures the ratio of the total variance contributed by the frequency sub-band $[\frac{1}{2^{m+1}}, \frac{1}{2^m}]$. So, the multi-scale tests lack the power if $S_y(f)$ is not flat but satisfies the relationship:

equation*[equation* omitted — 132 chars of source]

As a simple illustrating example, Figure\,(ref) plots $S_y(f)$ for a white noise process and a correlated process. By construction, the contribution of frequency sub-band $[\frac{1}{2^{m+1}}, \frac{1}{2^m}]$ to the total variance of each process is the same, and the multi-scale tests are thus unable to distinguish these two processes. To detect this correlated process, an intuitive way is to further decompose the high-frequency component $W_{m}$, so that more signals to reject $H_0$ can be found within the frequency sub-band $[\frac{1}{2^{m+1}}, \frac{1}{2^m}]$. However, the MODWT fails to do this, since it does not re-decompose $W_{m}$ any more.

figure[figure omitted — 257 chars of source]

This paper is motivated to propose a new family of frequency-domain-based tests for $H_0$ by using the maximum overlap discrete wavelet packet transform (MODWPT). The MODWPT decomposes the process $\{y_t\}$ into $2^{m}$ different components $\{W_{m,n}; n=0,...,2^{m}-1\}$ at each scale $m$, where $W_{m,n}\equiv \{W_{m,n,t}\}$ is related to the frequency sub-band $[\frac{n}{2^{m+1}}, \frac{n+1}{2^{m+1}}]$, and it is decomposed recursively from $2^{m-1}$ components $\{W_{m-1,n}\}$ at the previous scale; see the right panel in Figure\,(ref) for the decomposition way of MODWPT. Unlike the MODWT, the MODWPT re-composes each $W_{m,n}$ so that the entire frequency band $[0,\frac{1}{2}]$ is refined, and it thus provides us with an effective way to largely overcome the inconsistency problem in multi-scale tests. With $\{W_{m,n}; n=0,...,2^{m}-1\}$, our testing principle uses the fact that if $\{y_t\}$ is stationary white noise,

equation[equation omitted — 126 chars of source]

where $\mbox{var}(W_{m,n,t})$ is the MODWPT-based wavelet variance, and $\mbox{var}(W_{m,n,t})/\mbox{var}(y_t)$ is the MODWPT-based WVR. Hence, at each scale $m$, we can look for the rejection evidence by measuring the distance between the sample version of MODWPT-based WVR and $\frac{1}{2^m}$ jointly over $n=1,...,2^{m}-1$. Note that we do not consider the testing signal in $W_{m,0}$ (which is identical to $V_{m}$) as done in Gen\c{c}ay and Signori (2015). Our resulting tests are called the multi-frequency-band (MFB) tests, since they are constructed by collecting signals from all frequency sub-bands (except the first one) at each scale $m$. The MFB tests are shown to have simple chi-squared limiting null distributions, under conditions that allow for higher order dependence, heteroskedasticity, and trending moments. Hence, they are easy-to-implement with great generality. Simulation studies show that the MFB tests can have desirable empirical size and power even when the sample size is small, and they can perform better than the multi-scale tests and other competitors especially when the serial dependence of the examined data exists at large lags. Also, the simulation studies indicate that the multi-scale tests could serve as diagnostic tools for many non-stationary models, including, for example, the time-varying GARCH model in Subba Rao (2006), the non-stationary GARCH model in Francq and Zako\"{i}an (2012), and the ZD-GARCH model in Li, Zhang, Zhu and Ling (2018), whose model diagnostic checking methods are absent in the literature.

Finally, two applications are given to demonstrate the usefulness of the MFB tests. In the first application, our MFB tests show that although the entire S&P500 return series in 2006--2015 is not white noise, its sub-series in 2009--2015 is white noise. These results are informative for empirical researchers, since they indicate that the S&P500 stock market possibly is not predictable in 2009--2015 but predictable in 2006--2008. Since the S&P500 stock market is relatively more volatile in 2006--2008 than 2009--2015, our findings may suggest that the S&P500 stock market is more likely to be inefficient when it is more volatile. In the second application, we apply our MFB tests to four non-stationary stock return series in Francq and Zako\"{i}an (2012), and find that three of them are not white noises. Hence, it implies that these three non-white-noise series have some dynamical structures in their conditional mean, and they should not be directly fitted by the first-order non-stationary GARCH model as done in Francq and Zako\"{i}an (2012).

The remainder of this paper is organized as follows. Section 2 introduces the MODWPT-based WVR and gives the asymptotics of its estimator. Section 3 proposes our MFB tests and studies their asymptotics. Simulations are provided in Section 4 and applications are offered in Section 5. Technical proofs are deferred to the Appendix.

Wavelet variance ratio and its estimator

The wavelet variance ratio (WVR) plays an important role in our testing principle. Below, we introduce the WVR based on the maximum overlap discrete wavelet packet transform (MODWPT) and its estimator. For more discussions on MODWPT, we refer to Percival and Walden (2000).

MODWPT-based WVR

To elaborate the definition of the MODWPT-based WVR, we simply assume that $\{y_t\}_{t=1}^{T}$ is a stationary process with mean zero. The MODWPT-based WVR is defined in terms of the MODWPT component of $\{y_t\}_{t=1}^{T}$. To compute the MODWPT component, we need a wavelet filter $\{h_l\}_{l=0}^{L-1}$ and its associated scaling filter $\{g_l\}_{l=0}^{L-1}$, where $\{h_l\}_{l=0}^{L-1}$ satisfies that $h_l=0$ for $l<0$ or $l\geq L$, and

align*[align* omitted — 103 chars of source]

and $\{g_l\}_{l=0}^{L-1}$ satisfies that $g_l=(-1)^{l+1}h_{L-1-l}$ and

align*[align* omitted — 144 chars of source]

for all nonzero integers $n$. Some well-known choices of $h_l$ and $g_l$ are given as follows:

itemize• Haar wavelet: $\{h_l\}_{l=0}^{1}=(1/2, -1/2)$ and $\{g_l\}_{l=0}^{1}=(1/2, 1/2)$. • Daubechies wavelets ($D(L)$): $D(2)$ is just the Haar wavelet. The wavelet and scaling filters for $D(4)$ are defined as $$\{h_l\}_{l=0}^{3}=\left( \frac{1-\sqrt{3}}{8},\frac{-3+\sqrt{3}}{8}, \frac{3+\sqrt{3}}{8},\frac{-1-\sqrt{3}}{8}\right) $$ and $$\{g_l\}_{l=0}^{3}=\left( \frac{1+\sqrt{3}}{8},\frac{3+\sqrt{3}}{8},\frac{3-\sqrt{3}}{8},\frac{1-\sqrt{3}}{8}\right),$$respectively. The wavelet and scaling filters for $D(L)$ with $L>4$ can be found in Daubechies (1992).

Let $L_m=(2^m-1)(L-1)+1$ for some integer $m\geq 1$. Based on $\{h_l\}_{l=0}^{L-1}$ and $\{g_l\}_{l=0}^{L-1}$, we then compute $\{\widetilde{v}_{m,n,l}\}_{l=0}^{L_m-1}$ by $$\widetilde{v}_{m,n,l}=\frac{1}{2^{m/2}}v_{m,n,l}$$ for $n=0, 1, ..., 2^{m}-1$. Here, $v_{m,n,l}$ is defined recursively by

equation*[equation* omitted — 94 chars of source]

with $v_{1,0,l}=g_l$ and $v_{1,1,l}=h_l$, where $\left[\cdot\right]$ is the integer part operator, and

equation*[equation* omitted — 134 chars of source]

Using $\{\widetilde{v}_{m,n,l}\}_{l=0}^{L_m-1}$, the MODWPT components $W_{m,n}\equiv\{W_{m,n,t}\}_{t=1}^{T}$ at scale $m$ are computed with the MODWPT coefficients

equation*[equation* omitted — 82 chars of source]

Note that $W_{m,n,t}$ can be fast calculated by using the R package “wmtsa”. Generally speaking, the MODWPT at each scale $m$ decomposes the entire frequency band $[0, \frac{1}{2}]$ into $2^m$ equal sub-bands (see the right panel in Figure\,(ref)), and the resulting $W_{m,n}$ contains the characteristics of the original time series $\{y_t\}_{t=1}^{T}$ in each sub-band $[\frac{n}{2^{m+1}},\frac{n+1}{2^{m+1}}]$.

Similar to Gen\c{c}ay and Signori (2015), we next define the wavelet variance of $\{y_t\}$ in the frequency sub-band $[\frac{n}{2^{m+1}},\frac{n+1}{2^{m+1}}]$ by

equation[equation omitted — 73 chars of source]

With $\{{\rm wvar}_{m,n}(y)\}$, we can approximately decompose the variance of $\{y_t\}$ at scale $m$ by

equation[equation omitted — 89 chars of source]

where the result ((ref)) holds, because ${\rm wvar}_{m,n}(y)\approx {\rm var}_{m,n}(y)\equiv 2\int^{\frac{n+1}{2^{m+1}}}_{\frac{n}{2^{m+1}}}S_y(f)\mathrm{d}f$ by neglecting the leakage of the wavelet filter (see Gen\c{c}ay and Signori (2015)), and ${\rm var}(y)=2\int^{1/2}_{0}S_y(f)\mathrm{d}f=\sum_{n=0}^{2^m-1}{\rm var}_{m,n}(y).$ Here, $S_y(f)$ is the spectral density function of $\{y_t\}$, and ${\rm var}_{m,n}(y)$ can be viewed as the general variance of $\{y_t\}$ in the sub-band $[\frac{n}{2^{m+1}},\frac{n+1}{2^{m+1}}]$.

Now, we define the MODWPT-based WVR in the frequency sub-band $[\frac{n}{2^{m+1}},\frac{n+1}{2^{m+1}}]$ by

equation[equation omitted — 86 chars of source]

Clearly, the result ((ref)) implies that for the general stationary process $\{y_t\}$, $\sum_{n=0}^{2^m-1}\xi_{m,n}(y)\approx 1.$ Particularly, if $\{y_t\}$ is covariance stationary white noise, Theorem (ref) below shows that the approximation symbol “$\approx$” can be replaced by the equality symbol “$=$”.

thmSuppose $\{y_t\}$ is covariance stationary white noise. Then, \begin{equation*} \xi_{m,n}(y)=\frac{1}{2^m} \end{equation*} at each scale $m$, where $n=0,...,2^m-1$.

The preceding theorem demonstrates that if $\{y_t\}$ is covariance stationary white noise, the MODWPT-based wavelet variance at each sub-band $[\frac{n}{2^{m+1}},\frac{n+1}{2^{m+1}}]$ contributes a ratio of $\frac{1}{2^{m}}$ to the total variance. In the next section, we will apply this result to form a class of tests for $H_0$. Specifically, we will measure the distance between $\xi_{m,n}(y)$ and $\frac{1}{2^m}$ under certain norm, and a large value of this distance conveys the evidence of rejection for $H_0$.

The estimator of $\xi_{m,n}(y)$

To facilitate our testing idea, an estimator of $\xi_{m,n}(y)$ is needed. In this paper, we estimate $\xi_{m,n}(y)$ by $\widehat{\xi}_{m,n,T}$, where

equation[equation omitted — 169 chars of source]

Let $z_{m,n,t}=\sum_{i=0}^{L_m-1}\sum_{j>i}^{L_m}\widetilde{v}_{m,n,i}\widetilde{v}_{m,n,j}y_{t-i}y_{t-j}$ and

equation[equation omitted — 161 chars of source]

where $s_{m,n,T}^2(z)$ is the long run variance of $\frac{1}{\sqrt{T}}\sum_{t=1}^{T}z_{m,n,t}$. Theorem (ref) below shows that the consistency and asymptotic normality of $\widehat{\xi}_{m,n,T}$ hold even for the heteroskedastic white noise $\{y_t\}$.

thmSuppose $\{y_t\}$ is heteroskedastic white noise. For any given $m\geq 1$ and $n=1,...,2^{m}-1$, (i) if Assumption (ref) in the Appendix holds, \begin{equation*} \widehat{\xi}_{m,n,T}\xrightarrow{p}\frac{1}{2^m} as T\to\infty; \end{equation*} (ii) if $\lim_{T\to\infty}\frac{1}{T}\sum_{t=1}^{T}Ey_t^2=\sigma^2<\infty$ and Assumption (ref) in the Appendix holds, \begin{equation} WV_{m,n}\equiv\sqrt{\frac{T\sigma^4}{4{\rm avar}(z_{m,n})}}\left(\widehat{\xi}_{m,n,T}-\frac{1}{2^m}\right)\xrightarrow{d} N(0,1) as T\to\infty, \end{equation} where ${\rm avar}(z_{m,n})$ is the probability limit of $s_{m,n,T}^2(z)$ in ((ref)).

To implement Theorem (ref)(ii), we need either estimate $\sigma^2$ and ${\rm avar}(z_{m,n})$ consistently or calculate them explicitly. For the general cases, $\sigma^2$ can be consistently estimated by $\widehat{\sigma}^{2}\equiv \frac{1}{T}\sum_{t=1}^{T}y_t^2$ under some mixingale conditions in Andrews (1988), and ${\rm avar}(z_{m,n})$ can be consistently estimated by the conventional Newey--West (NW) estimator $\widehat{{\rm avar}}(z_{m,n})$. For a special case that

equation[equation omitted — 116 chars of source]

we can show that $4\sigma^{-4}{\rm avar}(z_{m,n})$ in ((ref)) has an explicit formula, which can be directly calculated from the wavelet filter $\{h_l\}$. Here, the cross-joint cumulants of order four for $\{y_t\}$ is defined as the coefficients $\kappa^{a,b,c,d}$ in the Taylor's expansion: $$\log M(\xi)=\sum_{a}\xi_{a}\kappa^{a}+\frac{1}{2!}\sum_{a,b}\xi_{a}\xi_{b}\kappa^{a,b}+\frac{1}{3!}\sum_{a,b,c}\xi_{a}\xi_{b}\xi_{c}\kappa^{a,b,c} +\frac{1}{4!}\sum_{a,b,c,d}\xi_{a}\xi_{b}\xi_{c}\xi_{d}\kappa^{a,b,c,d}+\cdots,$$ where $M(\xi)=E\exp(\xi'y_{t}^{ijkl})$ with $\xi\in\mathcal{R}^{4\times1}$ and $y_{t}^{ijkl}=(y_{t-i},y_{t-j},y_{t-k},y_{t-l})'\in\mathcal{R}^{4\times1}$ for any $i,j,k,l$, and each index in the summation is running from 1 to 4.

proSuppose $\{y_t\}$ is heteroskedastic white noise and the condition ((ref)) holds. Then, $WV_{m,n}$ defined in ((ref)) can be simplified as \begin{equation} WV_{m,n}=\sqrt{\frac{T}{a(\widetilde{v}_{m,n,n})}}\left(\widehat{\xi}_{m,n,T}-\frac{1}{2^m}\right), \end{equation} where \begin{equation*} a(\widetilde{v}_{m,n_1,n_2})=\sum_{s\inZ}\,\sum_{i=i_{\min}}^{i_{\max}}\sum_{j\geq i}^{j_{\max}}\widetilde{v}_{m,n_1,i}\widetilde{v}_{m,n_1,j}\widetilde{v}_{m,n_2,i-s}\widetilde{v}_{m,n_2,j-s} \end{equation*} with $i_{\min}=\max\{0,s\}$, $i_{\max}=\min\{L_m,L_m+s\}-2$ and $j_{\max}=\min\{L_m,L_m+s\}-1$.

Note that $WV_{m,n}$ aims to convey the testing signal expressed by the WODWPT-based WVR within the frequency sub-band $[\frac{n}{2^{m+1}}, \frac{n+1}{2^{m+1}}]$, and the results of $WV_{m,n}$ in Theorem (ref)(ii) and Proposition (ref) are key to form our test statistics below.

Multi-frequency-band tests

In this section, we propose some new test statistics based on the WODWPT-based WVR to detect the null hypothesis $H_0$ in ((ref)). Let ${\bf W}_{m}\equiv(WV_{m,1},\cdots,WV_{m,2^{m}-1})'\in\mathcal{R}^{(2^{m}-1)\times1}$, and $\Sigma_{m}\in\mathcal{R}^{(2^{m}-1)\times (2^{m}-1)}$ be the asymptotic covariance matrix of ${\bf W}_{m}$ under $H_0$ with its $(i,j)$th entry $$\Sigma_{m,i,j}=\frac{{\rm acov}(z_{m,i}z_{m,j})}{\sqrt{{\rm avar}(z_{m,i})}\sqrt{{\rm avar}(z_{m,j})}},$$ where ${\rm acov}(z_{m,i}z_{m,j})$ is the probability limit of the long run covariance of $\frac{1}{\sqrt{T}}\sum_{t=1}^{T}z_{m,i,t}$ and $\frac{1}{\sqrt{T}}\sum_{t=1}^{T}z_{m,j,t}$. Since our testing principle is to measure the distance between $\widehat{\xi}_{m,n,T}$ and $\frac{1}{2^m}$ for $n=1,...,2^m-1$, a straightforward way is to consider a joint multi-frequency-band test statistic:

equation[equation omitted — 81 chars of source]

at each scale $m$. By construction, we know that under $H_0$, $$MFB_{m}\xrightarrow{d} \chi^{2}_{2^m-1} \mbox{ as }T\to\infty.$$

Our test $MFB_{m}$ is similar to the multi-scale test $GSM_m$ based on the maximum overlap discrete wavelet transform (MODWT) in Gen\c{c}ay and Signori (2015), where $$GSM_{m}\equiv (GS_1,...,GS_m)\dot{\Sigma}_m^{-1}(GS_1,...,GS_m)',$$ and under $H_0$, $GSM_{m}\xrightarrow{d} \chi^{2}_{m}$ as $T\to\infty$. Here, $\dot{\Sigma}_{m}\in\mathcal{R}^{m\times m}$ is the asymptotic covariance matrix of $(GS_1,...,GS_m)$ with $$GS_{m}\equiv \sqrt{\frac{T\sigma^4}{4{\rm avar}(z_{m})}}\left(\widehat{\xi}_{m,T}-\frac{1}{2^m}\right),$$ where $\widehat{\xi}_{m,T}$ is defined as $\widehat{\xi}_{m,n,T}$ in ((ref)) with $W_{m,n,t}$ replaced by $W_{m,t}$, ${\rm avar}(z_{m})$ is defined as ${\rm avar}(z_{m,n})$ in Theorem (ref) with $z_{m,n,t}$ replaced by $z_{m,t}^{*}$, and $$z_{m,t}^{*}=\sum_{i=0}^{L_m-1}\sum_{j>i}^{L_m}h_{m,i}h_{m,j}y_{t-i}y_{t-j}.$$ Like $GSM_{m}$, $MFB_{m}$ can also consistently detect any finite ARMA alternatives and have non-trivial power to detect the local alternative of the form: $$H_{1T}: S_{T}(f)=\frac{1}{\sqrt{T}}\Big(S(f)-\frac{1}{2}\Big)+\frac{1}{2},$$ by using the similar arguments as in Gen\c{c}ay and Signori (2015), where $S(f)$ is the non-constant spectrum. However, the two tests have distinctions due to the different decomposition ways of MODWT and MODWPT as shown in Figure\,(ref). Specifically, $GSM_m$ looks for the rejection evidence from the components $\{W_{1},...,W_{m}\}$ at the first $m$ scales, while $MFB_{m}$ does it from the components $\{W_{m,1},...,W_{m,2^{m}-1}\}$ at a given scale $m$. When $m=1$, $GSM_m$ and $MFB_{m}$ are identical. However, when $m>1$, $MFB_{m}$ tends to find more adequate testing signals than $GSM_m$, since the MODWPT zooms in the high frequency sub-bands by further decomposing $W_{m,n}$, while the MODWT does not.

To use $MFB_{m}$ in practice, we need calculate $WV_{m,n}$ in ((ref)) and replace $\Sigma_m$ in ((ref)) by a known matrix. In general cases, $WV_{m,n}$ can be calculated by replacing $\sigma^2$ and ${\rm avar}(z_{m,n})$ with $\widehat{\sigma}^2$ and the NW estimator $\widehat{{\rm avar}}(z_{m,n})$, and $\Sigma_{m}$ can be replaced by its NW estimator $\widehat{\Sigma}_{m}$, where the $(i,j)$th entry of $\widehat{\Sigma}_{m}$ is $$\widehat{\Sigma}_{m,i,j}=\frac{\widehat{{\rm acov}}(z_{m,i}z_{m,j})}{\sqrt{\widehat{{\rm avar}}(z_{m,i})}\sqrt{\widehat{{\rm avar}}(z_{m,j})}},$$ and $\widehat{{\rm acov}}(z_{m,i}z_{m,j})$ is the NW estimator of ${\rm acov}(z_{m,i}z_{m,j})$. In a particular case, if $\{y_t\}$ satisfies the condition ((ref)), $WV_{m,n}$ can be calculated explicitly as in ((ref)), and $\Sigma_{m}$ can be simplified as $A_{m}$ by the similar arguments as for Proposition (ref), where the $(i,j)$th entry of $A_{m}$ is

equation[equation omitted — 137 chars of source]

Now, we consider three computational versions of $MFB_{m}$:

itemize$MFB_{m}^{g}$ calculates $WV_{m,n}$ as in ((ref)), and replaces $\Sigma_m$ by $A_m$ in ((ref)); • $MFB_{m}^{\vartriangle}$ calculates $WV_{m,n}$ with $\sigma^2$ and ${\rm avar}(z_{m,n})$ replaced by $\widehat{\sigma}^2$ and $\widehat{{\rm avar}}(z_{m,n})$, and replaces $\Sigma_m$ by $A_m$; • $MFB_{m}^{e}$ calculates $WV_{m,n}$ with $\sigma^2$ and ${\rm avar}(z_{m,n})$ replaced by $\widehat{\sigma}^2$ and $\widehat{{\rm avar}}(z_{m,n})$, and replaces $\Sigma_m$ by $\widehat{\Sigma}_{m}$.

Note that $MFB_{m}^{g}$, $MFB_{m}^{\vartriangle}$ and $MFB_{m}^{e}$ are constructed in a similar way as the multi-scale tests $GSM_{m}^{g}$, $GSM_{m}^{\vartriangle}$ and $GSM_{m}^{e}$ in Gen\c{c}ay and Signori (2015), where we use the notation $GSM_{m}^{e}$ to denote their test $GSM_{m}$ for the notational consistency. By construction, $MFB_{m}^{g}$ and $MFB_{m}^{\vartriangle}$ are feasible for the special case that condition ((ref)) holds, while $MFB_{m}^{e}$ is valid for general cases. The same conclusion holds for their multi-scale counterparts.

Simulation

In this section, we examine the finite-sample performance of our tests $MFB_{m}^{g}$, $MFB_{m}^{\vartriangle}$ and $MFB_{m}^{e}$ in comparison with the portmanteau tests $Q_{K}$ in Ljung and Box (1978), the automatic portmanteau test $AQ$ in Escanciano and Lobato (2009), and the multi-scale tests $GSM_{m}^{g}$, $GSM_{m}^{\vartriangle}$ and $GSM_{m}^{e}$ in Gen\c{c}ay and Signori (2015). Unless stated otherwise, all MFB and GSM tests are computed with Haar wavelet in the sequel.

Size study

Let $\epsilon_t\stackrel{i.i.d.}{\sim} N(0,1)$ unless specified. To examine the empirical size of all tests, we consider the following null models:

enumerate[label=N\arabic*] • $[\bf{N(0,1)}]$ a standard normal process: $y_t=\epsilon_t$; • $[\bf{N(0,1)}$-$\bf{GARCH}]$ a GARCH process with $N(0,1)$ innovations: $y_t=\sigma_t\epsilon_t$ and $\sigma_t^2=0.001+0.05y^2_{t-1}+0.90\sigma^2_{t-1}$; • $[\bf{t_5}$-$\bf{GARCH}]$ a GARCH process as in model N2 except $\epsilon_t\stackrel{i.i.d.}{\sim} t_5$; • $[\bf{EGARCH}]$ an EGARCH process with $N(0,1)$ innovations: $y_t=\sigma_t\epsilon_t$ and $\log\sigma_t^2=0.001+0.5|\epsilon_t|-0.2\epsilon_t+0.95\log\sigma^2_{t-1}$; • $[\mbox{\bf{Mixture of normals}}]$ a mixture of two normals $N(0,1/2)$ and $N(0,1)$ with mixing probability $1/2$; • $[\bf{N(0,t)}]$: a heteroskedastic normal with trending variance: $y_t=\sqrt{t}\epsilon_t$; • $[\mbox{\bf{Time-varying GARCH}}]$ a time-varying GARCH$(1,1)$ process with $N(0,1)$ innovations: $y_t=\tau(t/T)u_t$, $\tau(x)=I(0<x<0.5)+2I(0.5\leq x<1)$, $u_t=\sigma_t\epsilon_t$ and $\sigma_t^2=0.05+0.05u^2_{t-1}+0.90\sigma^2_{t-1}$; • $[\mbox{\bf{Non-stationary GARCH}}]$ a non-stationary GARCH$(1,1)$ process with $N(0,1)$ innovations: $y_t=\sigma_t\epsilon_t$ and $\sigma_t^2=0.001+0.1096508y^2_{t-1}+0.90\sigma^2_{t-1}$; • $[\mbox{\bf{ZD-GARCH}}]$ a ZD-GARCH$(1,1)$ process with $N(0,1)$ innovations: $y_t=\sigma_t\epsilon_t$ and $\sigma_t^2=0.1096508y^2_{t-1}+0.90\sigma^2_{t-1}$; • $[\mbox{\bf{All-pass ARMA}}]$ an All-pass ARMA$(1,1)$ process with $N(0,1)$ innovations: $y_t=0.8y_{t-1}+\epsilon_t-(1/0.8)\epsilon_{t-1}$; • $[\mbox{\bf{Bilinear}}]$ a bilinear process with $N(0,1)$ innovations: $y_t=\epsilon_t+0.5\epsilon_{t-1}y_{t-2}$; • $[\mbox{\bf{Nonlinear MA}}]$ a nonlinear MA model with $N(0,1)$ innovations: $y_t=\epsilon_t+0.5\epsilon_{t-1}\epsilon_{t-2}$.

Models N1--N6 were considered by Gen\c{c}ay and Signori (2015), and except model N6, the other five models are stationary MDS with constant variances. Models N7--N9 were studied by Subba Rao (2006), Francq and Zako\"{i}an (2012), and Li, Zhang, Zhu and Ling (2018), respectively. These three models are non-stationary MDS with time-varying variances. Unlike models N1--N9, models N10--N12 are uncorrelated but non-MDS as shown in Shao (2011b).

table*[table* omitted — 3,991 chars of source]

As the settings in Gen\c{c}ay and Signori (2015), Table (ref) reports the proportion (in percentage) of rejections at 5% nominal level for all MFB and GSM tests with $m=2$, the portmanteau tests $Q_{K}$ with $K=5, 10, 20$, and the automatic portmanteau test $AQ$, where 10000 replications are generated from each null model with the sample size $T=100$, 300 or 1000. From this table, our findings are as follows:

(i) Our three MFB tests have a similar size performance as their GSM counterparts in all examined cases. When the sample size is small (e.g., $T=100$), $MFB_2^g$ has an accurate size performance, except for models N4, N6--N9 and N11--N12. As the sample size becomes larger (e.g., $T=1000$), the over-sized problem for $MFB_2^g$ is even worse. In contrast, $MFB_{2}^{\vartriangle}$ and $MFB_{2}^{e}$ can always have accurate sizes when the sample size is large, although they (particularly $MFB_{2}^{e}$) tend to be slightly over-sized when the sample size is small.

(ii) All three portmanteau tests $Q_{K}$ show good size performances in models N1, N5 and N10, but they have the severe over-sized problem in models N3--N4, N6--N9 and N11, and this problem tends to exist in models N2 and N12 even when the sample size is large (e.g., $T=1000$).

(iii) The automatic portmanteau test $AQ$ exhibits a good size performance in all examined cases, except that it tends to have a slightly over-sized problem when the sample size is small, and this problem remains in models N10--N12 even when the sample size is large.

Overall, our findings are similar to those in Gen\c{c}ay and Signori (2015). On one hand, when the sample size is small, $MFB_2^g$ (or $GSM_2^g$) has a relatively better size performance than others for most of stationary MDS data, and $MFB_2^{\vartriangle}$ (or $GSM_2^{\vartriangle}$ and $AQ$) does this for most of non-stationary or non-MDS data. On the other hand, when the sample size is large, $MFB_{2}^{\vartriangle}$ (or $GSM_{2}^{\vartriangle}$) seems to have the best size performance in general.

Power study

To examine the empirical power of all tests, we consider the following four alternative models:

enumerate[label=A\arabic*] • $[\bf{N(0,1)}$-$\bf{AR(2)}]$ an AR(2) process with $N(0,1)$ innovations: $y_t=\beta_1y_{t-1}+\beta_2y_{t-2}+\epsilon_t$; • $[\bf{N(0,1)}$-$\bf{AR(3)}]$ an AR(3) process with $N(0,1)$ innovations: $y_t=\beta_1y_{t-1}+\beta_2y_{t-3}+\epsilon_t$; • $[\bf{N(0,t)}$-$\bf{AR(2)}]$ an AR(2) process with $N(0,t)$ innovations: $y_t=\beta_1y_{t-1}+\beta_2y_{t-2}+\sqrt{t}\epsilon_t$; • $[\bf{N(0,t)}$-$\bf{AR(3)}]$ an AR(3) process with $N(0,t)$ innovations: $y_t=\beta_1y_{t-1}+\beta_2y_{t-3}+\sqrt{t}\epsilon_t$,

where $\beta_1$ (or $\beta_2$) is set to be $-0.30, -0.20, ..., 0.20,$ and $0.30$.

Model A1 was considered in Gen\c{c}ay and Signori (2015), and models A2--A4 are designed to see how the tests perform when the data have the serial dependence at a larger lag or they are heteroskedastic.

As before, we follow the settings in Gen\c{c}ay and Signori (2015), and thus restrict our analysis to compare the (size-adjusted) power of $MFB_2^g$, $GSM_2^g$, $Q_{20}$, and $AQ$ when the sample size is small. Tables (ref) and (ref) report the power (in percentage) at 5% nominal level for $MFB_2^g$, where 10000 replications are generated from each alternative model with the sample size $T=100$. To make a comparison, Tables (ref) and (ref) also report the relative power gains of $MFB_2^g$ with respect to the other three tests. From these two tables, we can have the following findings:

(i) For model A1, $MFB_2^g$ is generally more powerful than $GSM_2^g$ when $\beta_1<0$, while $GSM_2^g$ outperforms $MFB_2^g$ when $\beta_1>0$. For model A2, the advantage of $GSM_2^g$ over $MFB_2^g$ largely disappears, but $MFB_2^g$ has a huge power improvement over $GSM_2^g$ up to 786%. This implies that the power advantage of $MFB_2^g$ over $GSM_2^g$ tends to be more substantial, when the serial dependence of data happens at larger lags. For models A3--A4 with heteroskedastic data, a similar conclusion can be drawn.

(ii) For all considered four models, $MFB_2^g$ is always more powerful than $Q_{20}$. The power performance between $MFB_2^g$ and $AQ$ is mixed. For models A1 and A3, $MFB_2^g$ (or $AQ$) shows its relative better performance when $\beta_1>0$ (or $\beta_1<0$). For model A2, $MFB_2^g$ has a clear power improvement over $AQ$ up to 88%, while $AQ$ is only slightly better than $MFB_2^g$ when $\beta_1<0$ and $\beta_2$ is close to 0. For model A4, a similar phenomenon as for model A2 can be observed. All these findings once again imply that $MFB_2^g$ has a more substantial power advantage over $AQ$, when the serial dependence of data happens at larger lags.

table*[table* omitted — 6,399 chars of source]
table*[table* omitted — 6,492 chars of source]

Robust analysis

In the previous two subsections, we focus on $m=2$ for our MFB tests. This subsection aims to do some robust analysis for our MFB tests, based on the settings as in Gen\c{c}ay and Signori (2015). First, we explore the finite sample performance of our MFB tests in terms of the choice of $m$. To illustrate it, we generate $10000$ replications with sample size $T=100, 300$ or 1000 from the following AR($k$) model: $$y_t=\beta y_{t-k}+\epsilon_t,$$ where $|\beta|<1$. Figures\,(ref) and (ref) plot the (size-adjusted) power of $MFB_m^g$ (for $m=1,...,5$) against AR($1$) and AR($5$) models at 5% nominal level, respectively. As a comparison, the (size-adjusted) power of $GSM_m^g$ is also plotted in these two figures. From Figure\,(ref), we can find that when $m=1, 2$ and $3$, all MFB and GSM tests have similar power performances, and when $m=4$ and $5$, the GSM tests perform better than the MFB tests especially for $\beta>0$. In contrast, Figure\,(ref) shows that when $m=3, 4$ and $5$, the MFB tests are clearly more powerful than the GSM tests, while all tests exhibit low power when $m=1$ and $2$. These findings suggest that when the serial dependence happens at the small lag, our MFB tests can perform stably over $m$, and when the serial dependence happens at the large lag, our MFB tests with a large $m$ can perform well, and they are generally more powerful than the GSM tests in this case.

figure[figure omitted — 326 chars of source]
figure[figure omitted — 323 chars of source]

Second, we check the finite sample performance of our MFB tests in terms of the choice of wavelets. As the settings in Gen\c{c}ay and Signori (2015), we report the size and (size-adjusted) power of $MFB_2^g$ for Haar wavelet and Daubechies wavelets D(4), D(6), D(8) and D(10) in Table (ref). From this table, we can see that there is no significant difference in terms of size, but the Haar wavelet has some marginal advantages in terms of power.

table*[table* omitted — 1,043 chars of source]

Applications

Application 1

Checking whether the market index returns are predictable has been a long standing problem in the literature. The empirical studies in Lo and MacKinlay (1988) and Hong and Lee (2005) found that the S&P500 index returns are predictable. However, their empirical studies overlooked a fact that the predictability conclusion made based on the entire period may not be true for some specific sub-periods. To relieve this concern, we examine whether the recent S&P500 return series as well as their sub-series are white noises, and if the white noise assumption is rejected, the examined series is predictable, therefore giving the empirical evidence against the efficient market hypothesis.

We consider the daily S&P500 index from January 2, 2006 to December 31, 2015, with 2515 observations in total. Denote the S&P500 return $y_t=100\log(P_t/P_{t-1})$, where $P_t$ is the closing S&P500 index at day $t$. We first apply the MFB tests, the GSM tests and the AQ test to the entire 10-year return series, and the results in Panel A of Table (ref) show a very strong evidence to reject the white noise assumption for this entire series. Although the entire series is not white noise, there has a chance that its sub-series may be white noise. To examine this, we then apply all tests to five 2-year sub-series, and the results reported in Panel B of Table (ref) indicate that both 2012-2013 and 2014-2015 sub-series are white noises at the level 5%, while the other three two-year sub-series are not. For these three non-white-noise sub-series, we further check whether their one-year sub-series are white noises. The results given in Panel C of Table (ref) show that among six 1-year sub-series, the 2009, 2010 and 2011 sub-series are indeed white noises at the level 5%. In all sub-series study, our MFB tests exhibit much more rejection evidence than the GSM tests, and the AQ test fails to do this for the 2010-2011 sub-series and the 2008 and 2011 sub-series.

Overall, our testing results imply that the S&P500 return series is not white noise during 2006--2008, while it is white noise during 2009--2015. Since the S&P500 stock market is relatively more volatile in 2006--2008 than 2009--2015, our findings may indicate that the S&P500 market is more likely to be inefficient when it is more volatile.

table*[table* omitted — 6,038 chars of source]

Application 2

This subsection re-visits daily stock returns of BTC, CCME, KV-A, and MCBF in Francq and Zako\"{i}an (2012). These four data sets range from June 29, 2007, March 31, 2009, March 31, 2006, and August 28, 2007, respectively, to February 7, 2011, with 907, 468, 1220, and 867, respectively, observations in total. In Francq and Zako\"{i}an (2012), all four stock return series are fitted by the non-stationary GARCH($1, 1$) model, while no investigation is given to check whether there exists serial dependence in their conditional mean. Intuitively, if these four stock return series are white noises, they can be directly fitted by the non-stationary GARCH($1, 1$) model, otherwise, they possibly have some conditional mean dynamics, which need be filtered out first.

We use our three MFB tests as well as three GSM tests and the automatic portmanteau test $AQ$ to examine whether these four stock return series are white noises. The testing results are summarized in Table (ref), from which we find that only CCME return series is white noise, while the other three return series are not at the level 5%. Specifically, our MFB tests get more rejection evidence than the GSM tests for the KV-A return series, and the GSM tests do it better especially at the scales $m=3$ and $4$ for the BTC return series. For the MCBF return series, the white noise hypothesis is strongly rejected by all tests. Compared with the MBF and GSM tests, the test $AQ$ can not find the significant evidence of rejection for BTC, CCME and KV-A return series.

In summary, our testing results imply that only CCME return series has no serial dependence on its conditional mean, and it is thus suitable to fit this series by the non-stationary GARCH($1, 1$) model. However, the other three return series (particularly, MCBF) most likely have serial dependence on their conditional mean, and without filtering out the conditional mean effect ahead, the fittings in Francq and Zako\"{i}an (2012) may be inappropriate for these three series.

table*[table* omitted — 2,182 chars of source]