EconBase
← Back to paper

Testing the martingale difference hypothesis in high dimension

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

218,956 characters · 20 sections · 118 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing the Martingale Difference Hypothesis in High Dimension

\if11 {

\affil[a]{\it Joint Laboratory of Data Science and Business Intelligence, Southwestern University of Finance and Economics, Chengdu, Sichuan Province, China} \affil[b]{\it Center for Statistics and Data Science, Beijing Normal University, Zhuhai, Guangdong Province, China} \affil[c]{\it Department of Statistics, University of Illinois at Urbana-Champaign, Champaign, IL, U.S.A. }

} \fi

\if01 {

center[center omitted — 92 chars of source]

} \fi

abstractIn this paper, we consider testing the martingale difference hypothesis for high-dimensional time series. Our test is built on the sum of squares of the element-wise max-norm of the proposed matrix-valued nonlinear dependence measure at different lags. To conduct the inference, we approximate the null distribution of our test statistic by Gaussian approximation and provide a simulation-based approach to generate critical values. The asymptotic behavior of the test statistic under the alternative is also studied. Our approach is nonparametric as the null hypothesis only assumes the time series concerned is martingale difference without specifying any parametric forms of its conditional moments. As an advantage of Gaussian approximation, our test is robust to the cross-series dependence of unknown magnitude. To the best of our knowledge, this is the first valid test for the martingale difference hypothesis that not only allows for large dimension but also captures nonlinear serial dependence. The practical usefulness of our test is illustrated via simulation and a real data analysis. The test is implemented in a user-friendly R-function.

{\it Key words:} $\alpha$-mixing, Gaussian approximation, high-dimensional statistical inference, martingale difference hypothesis, parametric bootstrap

{\it JEL code}: C12, C15, C55

\onehalfspacing

Introduction

Testing the martingale difference hypothesis is a fundamental problem in econometrics and time series analysis. The concept of martingale difference plays an important role in many areas of economics and finance. Several economic and financial theories such as the efficient markets hypothesis Fama1970,Fama1991,LeRoy1989,Lo1997, rational expectations Hall1978 and optimal asset pricing Cochrane2005,Fama2013, yield such dependence restrictions on the underlying economic and financial variables. More formally, let $\{{\mathbf x}_t\}$ be a $p$-dimensional time series with $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ for any $t\in\mathbb{Z}$. Write ${\mathbf x}_t=(x_{t,1},\ldots,x_{t,p})^{\scriptscriptstyle {\rm \top}}$ and denote by $\mathscr{F}_{s}$ the $\sigma$-field generated by $\{{\mathbf x}_t\}_{t\leqslant s}$. We call $\{{\mathbf x}_t\}_{t\in\mathbb{Z}}$ a martingale difference sequence (MDS) if and only if $\mathbb{E}({\mathbf x}_t\,|\,\mathscr{F}_{t-1})=\boldsymbol{0}$ for any $t\in\mathbb{Z}$. Given the observations $\{{\mathbf x}_t\}_{t=1}^{n}$, we are interested in the hypothesis testing problem:

align[align omitted — 164 chars of source]

The MDS hypothesis implies that the past information does not help to improve the prediction of future values of a MDS, so the best nonlinear predictor of the future values of a MDS given the current information set is just its unconditional expectation. The theme of the lack of predictability is of central interest in economics and finance and has stimulated a huge literature in both econometrics and time series analysis.

So far most of the work on MDS testing is restricted to the univariate case, i.e., $p=1$. In one strand of literature, the MDS testing problem is reduced to testing the uncorrelatedness in either time domain or spectral domain. See BP1970, LB1978, Durlauf1991, Hong1996, Deo2000, LNS2001, and Shao2011a,Shao2011b, among others. These tests target on serial correlation but are unable to capture nonlinear serial dependence. There are examples of uncorrelated processes that are not MDS such as certain bilinear processes and nonlinear moving average processes, see DL2003 for specific examples. Hence, it is important to develop tests that can go beyond linear serial dependence. In the specification testing literature, the exponential function based approach, pioneered by Bierens1984, Bierens1990, DeJong1996 and BP1997, is capable of detecting nonlinear serial dependence. Using the characteristic function, Hong1999 proposed the generalized spectral density as a new tool for specification testing in a nonlinear time series framework; see HL2003 and HL2005 for further developments. As an interesting extension of Hong1999, EV2006a developed a MDS test based on the generalized spectral distribution function to capture nonlinear serial dependence at all lags. Parallel to the exponential/characteristic function based approach, the indicator/distribution function based approach has been taken by Stute1997, KS1999, DL2003, and PW2005 among others. We refer to EL2009 for a comprehensive review.

For the multivariate time series, i.e., $p>1$, the literature for the MDS testing is scarce. Although it is expected that most of the above-mentioned tests can be extended to relatively low dimensional case, the theoretical and empirical properties of these tests are unknown. Recently, HLZ2017 proposed a multivariate extension of the classical univariate variance ratio test LM1988,PS1988,CD2006 to test a weak form of the efficient markets hypothesis, i.e., uncorrelatedness of ${\mathbf x}_t$. As argued in HLZ2017, the rationale to consider the MDS test for multivariate time series is that even if the MDS hypothesis holds for each component series $\{x_{t,j}\}_{t\in\mathbb{Z}}$, the MDS hypothesis could be violated at the multivariate level. In particular, the current return on the $i$th asset may be predicted by past observations of the $j$th asset. A univariate test may fail to detect this kind of cross-serial dependence, which can be captured by a multivariate test. Since it is well known that the variance ratio test only targets on serial correlation, the test of HLZ2017 is unable to capture nonlinear serial dependence.

Nowadays, time series of moderate or high dimension are routinely collected or generated owing to the advance in science and technology. For example, S&P 500 index measures the stock performance of 500 large companies listed on stock exchanges in the United States, and it is tempting to ask whether the stock returns of the 500 companies are predictable at the daily or weekly frequency for a given time period (say, 5 years). The same question can be asked for the stocks within the same sector, such as those in the real estate sector (see Section (ref) for data illustration). This naturally leads us to the regime where the dimension $p$ is comparable to or exceeds the sample size $n$. To the best of our knowledge, there is no MDS testing procedure available in the literature that allows the dimension $p$ to exceed the sample size $n$. Most of the aforementioned tests developed in the univariate setting require nontrivial modification to accommodate the high-dimensionality. The multivariate variance ratio test in HLZ2017 allows for growing dimension $p$ in their theory (i.e., $1/p+p/n=o(1)$) but is quite limited since their test cannot be implemented when $p>n$ and may encounter computational problems when $p$ is large (say, $p>120$); see Section (ref) for more details.

To fill this gap, we introduce a new test for the MDS hypothesis of multivariate and possibly high-dimensional time series. We first use the element-wise max-norm of a sample-based matrix to characterize the nonlinear dependence of underlying $p$-dimensional time series $\{{\mathbf x}_t\}$ at a given lag $j\geqslant1$, and then combine such information at different lags to propose our test statistic. Owing to the high-dimensionality and unknown temporal and cross-series dependence, the limiting null distribution of our test statistic is hard to derive, and it may not even have a closed form. To circumvent such difficulty, we employ the celebrated Gaussian approximation technique CCK2013, which has undergone a rapid development recently, to establish the asymptotic equivalence between the null distribution of our test statistic and that of a certain function of a multivariate Gaussian random vector. Our theoretical analysis shows that our proposed test works even if $p$ grows exponentially with respect to the sample size $n$, provided that some suitable regularity assumptions hold. To facilitate feasible inference, we propose a simulation-based approach to generate critical values. We also investigate the power behavior of our test under some local alternatives.

Since the seminal contribution of CCK2013, the literature on Gaussian approximation in the high-dimensional setting has been growing rapidly. For the sample mean of independent random vectors, we mention CCK2013, CCK2017, DZ2020, FK2020, KMB2020, CCKK2019, and CCK2020. For high-dimensional $U$-statistics and $U$-processes, see Chen2018 and CK2019 for recent developments. The applicability of Gaussian approximation has also been extended to high-dimensional time series setting by ZW2017, ZC2018, CCK2019 and CCW2020. Also see CYZ2017,CZZZ2017,CZZW2017,CQYZ2018, and YC2020 among others for the use of Gaussian approximation or variants in high-dimensional statistical inference.

ZW2017 and ZC2018 considered the Gaussian approximation for $$\max_{1\leqslant j\leqslant p}\frac{1}{\sqrt{n}}\sum_{t=1}^nx_{t,j}$$ with the physical dependence measure Wu2005 imposed on $\{{\mathbf x}_t\}$, and CCK2019 considered the same problem when $\{{\mathbf x}_t\}$ is a $\beta$-mixing sequence. CCW2020 studied the Gaussian approximations for $\mathbb{P}(n^{-1/2}\sum_{t=1}^n{\mathbf x}_t\in A)$ over some general classes of the set $A$ (hyper-rectangles, simple convex sets and sparsely convex sets) under three different dependency framework ($\alpha$-mixing, $m$-dependent, and physical dependence measure), which include the results obtained in ZW2017, ZC2018 and CCK2019 as special cases. Compared to the use of Gaussian approximation results for high-dimensional time series in the existing works, our test statistic is considerably more involved and motivates us to develop new techniques for establishing the asymptotic equivalence between the null distribution of our test statistic and that of a certain function of a multivariate Gaussian random vector. More specifically, the theoretical analysis in this paper targets on the Gaussian approximation for some function of the high-dimensional vector $(n-K)^{-1/2}\sum_{t=1}^{n-K}\boldsymbol \eta_t$, where $\boldsymbol \eta_t$ is a newly defined vector based on $\{{\mathbf x}_t,{\mathbf x}_{t+1},\ldots,{\mathbf x}_{t+K}\}$ and $K$ is the number of lags involved in our test statistic. Since $K$ is allowed to grow with the sample size $n$ in our setting, the dependence structure among $\{\boldsymbol \eta_t\}$ will vary with $K$ which cannot be covered in the frameworks of above mentioned works, and the existing Gaussian approximation results cannot be applied here. Some nontrivial technical challenges need to be addressed in our theoretical analysis.

From a methodological and practical viewpoint, we highlight a few appealing features of our proposed test:

(a) Our approach is nonparametric as the null hypothesis only assumes the time series concerned is martingale difference without specifying any parametric forms of its conditional moments. Hence, it is robust to second-order and higher-order conditional moments of unknown forms, including conditional heteroscedasticity, a prominent feature of many financial time series.

(b) It allows the dimension $p$ to grow exponentially with respect to the sample size $n$, and works well for a broad range of dimension $p$ even at a medium sample size (e.g., $n=300$) as shown in our simulation studies. We have developed an R-function {\verb"MartG_test"} in the package {\verb"HDTSA"} which implements the test in an automatic manner.

(c) There is no particular requirement on the strength of cross-series dependence in our theory, so our test is applicable to time series with cross-series dependence of unknown magnitude. Strong cross-series dependence has been commonly observed in many real high-dimensional time series data.

The rest of this paper is organized as follows. The methodology and theoretical analysis are given in Sections (ref) and (ref), respectively. Section (ref) extends the proposed test to more general settings. Section (ref) studies the finite sample performance of our proposed test. A real data analysis is presented in Section (ref). Section (ref) concludes the paper. Section (ref) includes the mathematical proofs of our main results. Some additional technical arguments and numerical studies are given in the supplementary material. At the end of this section, we introduce some notation that is used throughout the paper. For any positive integer $q\geqslant2$, we write $[q]=\{1,\ldots,q\}$ and denote by $\mathbb{S}^{q-1}$ the $q$-dimensional unit sphere. For any $q_1\times q_2$ matrix ${\mathbf M}=(m_{i,j})_{q_1\times q_2}$, let $|{\mathbf M}|_{\infty}=\max_{i\in[q_1],j\in[q_2]}|m_{i,j}|$ and $|{\mathbf M}|_0=\sum_{i=1}^{q_1}\sum_{j=1}^{q_2}I(m_{i,j}\neq0)$, where $I(\cdot)$ denotes the indicator function. Specifically, if $q_2=1$, we use $|{\mathbf M}|_\infty=\max_{i\in[q_1]}|m_{i,1}|$ and $|{\mathbf M}|_0=\sum_{i=1}^{q_1}I(m_{i,1}\neq 0)$ to denote the $L_\infty$-norm and $L_0$-norm of the $q_1$-dimensional vector ${\mathbf M}$, respectively. For any $q$-dimensional vector ${\mathbf a}=(a_1,\ldots,a_q)^{\scriptscriptstyle {\rm \top}}$, write $\psi({\mathbf a})$ as the $q$-dimensional vector $\{\psi(a_1),\ldots,\psi(a_q)\}^{\scriptscriptstyle {\rm \top}}$ for given function $\psi:\mathbb{R}\rightarrow\mathbb{R}$, and denote by ${\mathbf a}_{\mathcal{L}}$ the subvector of ${\mathbf a}$ collecting the components indexed by a given index set $\mathcal{L}\subset[q]$.

Methodology

Test statistic and the associated critical values

Let $\{{\mathbf x}_t\}$ be a $p$-dimensional time series with $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ for any $t$. Given the observations $\{{\mathbf x}_t\}_{t=1}^{n}$, we shall develop a martingale difference hypothesis test that can capture certain nonlinear dependence between ${\mathbf x}_t$ and ${\mathbf x}_{t+j}$, for $j\in \mathbb N_+$. To this end, we let $\boldsymbol{\phi}(\cdot): \mathbb{R}^p\rightarrow \mathbb{R}^d$ represent a map that is provided by the user. For example, $\boldsymbol{\phi}({\mathbf x})={\mathbf x}$ is the linear identity map; $\boldsymbol{\phi}({\mathbf x})=\{{\mathbf x}^{{\scriptscriptstyle {\rm \top}}},({\mathbf x}^2)^{{\scriptscriptstyle {\rm \top}}}\}^{{\scriptscriptstyle {\rm \top}}}$ includes both linear and quadratic terms, where ${\mathbf x}^2=(x_{1}^2, \ldots, x_{p}^2)^{{\scriptscriptstyle {\rm \top}}}$ with ${\mathbf x}=(x_1,\ldots,x_p)^{\scriptscriptstyle {\rm \top}}$; $\boldsymbol{\phi}({\mathbf x})=\cos({\mathbf x})$ captures certain type of nonlinear dependence, where $\cos({\mathbf x})=\{\cos(x_1),\ldots,\cos(x_p)\}^{\scriptscriptstyle {\rm \top}}$ with ${\mathbf x}=(x_1,\ldots,x_p)^{\scriptscriptstyle {\rm \top}}$.

Denote $\boldsymbol{\gamma}_j=(n-j)^{-1}\sum_{t=1}^{n-j} \mathbb{E}[{\rm vec}\{\boldsymbol{\phi}({\mathbf x}_{t}){\mathbf x}_{t+j}^{{\scriptscriptstyle {\rm \top}}}\}]$ for each $j\geqslant1$. Our proposal for testing the martingale difference hypothesis consists in checking all the pairwise covariance between $\boldsymbol{\phi}({\mathbf x}_t)$ and ${\mathbf x}_{t+j}$, namely, our null hypothesis is now

align[align omitted — 101 chars of source]

It is easy to see that $H_0$ in (ref) implies $H_0'$ in (ref) but not vice versa. In theory, it would be ideal to develop a test that is consistent with any violation of $H_0$ but this is very challenging in a model free setting, since the alternative we target is huge owing to the high-dimensionality and nonlinear serial dependence at all lags. As argued in PJ2014, “{\it Typically, the information set includes the infinite past history of the series,.... If a finite number of lagged values is included in the conditioning set, some dependence structure in the process may be missed due to omitted lags. However, tests that are designed to cope with the infinite lag case may have very low power (e.g., DeJong1996) and may not be feasible in empirical applications.}” Thus even in the low-dimensional setting, it is not clear whether there is a practical benefit for a test that is consistent with all alternatives. This motivates us to relax the null hypothesis $H_0$ and focus on the directional alternatives encoded by the function $\boldsymbol{\phi}(\cdot)$, which is pre-specified by the user and can incorporate some prior information.

Note that if the time series $\{{\mathbf x}_t\}$ is strictly stationary, then $\boldsymbol{\gamma}_j=\mathbb{E}[{\rm vec}\{\boldsymbol{\phi}({\mathbf x}_{0}){\mathbf x}_{j}^{{\scriptscriptstyle {\rm \top}}}\}]$ which represents the population-level nonlinear dependence measure at lag $j$. In our asymptotic theory, no stationarity assumption needs to be imposed. To test $H_0'$, it is natural to consider a test statistic with the following form

align[align omitted — 92 chars of source]

where $\hat{\boldsymbol{\gamma}}_j=(n-j)^{-1}\sum_{t=1}^{n-j} {\rm vec}\{\boldsymbol{\phi}({\mathbf x}_{t}){\mathbf x}_{t+j}^{{\scriptscriptstyle {\rm \top}}}\}$ is the estimator of $\boldsymbol{\gamma}_j$. Here $K=o(n)$ is a truncation lag and is allowed to grow with respect to the sample size $n$. This flexibility is important when there exists nonlinear serial dependence at large lags.

Intuitively, a large value of $T_n$ provides evidence against $H_0'$ in (ref) and then we can reject $H_0$ in (ref) if

align[align omitted — 52 chars of source]

where ${\rm cv}_\alpha>0$ is the critical value at the significance level $\alpha\in(0,1)$. To determine ${\rm cv}_\alpha$, we need to derive the distribution of $T_n$ under $H_0$. Write $\hat{\boldsymbol{\gamma}}=(\hat{\boldsymbol{\gamma}}_1^{\scriptscriptstyle {\rm \top}},\ldots,\hat{\boldsymbol{\gamma}}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ and $\boldsymbol{\gamma}=(\boldsymbol{\gamma}_1^{\scriptscriptstyle {\rm \top}},\ldots,\boldsymbol{\gamma}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$. For fixed $(p,d,K)$ and under suitable moment and weak dependence conditions, it follows from the central limit theorem that $\sqrt{n}(\hat{\boldsymbol{\gamma}}-\boldsymbol{\gamma})\rightarrow_d\mathcal{N}(\boldsymbol{0},\mathring{\boldsymbol{\Sigma}}_{K})$ as $n\rightarrow\infty$ for some positive definite matrix $\mathring{\boldsymbol{\Sigma}}_{K}\in\mathbb{R}^{(Kpd)\times(Kpd)}$. Let $\mathring{\mathbf g}:=(\mathring{g}_1,\ldots,\mathring{g}_{Kpd})^{\scriptscriptstyle {\rm \top}}\sim {\mathcal{N}}(\boldsymbol{0},\mathring{\boldsymbol{\Sigma}}_K)$. By the continuous mapping theorem, the distribution of $T_n$ under $H_0$ can be approximated by that of its Gaussian analogue $\mathring{G}_K = \sum_{j=1}^{K} |\mathring{\mathbf g}_{\mathcal{L}_j}|_\infty^2$ in the scenario with fixed $(p,d,K)$, where $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$. Write $\tilde{n}=n-K$ and let

align[align omitted — 354 chars of source]

for any $t\in[\tilde{n}]$. Define

align[align omitted — 151 chars of source]

which is the long-run covariance matrix of the sequence $\{\boldsymbol \eta_t\}_{t=1}^{\tilde{n}}$. For fixed $(p,d,K)$, the asymptotic covariance $\mathring{\boldsymbol{\Sigma}}_K$ of $\sqrt{n}(\hat{\boldsymbol{\gamma}}-\boldsymbol{\gamma})$ is essentially the limit of $\boldsymbol{\Sigma}_{n,K}$ specified in (ref) as $n\rightarrow\infty$. In the high-dimensional scenarios, i.e., when $(p,d,K)$ is diverging with respect to $n$, Proposition (ref) indicates that such approximation for the null distribution of $T_n$ is still valid even when $p$ and $d$ grow exponentially with respect to the sample size $n$.

propositionAssume Conditions {\rm (ref)--(ref)} in Section {\rm (ref)} hold and $G_K = \sum_{j=1}^{K} |\mathbf g_{\mathcal{L}_j}|_\infty^2$, where $\mathbf g=(g_1,\ldots,g_{Kpd})^{\scriptscriptstyle {\rm \top}} \\\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_{n,K})$ and $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$. Let $K=O(n^\delta)$ for some constant $0\leqslant \delta<f_1(\tau_1,\tau_2)$ with $f_1(\tau_1,\tau_2)$ defined as (ref) in Section {\rm (ref)}. Then it holds that $ \sup_{x>0}|\mathbb{P}_{H_0}( T_n >x) - \mathbb{P}(G_K > x)|=o(1)$ as $n\rightarrow\infty$, provided that $\log(pd)=o(n^{c})$ for some constant $c>0$ only depending on $(\tau_1,\tau_2,\delta)$.

Proposition (ref) reveals that the Kolmogorov-Smirnov distance between the null distribution of the proposed test statistic $T_n$ and the distribution of $G_K$ converges to zero, even when $p$ and $d$ diverge at some exponential rate of $n$. Letting

align[align omitted — 104 chars of source]

in (ref), Proposition (ref) yields that $\mathbb{P}_{H_0}(T_n>{\rm cv}_\alpha)\to \alpha$ as $n\to\infty$. Since the long-run covariance matrix $\boldsymbol{\Sigma}_{n,K}$ is usually unknown in practice, we need to replace it by some estimate $\widehat{\boldsymbol{\Sigma}}_{n,K}$ and then use $\hat{{\rm cv}}_\alpha$ defined below to approximate the desired critical value ${\rm cv}_{\alpha}$ specified in (ref):

align[align omitted — 135 chars of source]

where $\mathcal{X}_n=\{{\mathbf x}_1,\ldots,{\mathbf x}_n\}$ and $\hat{G}_K = \sum_{j=1}^{K} |\hat{\mathbf g}_{\mathcal{L}_j}|_\infty^2$ with $\hat{\mathbf g}:=(\hat{g}_1,\ldots,\hat{g}_{Kpd})^{\scriptscriptstyle {\rm \top}}\sim {\mathcal{N}}(\boldsymbol{0},\widehat\boldsymbol{\Sigma}_{n,K})$ and $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$. Then we reject the null hypothesis $H_0$ specified in (ref) if

align[align omitted — 62 chars of source]

We defer the details of $\widehat{\boldsymbol{\Sigma}}_{n,K}$ to Section (ref).

remark{\rm If we select the function $\boldsymbol{\phi}({\mathbf x})={\mathbf x}$, the test statistic $T_n$ defined in (ref) can also be applied to test the high-dimensional white noise hypothesis, i.e., $H_0:\{{\mathbf x}_t\}_{t\in\mathbb{Z}}$ is white noise versus $H_1:\{{\mathbf x}_t\}_{t\in\mathbb{Z}}$ is not white noise. CYZ2017 considered this hypothesis testing problem with $L_\infty$-type test statistic using the maximum absolute autocorrelations and cross-correlations of the component series in ${\mathbf x}_t$ over all lags $k\in[K]$. It is well known that the $L_\infty$-type test statistic is powerful against the sparse alternatives, that is, only a small fraction of the elements in $\boldsymbol{\gamma}=(\boldsymbol{\gamma}_1^{\scriptscriptstyle {\rm \top}},\ldots,\boldsymbol{\gamma}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ are nonzero, while it can be powerless for the dense but faint alternatives, i.e., when most elements in $\boldsymbol{\gamma}=(\boldsymbol{\gamma}_1^{\scriptscriptstyle {\rm \top}},\ldots,\boldsymbol{\gamma}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ are nonzero but with very small magnitudes. To remedy such weakness, our proposed $T_n$ in (ref) combines the signals from different lags together using the sum of squares and is expected to improve the power performance in case of dense but faint alternatives. On the technical side, constructing the Gaussian approximation to the null distribution of $T_n$ defined in (ref) is more challenging than that for the $L_\infty$-type statistic used in CYZ2017. CYZ2017 only considered the case with fixed $K$ under the $\beta$-mixing assumption. The null distribution of their test statistic can be easily obtained by the associated Gaussian approximation results developed in CCK2019. In this paper, we only impose the $\alpha$-mixing assumption on $\{{\mathbf x}_t\}$ and the corresponding $\alpha$-mixing coefficients of $\{\boldsymbol \eta_t\}$ become a triangular array owing to the divergence of $K$. To the best of our knowledge, our paper is the first attempt to derive the Gaussian approximation results in such a complex setting.}
remark{\rm As we mentioned earlier, the only paper that allows growing dimension for the martingale difference hypothesis testing is HLZ2017, which generalized the variance ratio test to multivariate time series. In their asymptotic theory, they considered both finite/fixed horizon (i.e., fixed $K$) and increasing horizon (i.e., $K\rightarrow\infty$ but $K^2/n\rightarrow 0$), which is also allowed in our theory. In their Theorem 7, they presented the limiting null distribution of a particular test statistic $Zd_{\rm tr}$ under the restriction that the dimension $p$ grows but $p/n\rightarrow 0$. Their another two test statistics $Z_{\rm tr}$ and $Z_{\rm det}$ for the setting of fixed $p$ cannot be implemented in practice when $p>\sqrt{n}$. By contrast, our test statistic can work for a much broader range of $p$, including the case $p\gg n$, and thus is advantageous in dealing with the martingale difference hypothesis testing for high-dimensional time series. In addition, we can capture nonlinear serial dependence owing to the flexibility of user-chosen $\boldsymbol{\phi}(\cdot)$, which yields a nonlinear dependence measure. In practice, we need to set the lags $K$ and the user-chosen map $\boldsymbol{\phi}(\cdot)$, which can incorporate some prior information we have. For example, if the time series is expected to exhibit seasonal dependence, then $K$ should be large enough to include some seasonal lags. If we are dealing with stock return data, then including quadratic terms in $\boldsymbol{\phi}(\cdot)$ might help to capture potential nonlinear dependence. }
remark{\rm If the time series $\{{\mathbf x}_t\}$ is strictly stationary, we know the transformed data $\{\boldsymbol \eta_t\}$ is also strictly stationary and our test statistic $T_n$ given in (ref) essentially converts the MDS testing problem for ${\mathbf x}_t$ to testing zero mean for the transformed data $\boldsymbol \eta_t$. There are indeed several papers in the literature of Gaussian approximation that tackle the mean testing problem for high-dimensional time series; see ZW2017, ZC2018, CCK2019 and CCW2020. ZW2017 and ZC2018 considered the Gaussian approximation theory in the framework that assumes the physical dependence Wu2005 among $\{\boldsymbol \eta_t\}$. CCK2019 and CCW2020 considered the Gaussian approximation theory, respectively, in the frameworks that assume the $\beta$-mixing assumption and $\alpha$-mixing assumption for $\{\boldsymbol \eta_t\}$. Notice that the dependence structure among $\{\boldsymbol \eta_t\}$ will vary with $K$. The dependence framework for $\{\boldsymbol \eta_t\}$ assumed in these existing works do not cover our current setting, thus the existing Gaussian approximation results cannot be used for approximating the null distribution of our proposed test statistic $T_n$. }

Estimation of long-run covariance matrix

In the low-dimensional setting, long-run covariance matrix estimation (or heteroscedastic-autocorrelation-consistent estimation) is a classic problem in econometrics and time series analysis and there is a rich literature. We refer the readers to two foundational papers by NW1987 and Andrews1991. In the high-dimensional setting, the estimator proposed in the low-dimensional environment can still be used, but establishing the proper probabilistic bounds for the difference is very challenging. Recall $\tilde{n}=n-K$. Following CYZ2017, we adopt the following estimate for the long-run covariance matrix $\boldsymbol{\Sigma}_{n,K}$:

align[align omitted — 157 chars of source]

where $\widehat \mathbf H_j=\tilde{n}^{-1}\sum_{t=j+1}^{\tilde{n}}(\boldsymbol \eta_t-\bar{\boldsymbol \eta})(\boldsymbol \eta_{t-j}-\bar{\boldsymbol \eta})^{{\scriptscriptstyle {\rm \top}}}$ if $j\geqslant0$ and $\widehat \mathbf H_j=\tilde{n}^{-1}\sum_{t=-j+1}^{\tilde{n}}({\boldsymbol \eta}_{t+j}-\bar{\boldsymbol \eta})({\boldsymbol \eta}_{t}-\bar{\boldsymbol \eta})^{{\scriptscriptstyle {\rm \top}}}$ otherwise, with $\bar{{\boldsymbol \eta}}=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}} {\boldsymbol \eta}_t$. Here $\mathcal{K}(\cdot)$ is a symmetric kernel function that is continuous at 0, and $b_n$ is the bandwidth diverging with $n$. The theoretical property of $\widehat{\boldsymbol{\Sigma}}_{n,K}$ defined as (ref) is summarized in Proposition (ref) in Section (ref). As indicated in Andrews1991, to make $\widehat{\boldsymbol{\Sigma}}_{n,K}$ given in (ref) be positive semi-definite, we can require the kernel function $\mathcal{K}(\cdot)$ to satisfy $\int_{-\infty}^\infty \mathcal{K}(x)e^{-{\rm i}x\lambda}\,{\rm d}x\geqslant0$ for any $\lambda\in\mathbb{R}$, where ${\rm i}=\sqrt{-1}$. The Bartlett kernel, Parzen kernel and Quadratic Spectral kernel all satisfy this requirement. See Section (ref) for the explicit forms of these kernels.

Given $\widehat{\boldsymbol{\Sigma}}_{n,K}$, to compute $\hat{{\rm cv}}_{\alpha}$ given in (ref), we need to generate $\hat{\mathbf g}:=(\hat{g}_1,\ldots,\hat{g}_{Kpd})^{\scriptscriptstyle {\rm \top}}\sim {\mathcal{N}}(\boldsymbol{0},\widehat\boldsymbol{\Sigma}_{n,K})$. Notice that $\widehat{\boldsymbol{\Sigma}}_{n,K}$ is a $(Kpd)\times(Kpd)$ matrix. The standard procedure is based on the Cholesky decomposition of $\widehat{\boldsymbol{\Sigma}}_{n,K}$ but generating $\hat{\mathbf g}$ is a computationally $(nK^2p^2d^2+K^3p^3d^3)$-hard problem that requires a large storage space for $\widehat{\boldsymbol{\Sigma}}_{n,K}$. In practice, $p$ and $d$ can be quite large. As suggested in CYZ2017, we can generate $\hat{\mathbf g}$ as follows:

algorithm[algorithm omitted — 681 chars of source]

We can show that $\hat{\mathbf g}$ obtained in Algorithm (ref) satisfies $\hat{\mathbf g} \,|\, \mathcal{X}_n\sim \mathcal{N}(\boldsymbol{0},\widehat{\boldsymbol{\Sigma}}_{n,K})$. The computational complexity of Step 2 in Algorithm (ref) is just $O(n^3)$ which is independent of $(p,d)$. When $p$ and $d$ are large, the required storage space of Algorithm (ref) is also much smaller than that of the standard procedure since it only requires to store $\{\boldsymbol \eta_t\}_{t=1}^{\tilde{n}}$ and $\bar{\boldsymbol \eta}$ rather than $\widehat{\boldsymbol{\Sigma}}_{n,K}$. In practice, we can draw $\hat{\mathbf g}_1,\ldots,\hat{\mathbf g}_B$ independently by Algorithm (ref) for some large integer $B$ and then take the $\lfloor B\alpha\rfloor$th largest value among $\hat{G}_{K,1},\ldots,\hat{G}_{K,B}$ to approximate $\hat{\rm cv}_\alpha$ defined as (ref), where $\hat{G}_{K,i} = \sum_{j=1}^K |\hat{\mathbf g}_{i,\mathcal{L}_j}|_\infty^2$ with $\hat{\mathbf g}_i=(\hat{g}_{i,1},\ldots,\hat{g}_{i,Kpd})^{\scriptscriptstyle {\rm \top}}$ and $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$.

Theoretical property

Recall $T_n=n\sum_{j=1}^{K}|\hat{\boldsymbol{\gamma}}_j|_{\infty}^2$. Since the distribution of $\hat{\boldsymbol{\gamma}}=(\hat{\boldsymbol{\gamma}}_1^{\scriptscriptstyle {\rm \top}},\ldots,\hat{\boldsymbol{\gamma}}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ can be well approximated by that of $\bar{\boldsymbol \eta}=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\boldsymbol \eta_t$ with $\tilde{n}=n-K$, the difference between the distributions of $T_n$ and $\tilde{T}_n:=\tilde{n}\sum_{j=1}^K|\bar{\boldsymbol \eta}_{\mathcal{L}_j}|_\infty^2$ is expected to be asymptotically negligible. See Lemma L2 in Section (ref). The key step in our theoretical analysis is to approximate the null distribution of $\tilde{T}_n$ by Gaussian approximation.

For any $j_1,\ldots,j_K\in[pd]$ and $x>0$, let $\mathcal A_{j_1,\ldots,j_K}(x)=\{{\mathbf b} \in\mathbb{R}^{Kpd}: {\mathbf b}_{S_{j_1,\ldots,j_K}}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}_{S_{j_1,\ldots,j_K}}\leqslant x\}$ with $S_{j_1,\ldots,j_K}=\{j_1,j_2+pd,\ldots,j_K+(K-1)pd\}$. Define $\mathcal A(x;K) = \bigcap_{j_1=1}^{pd} \cdots \bigcap_{j_K=1}^{pd} \mathcal A_{j_1,\ldots,j_K}(x)$. We then have $\{\tilde T_n\leqslant x\}=\{\tilde{n}^{1/2}\bar{\boldsymbol \eta}\in\mathcal A(x;K)\}$. Note that the set $\mathcal A_{j_1,\ldots,j_K}(x)$ is convex that only depends on the components in $S_{j_1,\ldots,j_K}$. We can reformulate $\mathcal A_{j_1,\ldots,j_K}(x)$ as follows:

align*[align* omitted — 258 chars of source]

Define $\mathcal{F} =\bigcup_{j_1=1}^{pd}\cdots\bigcup_{j_K=1}^{pd}\{{\mathbf a}\in\mathbb{S}^{Kpd-1}:{\mathbf a}_{S_{j_1,\ldots,j_K}}\in\mathbb{S}^{K-1}\}$. Then $ \mathcal A(x;K)=\bigcap_{{\mathbf a}\in\mathcal{F}}\{{\mathbf b}\in\mathbb{R}^{Kpd}: {\mathbf a}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}\leqslant x^{1/2}\}$ and

align[align omitted — 241 chars of source]

for any $x>0$. As indicated in (ref), to construct the Gaussian approximation of $\mathbb{P}_{H_0}(\tilde{T}_n\leqslant x)$, we need to impose the following assumption on the tail behavior of ${\mathbf a}^{\scriptscriptstyle {\rm \top}}\boldsymbol \eta_t$. See also CCK2017 and CCW2020.

conditionThere exist some universal constants $C_1>1$, $C_2>0$ and $\tau_1\in(0,1]$ independent of $(K,p,d,n)$ such that \begin{align*} \sup_{t\in[n]}\sup_{{\mathbf a}\in\cal F}\mathbb{P}(|{\mathbf a}^{{\scriptscriptstyle {\rm \top}}}\boldsymbol \eta_t|>x) \leqslant C_1\exp(-C_2 x^{\tau_1}) \end{align*} for any $x>0$.

Condition (ref) is stronger than necessary for the theoretical justification of our proposed method, and it can be weakened at the expense of much lengthier proofs. For example, Condition (ref) can be replaced by the assumption:

equation[equation omitted — 130 chars of source]

for any $x>0$. Recall $\eta_{t,\ell}=\phi_{l_1}({\mathbf x}_t)x_{t+k,l_2}$ for some $l_1\in[d], l_2\in[p]$ and $k\in[K]$. If $\boldsymbol{\phi}(\cdot)$ is selected as some bounded functions, then (ref) holds provided that $\max_{t\in[n]}\max_{l_2\in[p]}\mathbb{P}(|x_{t,l_2}|>x) \leqslant C_*\exp(-C_{**}x^{\tau_1})$ for any $x>0$. If $\boldsymbol{\phi}(\cdot)$ and ${\mathbf x}_t$ satisfy $\max_{t\in[n]}\max_{l_1\in[d]}\mathbb{P}\{|\phi_{l_1}({\mathbf x}_t)|>x\} \leqslant C_*\exp(-C_{**}x^{\tau_*})$ and $\max_{t\in[n]}\max_{l_2\in[p]}\mathbb{P}(|x_{t,l_2}|>x) \leqslant C_*\exp(-C_{**}x^{\tau_{**}})$ for any $x>0$, by Lemma 2 of CTW2013, we know (ref) holds with $\tau_1=\tau_*\tau_{**}/(\tau_*+\tau_{**})$. For any ${\mathbf a}\in\cal F$, there exists $(j_1,\ldots,j_K)\in[pd]^K$ such that $\sum_{\ell=1}^K a_{j_\ell+(\ell-1)pd}^2=1$ and $a_{j}=0$ for $j\notin S_{j_1,\ldots,j_K}$, which implies $\sum_{\ell=1}^K|a_{j_\ell+(\ell-1)pd}|\leqslant \sqrt{K}$. By Bonferroni inequality and (ref), for any given ${\mathbf a}\in\cal F$, it holds that

align[align omitted — 425 chars of source]

for any $x>0$, which provides a rough upper bound for $\max_{t\in[n]}\sup_{{\mathbf a}\in\mathcal{F}}\mathbb{P}(|{\mathbf a}^{{\scriptscriptstyle {\rm \top}}}\boldsymbol \eta_t|>x)$. When $K$ is a fixed positive integer, by (ref), we know Condition (ref) is satisfied provided that (ref) holds. If we only assume (ref), we can still establish the associated Gaussian approximation results based on (ref) rather than Condition (ref) but the associated arguments will be quite cumbersome.

conditionAssume that $\{{\mathbf x}_t\}$ is $\alpha$-mixing in the sense that \begin{align*} \alpha(k):=\sup_t\sup_{(A,B)\in \mathscr{F}_{-\infty}^t \times \mathscr{F}_{t+k}^{+\infty}} |\mathbb{P}(A\cap B) - \mathbb{P}(A)\mathbb{P}(B)| \to 0 as k\to \infty \,, \end{align*} where $\mathscr{F}_{-\infty}^u$ and $\mathscr{F}_{u+k}^{+\infty}$ are the $\sigma$-fields generated respectively by $\{{\mathbf x}_t\}_{t\leqslant u}$ and $\{{\mathbf x}_t\}_{t\geqslant u+k}$. Furthermore, there exist some universal constants $C_3>1$, $C_4>0$ and $\tau_2\in(0,1]$ independent of $(K,p,d,n)$ such that $ \alpha(k) \leqslant C_3\exp(-C_4 k^{\tau_2})$ for all $k\geqslant 1$.

The $\alpha$-mixing assumption in Condition (ref) is weaker than the $\beta$-mixing assumption considered in CCK2019. Restricting $\tau_2\in(0,1]$ is just to simplify the presentation. If the $\alpha$-mixing coefficients satisfy Condition (ref) with some constant $\tau_2>1$, then Condition (ref) will be satisfied automatically with $\tau_2=1$. Under certain conditions, VAR processes, multivariate ARCH processes, and multivariate GARCH processes all satisfy Condition (ref) with $\tau_2=1$; see Hafner:Preminger:2009, Boussama_etal:2011 and Wong_etal:2020. In addition, if we only require $\sup_{t\in[n]}\sup_{{\mathbf a}\in\cal F} \mathbb{P}(|{\mathbf a}^{{\scriptscriptstyle {\rm \top}}}\boldsymbol \eta_t|>x)=O\{x^{-(\nu+\epsilon)}\}$ for any $x>0$ in Condition 1 and $\alpha(k)=O\{k^{-\nu(\nu+\epsilon)/(2\epsilon)}\}$ for all $k\geqslant 1$ in Condition 2 with some constants $\nu>2$ and $\epsilon>0$, we can also apply the Fuk-Nagaev-type inequalities to construct the upper bounds for the tail probabilities of certain statistics for which our testing procedure still works for $Kpd$ diverging at some polynomial rate of $n$. We refer to Section 3.2 of CGY2018 for the implementation of the Fuk-Nagaev-type inequalities in such a scenario.

conditionThere exists a universal constant $C_5>0$ independent of $(K,p,d,n)$ such that \begin{align*} \inf_{{\mathbf a}\in\cal F} {\rm Var}\bigg(\frac{1}{\sqrt{\tilde{n}}}\sum_{t=1}^{\tilde{n}} {\mathbf a}^{{\scriptscriptstyle {\rm \top}}}\boldsymbol \eta_t \bigg) \geqslant C_5\,. \end{align*}
conditionThe kernel function $\mathcal K(\cdot)$ is continuously differentiable with bounded derivatives on $\mathbb{R}$ satisfying {\rm(i)} $\mathcal K(0)=1$, {\rm(ii)} $\mathcal K(x)=\mathcal K(-x)$ for any $x\in\mathbb{R}$, and {\rm(iii)} $|\mathcal K(x)|\leqslant C_6|x|^{-\vartheta}$ as $|x|\to\infty$ for some universal constants $C_6>0$ and $\vartheta>1$.

Condition (ref) is a mild technical assumption for the validity of the Gaussian approximation which requires the long-run variance of the sequence $\{{\mathbf a}^{\scriptscriptstyle {\rm \top}}\boldsymbol \eta_t\}$ to be non-degenerate. Note that there are no explicit requirements on the cross-series dependence, and both weak and strong cross-series dependence are allowed in our theory. Condition (ref) is commonly used for the nonparametric estimation of the long-run covariance matrix; see NW1987 and Andrews1991. For the kernel functions with bounded support such as Parzen kernel and Bartlett kernel, we have $\vartheta=\infty$ in Condition (ref).

For $\tau_1$ and $\tau_2$ specified in Conditions (ref) and (ref), we define

align[align omitted — 166 chars of source]

Such defined $f_1(\tau_1,\tau_2)$ is used to control the divergence rate of $K$ which is determined from the technical proofs of Gaussian approximation theory. See Proposition (ref) in Section (ref). Notice that $\tau_1, \tau_2\in(0,1]$. When $\tau_1=\tau_2=1$, then $f_1(\tau_1,\tau_2)=1/15$.

Assume that the bandwidth $b_n$ involved in (ref) satisfies $b_n\asymp n^\rho$ for some constant $0<\rho< (\vartheta-1)/(3\vartheta-2)$ with $\vartheta$ specified in Condition (ref). Let

align[align omitted — 138 chars of source]

Such defined $f_2(\rho,\vartheta)$ is also used to control the divergence rate of $K$ which is obtained from the estimation of long-run covariance matrix $\boldsymbol{\Sigma}_{n,K}$. See Proposition (ref) below. For given kernel function $\mathcal{K}(\cdot)$, the parameter $\vartheta$ is determined. Since $\vartheta=\infty$ if $\mathcal{K}(\cdot)$ is selected as the kernel functions with bounded support such as Parzen kernel and Bartlett kernel, then $f_2(\rho,\infty)=\min\{\rho/5,(1-3\rho)/6\}$. For given $\vartheta>1$, the optimal selection of $\rho$ that maximizes $f_2(\rho,\vartheta)$ with respect to $\rho$ is $(5\vartheta-5)/(21\vartheta-13)$ and the associated $f_2(\rho,\vartheta)=(\vartheta-1)/(21\vartheta-13)$.

propositionAssume that Conditions {\rm(ref)}, {\rm(ref)} and {\rm(ref)} hold. Let $b_n\asymp n^\rho$ for some constant $0<\rho<(\vartheta-1)/(3\vartheta-2)$, and $K=O(n^\delta)$ for some constant $0\leqslant \delta<f_2(\rho,\vartheta)$ with $f_2(\rho,\vartheta)$ defined as (ref). Then $ |\widehat{\boldsymbol{\Sigma}}_{n,K}-\boldsymbol{\Sigma}_{n,K}|_\infty =o_{\rm p}[K^{-3}\{\log(npd)\}^{-2}]$ provided that $\log(pd)=o(n^{c})$ for some constant $c>0$ only depending on $(\tau_1,\tau_2,\rho,\vartheta,\delta)$.

Different from the existing literature of high-dimensional covariance matrix estimation, our procedure does not require $\widehat{\boldsymbol{\Sigma}}_{n,K}$ to be consistent under the matrix $L_2$-operator norm and therefore it can work without imposing any structural assumptions on the underlying long-run covariance matrix $\boldsymbol{\Sigma}_{n,K}$. More specifically, our procedure only requires $|\widehat{\boldsymbol{\Sigma}}_{n,K}-\boldsymbol{\Sigma}_{n,K}|_\infty=o_{\rm p}[K^{-3}\{\log(npd)\}^{-2}]$, which is a quite mild requirement and our proposed $\widehat{\boldsymbol{\Sigma}}_{n,K}$ in Section (ref) satisfies this even when $p$ and $d$ grow exponentially with $n$. Now we are ready to present the theoretical guarantees of the testing procedure (ref).

theoremAssume Conditions {\rm (ref)--(ref)} hold. Let $b_n\asymp n^\rho$ for some constant $0<\rho<(\vartheta-1)/(3\vartheta-2)$. Select $K=O(n^\delta)$ for some constant $0\leqslant \delta<\min\{f_1(\tau_1,\tau_2)\,,f_2(\rho,\vartheta)\}$ with $f_1(\tau_1,\tau_2)$ and $f_2(\rho,\vartheta)$ defined as (ref) and (ref), respectively. Then $ \mathbb{P}_{H_0}( T_n > \hat{\rm cv}_\alpha) \to \alpha$ as $n\rightarrow\infty$, provided that $\log(pd)=o(n^{c})$ for some constant $c>0$ only depending on $(\tau_1,\tau_2,\rho,\vartheta,\delta)$.

Theorem (ref) reveals the validity of our proposed test in the sense that the testing procedure maintains the nominal significance level asymptotically under the null hypothesis, where $pd$ is allowed to diverge exponentially with respect to the sample size $n$. In Theorem (ref), the asymptotic power of the proposed tests is analyzed.

theoremAssume the conditions of Theorem {\rm(ref)} hold. Let $\varrho$ be the largest element in the main diagonal of $\boldsymbol{\Sigma}_{n,K}$, and write $\lambda(K,p,d,\alpha)=\{2\log(pd)\}^{1/2} +\{2\log(4K/\alpha)\}^{1/2}$. If $ \sum_{j=1}^{K}|\boldsymbol{\gamma}_j|_\infty^2 \geqslant n^{-1}K\varrho\lambda^2(K,p,d,\alpha)(1+\epsilon_n)^2$ under the alternative hypothesis for some $\epsilon_n>0$ satisfying $\epsilon_n\rightarrow 0$ and $\varrho\lambda^2(K,p,d,\alpha)K^{-1}(\log K)^{-1}\epsilon_n^2\rightarrow\infty$, then $ \mathbb{P}_{H_1}(T_n>\hat{\rm cv}_\alpha)\to 1$ as $n\to\infty$.

Theorem (ref) shows that our proposed test is consistent under local alternatives. Recall $\boldsymbol{\gamma}=(\boldsymbol{\gamma}_1^{{\scriptscriptstyle {\rm \top}}},\ldots,\boldsymbol{\gamma}_K^{{\scriptscriptstyle {\rm \top}}})^{{\scriptscriptstyle {\rm \top}}}$ with each $\boldsymbol{\gamma}_j\in\mathbb{R}^{pd}$. When $K$ is fixed and $\varrho=O(1)$, the latter of which holds under suitable assumptions on the data generating process, the condition that $|\boldsymbol{\gamma}|_\infty \geqslant Cn^{-1/2}\{\log(Kpd)\}^{1/2}$ for some positive constant $C$, is sufficient for $\sum_{j=1}^K|\boldsymbol{\gamma}_j|^2_\infty \geqslant n^{-1}K\varrho\lambda^2(K,p,d,\alpha)(1+\epsilon_n)^2$. As we have discussed in Remark (ref), if the time series $\{{\mathbf x}_t\}$ is strictly stationary, we know the transformed data $\{\boldsymbol \eta_t\}$ is also strictly stationary and the proposed test statistic $T_n$ given in (ref) essentially tests whether $\boldsymbol{\gamma}=\mathbb{E}(\boldsymbol \eta_t)=\boldsymbol{0}$ or not. As shown in Theorem 3 of CLX2013, $n^{-1/2}\{\log(Kpd)\}^{1/2}$ is the minimax optimal separation rate of any tests for the $(Kpd)$-dimensional mean vector hypothesis testing problem $H_0:\boldsymbol{\gamma}=\boldsymbol{0}$ versus $H_1:\boldsymbol{\gamma}\neq\boldsymbol{0}$ based on the data $\{\boldsymbol \eta_t\}_{t=1}^n$ if the smallest eigenvalues of ${\rm Var}(\boldsymbol \eta_t)$ are uniformly bounded away from zero. That is, for any $\alpha,\, \beta>0$ satisfying $\alpha+\beta<1$, there exists a constant $\delta_0>0$ such that $\inf_{\boldsymbol{\gamma}\in\mathcal{M}(\delta_0)}\sup_{\xi_\alpha\in\mathcal{T}_\alpha}\mathbb{P}_{H_1}(\mbox{reject $H_0$ based on $\xi_\alpha$})\leqslant 1-\beta$ for all sufficiently large $n$, $p$ and $d$, where $\mathcal{M}(\delta_0)=\{\boldsymbol{\gamma}\in\mathbb{R}^{Kpd}:|\boldsymbol{\gamma}|_\infty\geqslant \delta_0 n^{-1/2}\{\log(Kpd)\}^{1/2}\}$, and $\mathcal{T}_\alpha$ is the set of all $\alpha$-level tests for the test $H_0:\boldsymbol{\gamma}=\boldsymbol{0}$ versus $H_1:\boldsymbol{\gamma}\neq\boldsymbol{0}$. Hence, if the time series $\{{\mathbf x}_t\}$ is strictly stationary, our proposed testing procedure with fixed $K$ will share some minimax optimal property.

General martingale difference hypothesis and specification testing

Our test procedure can also be extended to a more general martingale difference hypothesis, that is

align[align omitted — 139 chars of source]

where $\boldsymbol \mu_x\in\mathbb{R}^p$ is an unknown vector. In this scenario, we can consider the test statistic

align[align omitted — 110 chars of source]

where $\hat{\boldsymbol{\gamma}}_j^{\rm new} =(n-j)^{-1}\sum_{t=1}^{n-j}{\rm vec}\{\boldsymbol{\phi}({\mathbf x}_t)({\mathbf x}_{t+j}-\bar{{\mathbf x}})^{{\scriptscriptstyle {\rm \top}}}\}$ with $\bar{{\mathbf x}}=n^{-1}\sum_{t=1}^n{\mathbf x}_t$. Write $\mathring{{\mathbf x}}_t={\mathbf x}_t-\boldsymbol \mu_x$. In comparison to $T_n$ given in (ref), we replace ${\mathbf x}_{t+j}$ there by its mean-centered version ${\mathbf x}_{t+j}-\bar{{\mathbf x}}$ in $T_n^{\rm new}$. Notice that

align*[align* omitted — 942 chars of source]

Since $K=o(n)$ and $j\in[K]$, ${\rm I}_j$ is the leading term of $\hat{\boldsymbol{\gamma}}_j^{\rm new}$, and ${\rm II}_j$ and ${\rm III}_j$ are the negligible terms in comparison to ${\rm I}_j$. Define

align*[align* omitted — 625 chars of source]

Write $\tilde{n}=n-K$. If Conditions (ref) and (ref) hold for $\boldsymbol \eta_t^{\rm new}$, together with Condition (ref), we know the null distribution of $T_n^{\rm new}$ can be approximated by that of its Gaussian analogue $G_K^{\rm new} = \sum_{j=1}^{K} |\mathbf g^{\rm new}_{\mathcal{L}_j}|_\infty^2$, where $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$ and $\mathbf g^{\rm new}=(g_1^{\rm new},\ldots,g_{Kpd}^{\rm new})^{\scriptscriptstyle {\rm \top}} \sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_{n,K}^{\rm new})$ with $\boldsymbol{\Sigma}_{n,K}^{\rm new}={\rm Cov} (\tilde{n}^{-1/2}\sum_{t=1}^{\tilde{n}}\boldsymbol \eta_t^{\rm new})$. Write $\hat{\mathring{{\mathbf x}}}_t={\mathbf x}_t-\bar{{\mathbf x}}$ and

align*[align* omitted — 629 chars of source]

Identical to (ref), we can adopt the following estimate for $\boldsymbol{\Sigma}_{n,K}^{\rm new}$:

align*[align* omitted — 172 chars of source]

where $\widehat{\mathbf H}_j^{\rm new}=\tilde{n}^{-1}\sum_{t=j+1}^{\tilde{n}}(\hat\boldsymbol \eta_t^{\rm new}-\bar{\hat{\boldsymbol \eta}}^{\rm new})(\hat\boldsymbol \eta_{t-j}^{\rm new}-\bar{\hat{\boldsymbol \eta}}^{\rm new})$ if $j\geqslant 0$ and $\widehat{\mathbf H}_j^{\rm new} =\tilde{n}^{-1} \sum_{t=-j+1}^{\tilde{n}} (\hat\boldsymbol \eta_{t+j}^{\rm new}-\bar{\hat{\boldsymbol \eta}}^{\rm new})(\hat\boldsymbol \eta_{t}^{\rm new} -\bar{\hat{\boldsymbol \eta}}^{\rm new})$ otherwise, with $\bar{\hat{\boldsymbol \eta}}^{\rm new} =\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\hat\boldsymbol \eta_t^{\rm new}$. Algorithm (ref) states how to implement the proposed general martingale difference hypothesis test in practice.

algorithm[algorithm omitted — 1,713 chars of source]

Below we shall provide some detailed discussion about potential extension of our test to the specification testing framework. Let $\mathbf y_t$ and $\mathbf u_t$ be observable $p$-dimensional and $q$-dimensional time series, respectively. Consider the time series model

equation[equation omitted — 103 chars of source]

where ${\mathbf x}_t$ is the error process, and $\mathbf h(\cdot;\cdot)\in\mathbb{R}^p$ is a known link function with unknown truth $\boldsymbol \theta_0\in\mathbb{R}^m$. Without loss of generality, we assume $\mathbb{E}({\mathbf x}_t\,|\,\mathbf u_t)=\boldsymbol{0}$. Model (ref) is quite general for our analysis where we can select $\mathbf u_t$ as $\mathbf y_{t-1},\ldots,\mathbf y_{t-\ell}$ for some integer $\ell\geqslant1$. For the model diagnosis, we are interested in the hypothesis testing problem:

equation[equation omitted — 172 chars of source]

Based on the conditional moment restrictions $\mathbb{E}({\mathbf x}_t\,|\,\mathbf u_t)=\boldsymbol{0}$, for given basis functions $\boldsymbol \psi(\cdot):\mathbb{R}^{q}\rightarrow\mathbb{R}^l$ with $pl\geqslant m$, we can identify the unknown truth $\boldsymbol \theta_0$ by the $pl$ unconditional moment restrictions

equation*[equation* omitted — 141 chars of source]

where $\otimes$ denotes the Kronecker product.

{\it Case 1}. If $m$ is fixed or diverges slowly with the sample size $n$, applying the estimation procedure suggested in CCC2015, we can obtain a consistent estimator $\hat{\boldsymbol \theta}_n$ for $\boldsymbol \theta_0$ and it admits the following asymptotic expansion:

equation[equation omitted — 163 chars of source]

where $\mathbf w(\cdot)$ is the influence function such that $\mathbb{E}\{\mathbf w(\mathbf y_t,\mathbf u_t)\}=\boldsymbol{0}$. Write $\hat{{\mathbf x}}_t=\mathbf y_t-\mathbf h(\mathbf u_t;\hat{\boldsymbol \theta}_n)$. Together with (ref), it holds that

equation*[equation* omitted — 210 chars of source]

Based on obtained $\{\hat{{\mathbf x}}_t\}_{t=1}^n$, we can propose the following test statistic for (ref):

equation[equation omitted — 108 chars of source]

where $\boldsymbol{\gamma}_j^{\natural}=(n-j)^{-1}\sum_{t=1}^{n-j}{\rm vec}\{\boldsymbol{\phi}(\hat{{\mathbf x}}_t)\hat{{\mathbf x}}_{t+j}^{{\scriptscriptstyle {\rm \top}}}\}$. In comparison to the original test statistic $T_n$ given in (3) based on observed $\{{\mathbf x}_t\}_{t=1}^n$, we replace ${\mathbf x}_t$ there by its estimate $\hat{{\mathbf x}}_t$. By Taylor expansion, under some regularity conditions, it holds that

align*[align* omitted — 276 chars of source]

where $\mathbf A_j=(n-j)^{-1}\sum_{t=1}^{n-j}\mathbb{E}\{{\mathbf x}_{t+j}\otimes[\nabla_{{\mathbf x}}\boldsymbol{\phi}({\mathbf x}_t)\nabla_{\boldsymbol \theta}\mathbf h(\mathbf u_t;\boldsymbol \theta_0)] +[\nabla_{\boldsymbol \theta}\mathbf h(\mathbf u_{t+j};\boldsymbol \theta_0)]\otimes\boldsymbol{\phi}({\mathbf x}_t)\}$. Define

align*[align* omitted — 434 chars of source]

Recall $\tilde{n}=n-K$. Following the same arguments in Section 2.1, the null distribution of $T_n^{\natural}$ can be approximated by that of its Gaussian analogue $ G_K ^{\natural}= \sum_{j=1}^{K} |\mathbf g_{\mathcal{L}_j}^{\natural}|_\infty^2$, where $\mathcal{L}_j=\{(j-1)pd+1,\ldots,jpd\}$ and $\mathbf g^{\natural}=( g_1^{\natural},\ldots, g_{Kpd}^{\natural})^{\scriptscriptstyle {\rm \top}} \sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_{n,K}^{\natural})$ with $\boldsymbol{\Sigma}_{n,K}^{\natural}={\rm Cov} (\tilde{n}^{-1/2}\sum_{t=1}^{\tilde{n}}\boldsymbol \eta_t^{\natural})$. The key challenge here is to construct a valid estimate $\widehat{\boldsymbol{\Sigma}}_{n,K}^{\natural}$ satisfying $|\widehat{\boldsymbol{\Sigma}}_{n,K}^{\natural}-{\boldsymbol{\Sigma}}_{n,K}^{\natural}|_\infty=o_{{\rm p}}[K^{-3}\{\log(npd)\}^{-2}]$ with unknown $\mathbf A_1,\ldots,\mathbf A_K$ and unobserved $\{{\mathbf x}_t\}$.

{\it Case 2}. If $m\gg n$, we need to assume the unknown truth $\boldsymbol \theta_0=(\theta_{0,1},\ldots,\theta_{0,m})^{{\scriptscriptstyle {\rm \top}}}$ in (ref) is sparse. Let $\mathcal{S}=\{k\in[m]:\theta_{0,k}\neq0\}$. Using the penalized estimation procedure, for example, CTW2018, we can obtain a sparse estimate $\hat{\boldsymbol \theta}_n$ for $\boldsymbol \theta_0$ satisfying the oracle property: (i) $\mathbb{P}(\hat{\boldsymbol \theta}_{n,\mathcal{S}^{\rm c}}=\boldsymbol{0})\rightarrow1$ as $n\rightarrow\infty$, and (ii) $\hat{\boldsymbol \theta}_{n,\mathcal{S}}$ follows the asymptotic expansion:

equation[equation omitted — 219 chars of source]

where $\tilde{\mathbf w}(\cdot)$ is the influence function such that $\mathbb{E}\{\tilde{\mathbf w}(\mathbf y_t,\mathbf u_t)\}=\boldsymbol{0}$, and $\boldsymbol \xi_n$ is the asymptotic bias satisfying $|\boldsymbol \xi_n|_\infty=O_{{\rm p}}(\delta_n)$ for some $\delta_n=o(1)$ but $\delta_n\gg n^{-1/2}$. To propose the testing procedure in the setting with $m\gg n$, we need to do the next three steps first: (a) identify the index set $\mathcal{S}$, (b) estimate the asymptotic bias $\boldsymbol \xi_n$, (c) obtain the bias-corrected estimate $\tilde{\boldsymbol \theta}_n$ for $\boldsymbol \theta_0$ based on the estimate of $\boldsymbol \xi_n$. Write $\hat{{\mathbf x}}_t=\mathbf y_t-\mathbf h(\mathbf u_t;\tilde{\boldsymbol \theta}_n)$. We can still use the test statistic $T_n^{\natural}$ given in (ref) in current setting. To determine the associated critical value, we only need to replace $\mathbf w(\cdot)$ and $\nabla_{\boldsymbol \theta}\mathbf h(\cdot;\boldsymbol \theta_0)$ by $\tilde{\mathbf w}(\cdot)$ and $\nabla_{\boldsymbol \theta_{\mathcal{S}}}\mathbf h(\cdot;\boldsymbol \theta_0)$, respectively, in the procedure for the setting with fixed or slowly diverging $m$. However, as commented in CCTW2021, if $\mathbf h(\cdot;\boldsymbol \theta)$ is a nonlinear function of $\boldsymbol \theta$, the asymptotic bias $\boldsymbol \xi_n$ may include some unknown information which makes the estimation of $\boldsymbol \xi_n$ extremely difficult (if not impossible). How to address this problem requires further study.

Simulation studies

In this section, we examine the finite sample performance of our proposed test in comparison with the ones proposed by HLZ2017. All tests in our simulation are implemented at the $5\%$ significance level using $4000$ Monte Carlo replications, and the number of bootstrap replications used to determine the critical value $\hat{{\rm cv}}_\alpha$ in our procedure is chosen as $B=2000$. We set the sample size $n\in\{100,300\}$ and lags $K\in\{2,4,6,8\}$. The dimension $p$ is set according to the ratio $p/n \in \{0.04, 0.08, 0.15, 0.4, 1.2\}$, which covers low-, moderate- and high-dimensional scenarios. Two types of maps are considered, i.e., (i) linear function ($d=p$), $\boldsymbol{\phi}({\mathbf x}_t)={\mathbf x}_t$; (ii) both linear and quadratic functions ($d=2p$), $\boldsymbol{\phi}({\mathbf x}_t)=\{{\mathbf x}_t^{{\scriptscriptstyle {\rm \top}}}, ({\mathbf x}_t^2)^{{\scriptscriptstyle {\rm \top}}}\}^{{\scriptscriptstyle {\rm \top}}}$. Furthermore, we use three kernel functions for the estimation of long-run covariance matrix $\boldsymbol{\Sigma}_{n,K}$, i.e.,

itemize• Quadratic Spectral (QS) kernel: $\mathcal{K}_{\rm{QS}}(x)=25(12\pi^2x^2)^{-1}\{ (6\pi x/5)^{-1}\sin(6\pi x/5) - \cos(6\pi x/5)\}$. • Parzen (PR) kernel: $\mathcal{K}_{\rm{PR}}(x)=(1-6x^2+6|x|^3)I(0\leqslant|x|\leqslant 1/2)+2(1-|x|)^3I(1/2 < |x|\leqslant 1)$. • Bartlett (BT) kernel: $\mathcal{K}_{\rm{BT}}(x)=(1-|x|)I(|x|\leqslant 1)$.

Recall $\tilde{n}=n-K$. We use the data-driven bandwidth formulas developed in Andrews1991 to determine the associated bandwidth $b_n$ involved in these three kernel functions, that is, $ b_{{\rm QS}}=1.3221\{\hat a(2) \tilde{n}\}^{1/5}$, $b_{{\rm PR}}=2.6614\{\hat a(2) \tilde{n}\}^{1/5}$ and $b_{{\rm BT}}=1.1447\{\hat a(1) \tilde{n}\}^{1/3}$, where $\hat a(2)=\{\sum_{\ell=1}^{Kpd} 4\hat\rho_\ell^2 \hat\sigma_\ell^4 (1-\hat\rho_\ell)^{-8}\}\{\sum_{\ell=1}^{Kpd}\hat\sigma_\ell^4(1-\hat\rho_\ell)^{-4}\}^{-1}$ and $\hat a(1)=\{\sum_{\ell=1}^{Kpd} 4\hat\rho_\ell^2 \hat\sigma_\ell^4 (1-\hat\rho_\ell)^{-6}(1+\hat\rho_\ell)^{-2}\} \{\sum_{\ell=1}^{Kpd}\hat\sigma_\ell^4(1-\hat\rho_\ell)^{-4}\}^{-1}$, with $\hat\rho_\ell$ and $\hat\sigma_\ell^2$ being, respectively, the estimated autoregressive coefficient and innovation variance from fitting an AR(1) model to time series $\{\eta_{t,\ell}\}_{t=1}^{\tilde{n}}$, the $\ell$th component sequence of $\{\boldsymbol \eta_t\}_{t=1}^{\tilde{n}}$ defined in (ref). Denote the test statistics based on the three kernels with linear map by $T_{{\rm{QS}}}^{l}$, $T_{{\rm{PR}}}^{l}$ and $T_{{\rm{BT}}}^{l}$, respectively, and denote the ones with both linear and quadratic map by $T_{{\rm{QS}}}^{q}$, $T_{{\rm{PR}}}^{q}$ and $T_{{\rm{BT}}}^{q}$, respectively. Note that the data-driven formulas by Andrews1991 are based on AR(1) model assumption and also deliver an estimation-optimal bandwidth in the low-dimensional setting. Here we apply it to determine the associated bandwidth $b_n$ in both moderate- and high-dimensional settings since there are no other known formulas and the numerical studies in CYZ2017 show such formula seems to work well when the dimension is large. We also include three tests proposed by HLZ2017 in our simulation comparison, i.e., the trace-based test $Z_{\rm{tr}}$, the determinant-based test $Z_{\rm{det}}$, and the large-dimensional test $Zd_{\rm tr}$. Note that HLZ2017 only examined the finite sample performance of $Z_{\rm{tr}}$ and $Z_{\rm{det}}$, which cannot be implemented when $p>\sqrt{n}$, whereas $Zd_{\rm tr}$ is shown to be valid under the assumption $p/n\rightarrow 0$ and its implementation becomes infeasible when $p>n$. The tests of HLZ2017 require the matrix normalization which is computationally prohibitive in the high-dimensional setting. See Section (ref) in the supplementary material for the comparison of computational cost between our test and the tests of HLZ2017.

Empirical size

To examine the empirical size, we consider the following models:

itemize[leftmargin=1.8cm] • i.i.d. normal sequence: ${\mathbf x}_t\overset{{\rm i.i.d.}}{\sim} \mathcal{N}(\boldsymbol{0}, \mathbf A)$ where $\mathbf A=(a_{kl})_{p\times p}$ with $a_{kl}=0.995^{|k-l|}$ for any $k,l\in [p]$. • Stochastic volatility model: ${\mathbf x}_t=\boldsymbol{\varepsilon}_t\exp(\boldsymbol{\sigma}_t)$ with $\boldsymbol{\sigma}_t=0.25\boldsymbol{\sigma}_{t-1}+0.05{\bf u}_t$, $\boldsymbol{\varepsilon}_t\overset{{\rm i.i.d.}}{\sim}{\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Omega}_\varepsilon)$ and ${\bf u}_t\overset{{\rm i.i.d.}}{\sim}{\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Omega}_u)$, where $\boldsymbol{\Omega}_\varepsilon=(\omega_{\varepsilon,kl})_{p\times p}$ and $\boldsymbol{\Omega}_{u}=(\omega_{u,kl})_{p\times p}$ with $\omega_{\varepsilon,kl}=I(k=l)+0.4I(k\neq l)$ and $\omega_{u,kl}=0.9^{|k-l|}$ for any $k,l\in[p]$. • Bivariate constant conditional correlation GARCH(1,1) model: $ {\mathbf x}_t ={\bf b}_t^{1/2}\circ \boldsymbol{\varepsilon}_t$ with ${\bf b}_t = {{\mathbf a}}_0 + \mathbf A_1{\bf b}_{t-1} + \mathbf A_2{\mathbf x}_{t-1}^2$ and $\boldsymbol{\varepsilon}_t \overset{{\rm i.i.d.}}{\sim} {\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Omega}_\varepsilon)$, where $\circ$ denotes the Hadamard product, ${\mathbf a}_0=(0.2,0.1\,{\bf 1}_{p-1}^{{\scriptscriptstyle {\rm \top}}})^{{\scriptscriptstyle {\rm \top}}}$, $\mathbf A_1=0.9\,{\bf I}_p$, $\mathbf A_2={\rm diag}(0.05,0.08,0.03\, {\bf 1}_{p-2}^{{\scriptscriptstyle {\rm \top}}})$, and $\boldsymbol{\Omega}_\varepsilon=(\omega_{\varepsilon,kl})_{p\times p}$ with $\omega_{\varepsilon,kl}=I(k=l)+0.5I(k\neq l)$ for any $k,l\in[p]$. Here ${\bf 1}_q$ and ${\bf I}_q$ denote, respectively, the $q$-dimensional vector with all components being $1$ and $q$-dimensional identity matrix for any given integer $q$.

A few comments are in order. Model 1 was used by CYZ2017 in their simulation for high-dimensional white noise testing problem. Model 2 is the multivariate extension of the univariate stochastic volatility model considered in EV2006a for the univariate martingale difference hypothesis testing problem. Model 3 is motivated from HLZ2017, which reduces to the bivariate GARCH model considered in HLZ2017 when $p=2$.

As seen from Table (ref), our tests have quite accurate size when the dimension $p$ is low for all models. For a fixed sample size $n$, the rejection rates tend to decrease as the dimension $p$ increases, showing the impact on the bootstrap-based approximation from the dimension $p$. For a fixed dimension $p$, enlarging sample size from $n=100$ to $n=300$ helps to bring down the size distortion to some extent for most kernels and maps, e.g., the empirical sizes for Model 1--3 are undersized when $n=100$ and $p/n=1.2~(p=120)$, and the empirical sizes increase and become much closer to the 5% nominal level when $n=300$ and $p/n=0.4~(p=120)$. Overall our tests show reasonably good size control and the undersize phenomenon for the moderate- and high-dimensional scenarios could be due to the bandwidth choice, which is always a difficult issue in practice. The three tests of HLZ2017 also show quite accurate size for Models 1 and 2, and there is some noticeable over-rejection for Model 3 when $n=100$. When $n=300$ and $p/n=0.4$, we are unable to implement the test $Zd_{\rm tr}$ even though $p<n$. The reason is that the computation of $Zd_{\rm tr}$ requires to store five $120^2\times 120^2$ matrices, and product of three $120^2\times 120^2$ matrices during the calculation, which results in running out of the memory (RAM: 8158 MB). This indicates the difficulty of implementing their tests for $p=120$ and beyond.

In order to investigate the influence of the data-driven bandwidth used in our simulation, we examine the sensitivity of our size and power results by replacing the data-driven bandwidth $b_n$ by its scaled version $c\cdot b_n$ with $c\in\{2^{-3},2^{-2},2^{-1},2^1,2^2,2^3\}$. Simulation results for Bartlett kernel are displayed in Tables (ref) and (ref). Simulation results for Quadratic Spectral kernel and Parzen kernel are reported in the supplementary material. For different multiplies $c$, the sizes and powers are relatively robust. In addition, we find that the results for $c<1$ perform a little better than these for $c>1$ in general, but not by much. Therefore, the choice of $c=1$ in our simulation is reasonable.

sidewaystable[htp] \scriptsize \caption{Empirical sizes ($\%$) of the tests $T_{\rm QS}^l$, $T_{\rm PR}^l$, $T_{\rm BT}^l$, $T_{\rm QS}^q$, $T_{\rm PR}^q$, $T_{\rm BT}^q$, $Z_{\rm tr}$, $Z_{\rm det}$ and $Zd_{\rm tr}$ for Models 1--3 at the 5% nominal level. } \resizebox{!}{6.3cm}{ \begin{tabular}{ccc|ccccccccc| ccccccccc| ccccccccc} & & & \multicolumn{9}{c|}{Model 1} & \multicolumn{9}{c|}{Model 2} & \multicolumn{9}{c}{Model 3} \\[0.6em] $n$ & $p/n$ & $K$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ \\[0.5em] \hline 100 & 0.04 & 2 & 4.2 & 4.5 & 4.5 & 4.3 & 4.3 & 4.4 & 5.2 & 4.4 & 5.5 & 4.2 & 4.4 & 4.7 & 2.2 & 2.4 & 2.5 & 5.2 & 4.7 & 4.8 & 3.5 & 3.6 & 4.2 & 2.9 & 2.8 & 3.2 & 6.7 & 6.3 & 5.2 \\ & & 4 & 5.1 & 5.0 & 5.2 & 4.6 & 4.3 & 4.5 & 5.2 & 5.2 & 5.4 & 3.1 & 3.3 & 3.5 & 2.9 & 2.8 & 3.2 & 4.3 & 4.9 & 6.3 & 3.2 & 3.2 & 3.5 & 3.0 & 3.1 & 3.4 & 6.9 & 6.9 & 6.3 \\ & & 6 & 4.6 & 4.4 & 4.8 & 4.5 & 4.5 & 4.6 & 5.0 & 4.7 & 6.0 & 3.0 & 2.9 & 3.5 & 2.8 & 2.7 & 2.9 & 5.4 & 4.4 & 6.0 & 3.1 & 3.1 & 3.6 & 3.9 & 4.0 & 4.1 & 6.5 & 6.8 & 6.1 \\ & & 8 & 4.4 & 4.4 & 4.7 & 5.0 & 4.9 & 5.1 & 5.6 & 4.9 & 6.2 & 2.8 & 2.8 & 3.3 & 2.9 & 2.8 & 3.1 & 5.4 & 4.5 & 6.5 & 3.3 & 3.3 & 4.0 & 4.5 & 4.4 & 4.9 & 6.7 & 7.9 & 6.9 \\[0.3em] & 0.08 & 2 & 4.1 & 4.1 & 4.1 & 3.1 & 3.2 & 3.4 & 4.8 & 4.9 & 5.2 & 3.8 & 3.9 & 4.0 & 1.9 & 1.8 & 1.8 & 4.4 & 4.9 & 5.2 & 3.3 & 3.2 & 3.5 & 2.7 & 2.5 & 2.7 & 5.9 & 8.0 & 5.2 \\ & & 4 & 3.9 & 3.8 & 4.0 & 4.4 & 4.3 & 4.7 & 6.0 & 4.8 & 5.2 & 3.0 & 3.0 & 3.5 & 1.8 & 1.9 & 2.1 & 5.4 & 5.1 & 6.0 & 2.6 & 2.5 & 3.0 & 2.7 & 2.7 & 2.7 & 7.1 & 8.7 & 5.2 \\ & & 6 & 4.6 & 4.4 & 4.6 & 4.4 & 4.4 & 4.7 & 6.7 & 6.3 & 6.2 & 2.2 & 2.1 & 2.5 & 2.2 & 2.2 & 2.4 & 5.5 & 5.9 & 5.6 & 2.2 & 2.3 & 2.8 & 4.4 & 4.2 & 4.5 & 7.5 & 8.2 & 7.1 \\ & & 8 & 4.8 & 4.9 & 5.2 & 4.0 & 3.9 & 4.2 & 7.4 & 5.7 & 5.4 & 2.3 & 2.0 & 3.0 & 2.1 & 2.2 & 2.4 & 7.2 & 5.5 & 6.9 & 2.8 & 2.8 & 3.3 & 4.4 & 4.4 & 4.9 & 8.2 & 8.9 & 7.6 \\[0.3em] & 0.15 & 2 & 4.2 & 4.4 & 4.4 & 4.0 & 3.8 & 4.2 & NA & NA & 4.6 & 3.4 & 3.4 & 3.7 & 1.9 & 1.9 & 2.3 & NA & NA & 4.8 & 3.4 & 3.4 & 4.1 & 1.9 & 1.9 & 1.9 & NA & NA & 5.1 \\ & & 4 & 4.3 & 4.2 & 4.6 & 3.2 & 3.2 & 3.5 & NA & NA & 4.3 & 2.5 & 2.5 & 2.7 & 1.7 & 1.8 & 2.1 & NA & NA & 5.9 & 2.2 & 2.2 & 2.7 & 2.5 & 2.4 & 2.7 & NA & NA & 5.8 \\ & & 6 & 4.2 & 4.0 & 4.4 & 3.8 & 3.8 & 4.0 & NA & NA & 5.2 & 2.7 & 2.9 & 3.3 & 1.7 & 1.6 & 1.9 & NA & NA & 5.7 & 2.2 & 2.1 & 2.8 & 2.8 & 2.7 & 3.0 & NA & NA & 7.4 \\ & & 8 & 4.0 & 4.1 & 4.5 & 4.4 & 4.5 & 4.8 & NA & NA & 6.3 & 2.7 & 2.6 & 3.1 & 2.1 & 2.2 & 2.4 & NA & NA & 6.4 & 1.8 & 1.8 & 2.6 & 3.4 & 3.6 & 3.6 & NA & NA & 8.3 \\[0.3em] & 0.40 & 2 & 3.8 & 4.0 & 4.1 & 2.1 & 2.3 & 2.6 & NA & NA & 4.8 & 3.4 & 3.5 & 3.9 & 2.0 & 2.3 & 2.5 & NA & NA & 4.7 & 2.7 & 2.5 & 3.0 & 1.8 & 1.8 & 1.9 & NA & NA & 5.5 \\ & & 4 & 2.8 & 2.9 & 3.2 & 2.5 & 2.6 & 2.6 & NA & NA & 5.3 & 3.2 & 3.2 & 3.6 & 2.1 & 2.1 & 2.3 & NA & NA & 5.4 & 1.7 & 1.7 & 2.1 & 2.0 & 2.2 & 1.9 & NA & NA & 6.6 \\ & & 6 & 2.9 & 3.0 & 3.4 & 2.8 & 2.8 & 3.2 & NA & NA & 5.6 & 2.6 & 2.6 & 3.1 & 2.1 & 2.0 & 2.1 & NA & NA & 6.0 & 1.1 & 1.1 & 1.7 & 2.4 & 2.6 & 2.3 & NA & NA & 7.9 \\ & & 8 & 3.2 & 3.0 & 3.4 & 3.2 & 3.1 & 3.4 & NA & NA & 5.8 & 2.8 & 2.8 & 3.2 & 2.6 & 2.5 & 2.8 & NA & NA & 5.2 & 1.5 & 1.4 & 2.2 & 3.3 & 3.6 & 3.4 & NA & NA & 9.2 \\[0.3em] & 1.20 & 2 & 2.2 & 2.3 & 2.5 & 1.1 & 1.2 & 1.3 & NA & NA & NA & 3.8 & 3.8 & 4.3 & 2.6 & 2.8 & 2.7 & NA & NA & NA & 1.6 & 1.6 & 2.0 & 2.9 & 3.3 & 2.5 & NA & NA & NA \\ & & 4 & 2.0 & 2.1 & 2.6 & 1.1 & 1.2 & 1.2 & NA & NA & NA & 2.9 & 3.1 & 3.2 & 2.5 & 2.5 & 2.7 & NA & NA & NA & 1.1 & 1.1 & 1.8 & 3.9 & 4.1 & 3.1 & NA & NA & NA \\ & & 6 & 2.0 & 2.1 & 2.4 & 1.1 & 1.1 & 1.3 & NA & NA & NA & 3.3 & 3.3 & 3.9 & 2.3 & 2.4 & 2.4 & NA & NA & NA & 1.1 & 1.1 & 1.5 & 4.9 & 5.3 & 3.9 & NA & NA & NA \\ & & 8 & 1.4 & 1.7 & 2.0 & 1.2 & 1.2 & 1.4 & NA & NA & NA & 3.1 & 3.2 & 3.8 & 2.9 & 3.1 & 3.2 & NA & NA & NA & 1.1 & 1.0 & 1.5 & 6.3 & 6.7 & 5.5 & NA & NA & NA \\[0.3em] \hline 300 & 0.04 & 2 & 5.6 & 5.5 & 5.8 & 4.2 & 4.1 & 4.4 & 5.1 & 5.5 & 5.8 & 4.1 & 4.1 & 4.2 & 3.8 & 3.7 & 3.9 & 4.9 & 5.6 & 4.7 & 4.0 & 4.0 & 4.0 & 3.6 & 3.4 & 3.8 & 6.2 & 7.5 & 5.2 \\ & & 4 & 3.9 & 4.2 & 4.5 & 4.7 & 4.6 & 5.0 & 5.9 & 5.4 & 5.0 & 3.7 & 3.9 & 4.2 & 2.9 & 2.8 & 3.2 & 6.1 & 5.9 & 5.6 & 3.8 & 3.7 & 4.1 & 3.3 & 3.4 & 3.9 & 6.2 & 6.6 & 5.5 \\ & & 6 & 4.2 & 4.1 & 4.2 & 5.5 & 5.2 & 5.5 & 6.4 & 6.7 & 5.6 & 3.9 & 3.6 & 3.9 & 3.9 & 4.0 & 4.2 & 6.6 & 6.8 & 6.4 & 3.7 & 3.7 & 4.1 & 4.7 & 4.7 & 4.9 & 7.3 & 7.9 & 5.6 \\ & & 8 & 4.7 & 4.8 & 5.0 & 6.0 & 6.0 & 6.3 & 7.1 & 6.9 & 5.8 & 3.7 & 3.8 & 4.0 & 4.1 & 4.0 & 4.3 & 7.1 & 6.4 & 5.1 & 3.2 & 3.0 & 3.4 & 4.4 & 4.4 & 4.8 & 8.6 & 8.0 & 6.3 \\[0.3em] & 0.08 & 2 & 4.8 & 4.8 & 5.0 & 4.0 & 4.0 & 4.1 & NA & NA & 5.5 & 4.2 & 4.3 & 4.4 & 3.5 & 3.6 & 3.8 & NA & NA & 4.8 & 3.6 & 3.5 & 3.8 & 3.2 & 3.2 & 3.2 & NA & NA & 5.6 \\ & & 4 & 3.8 & 3.8 & 3.9 & 4.1 & 4.0 & 4.2 & NA & NA & 5.2 & 3.8 & 3.5 & 3.8 & 3.7 & 3.6 & 3.9 & NA & NA & 5.0 & 3.7 & 3.5 & 4.0 & 3.2 & 3.0 & 3.4 & NA & NA & 5.4 \\ & & 6 & 4.6 & 4.4 & 5.0 & 5.0 & 5.0 & 5.4 & NA & NA & 5.0 & 3.6 & 3.3 & 3.8 & 4.1 & 4.2 & 4.4 & NA & NA & 5.4 & 3.3 & 3.2 & 3.7 & 3.4 & 3.3 & 3.7 & NA & NA & 5.3 \\ & & 8 & 3.9 & 4.2 & 4.3 & 5.7 & 5.9 & 6.1 & NA & NA & 5.4 & 3.7 & 3.6 & 4.1 & 3.8 & 3.7 & 4.0 & NA & NA & 5.4 & 3.0 & 3.0 & 3.4 & 3.6 & 3.7 & 4.1 & NA & NA & 5.9 \\[0.3em] & 0.15 & 2 & 4.7 & 4.6 & 4.8 & 3.8 & 3.9 & 4.3 & NA & NA & 6.0 & 4.4 & 4.4 & 4.7 & 3.7 & 3.6 & 3.9 & NA & NA & 4.6 & 3.9 & 4.0 & 4.2 & 3.1 & 2.9 & 3.3 & NA & NA & 5.0 \\ & & 4 & 4.4 & 4.4 & 4.6 & 4.4 & 4.3 & 4.6 & NA & NA & 4.6 & 3.7 & 3.9 & 3.9 & 4.2 & 4.2 & 4.4 & NA & NA & 5.1 & 3.2 & 3.2 & 3.6 & 3.2 & 3.0 & 3.3 & NA & NA & 5.3 \\ & & 6 & 3.9 & 3.9 & 4.1 & 4.6 & 4.4 & 4.8 & NA & NA & 5.1 & 3.5 & 3.5 & 3.8 & 3.4 & 3.7 & 3.8 & NA & NA & 5.3 & 3.2 & 3.0 & 3.4 & 3.6 & 3.5 & 3.8 & NA & NA & 6.8 \\ & & 8 & 3.9 & 4.0 & 4.2 & 4.4 & 4.4 & 4.7 & NA & NA & 5.6 & 3.5 & 3.5 & 3.7 & 4.2 & 4.3 & 4.4 & NA & NA & 5.6 & 3.0 & 3.0 & 3.4 & 3.6 & 3.4 & 4.1 & NA & NA & 5.9 \\[0.3em] & 0.40 & 2 & 4.2 & 4.2 & 4.3 & 2.6 & 2.6 & 2.8 & NA & NA & NA & 4.5 & 4.6 & 4.8 & 3.1 & 3.0 & 3.4 & NA & NA & NA & 3.8 & 3.8 & 4.1 & 2.8 & 2.7 & 3.1 & NA & NA & NA \\ & & 4 & 3.5 & 3.5 & 3.6 & 3.2 & 3.2 & 3.5 & NA & NA & NA & 4.2 & 4.3 & 4.5 & 3.8 & 3.7 & 4.0 & NA & NA & NA & 3.1 & 3.1 & 3.5 & 2.7 & 2.6 & 3.0 & NA & NA & NA \\ & & 6 & 3.7 & 3.9 & 4.2 & 4.1 & 4.1 & 4.7 & NA & NA & NA & 4.1 & 4.0 & 4.4 & 3.9 & 4.0 & 4.2 & NA & NA & NA & 2.7 & 2.6 & 3.0 & 2.4 & 2.3 & 2.7 & NA & NA & NA \\ & & 8 & 3.2 & 3.3 & 3.8 & 4.1 & 4.0 & 4.6 & NA & NA & NA & 4.1 & 4.1 & 4.2 & 4.2 & 4.2 & 4.4 & NA & NA & NA & 2.3 & 2.3 & 2.8 & 2.8 & 2.9 & 3.1 & NA & NA & NA \\[0.3em] & 1.20 & 2 & 3.1 & 3.0 & 3.4 & 1.8 & 1.7 & 2.0 & NA & NA & NA & 4.0 & 4.2 & 4.2 & 3.9 & 3.9 & 4.0 & NA & NA & NA & 3.8 & 3.8 & 4.0 & 2.4 & 2.4 & 2.8 & NA & NA & NA \\ & & 4 & 2.3 & 2.2 & 2.4 & 1.8 & 1.8 & 2.0 & NA & NA & NA & 3.8 & 3.9 & 4.0 & 3.9 & 3.9 & 4.2 & NA & NA & NA & 2.7 & 2.7 & 3.1 & 1.8 & 1.9 & 2.0 & NA & NA & NA \\ & & 6 & 1.3 & 1.2 & 1.7 & 1.7 & 1.7 & 1.9 & NA & NA & NA & 3.8 & 3.5 & 3.9 & 4.2 & 4.4 & 4.7 & NA & NA & NA & 2.1 & 2.2 & 2.6 & 2.1 & 2.1 & 2.2 & NA & NA & NA \\ & & 8 & 1.1 & 1.3 & 1.8 & 1.7 & 1.8 & 2.1 & NA & NA & NA & 4.2 & 4.4 & 4.6 & 4.0 & 3.9 & 4.2 & NA & NA & NA & 1.8 & 1.7 & 2.4 & 2.0 & 2.0 & 2.5 & NA & NA & NA \\ \end{tabular} }
sidewaystable[htp] \scriptsize \caption{Empirical sizes ($\%$) of the tests $T_{\rm BT}^l$ and $T_{\rm BT}^q$ for Models 1--3 at the 5% nominal level, where $c$ represents the constant which is multiplied by Andrews' bandwidth. } \resizebox{!}{5.7cm}{ \begin{tabular}{ccc|ccccccc|ccccccc|ccccccc|ccccccc|ccccccc|ccccccc} & & & \multicolumn{7}{c|}{Model 1 with $T_{\rm BT}^l$} & \multicolumn{7}{c|}{Model 1 with $T_{\rm BT}^q$} & \multicolumn{7}{c|}{Model 2 with $T_{\rm BT}^l$} & \multicolumn{7}{c|}{Model 2 with $T_{\rm BT}^q$} & \multicolumn{7}{c|}{Model 3 with $T_{\rm BT}^l$} & \multicolumn{7}{c}{Model 3 with $T_{\rm BT}^q$} \\[0.4em] \hline & & & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c}{$c$} \\[0.2em] $n$ & $p/n$ & $K$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ \\[0.3em] \hline 100 & 0.04 & 2 & 4.6 & 4.9 & 4.7 & 4.3 & 5.3 & 6.1 & 9.4 & 4.2 & 4.3 & 4.2 & 4.0 & 4.3 & 4.5 & 7.9 & 4.2 & 4.4 & 4.1 & 4.0 & 4.0 & 4.1 & 5.5 & 3.0 & 3.0 & 2.8 & 2.5 & 2.4 & 2.8 & 4.3 & 4.1 & 4.0 & 3.6 & 3.6 & 3.7 & 3.9 & 5.0 & 3.9 & 3.7 & 2.8 & 2.8 & 2.9 & 3.0 & 3.9 \\ & & 4 & 5.1 & 5.0 & 4.8 & 4.5 & 4.3 & 5.3 & 6.6 & 5.0 & 4.2 & 4.4 & 4.6 & 3.8 & 4.8 & 6.4 & 4.3 & 4.3 & 3.4 & 3.8 & 2.9 & 2.4 & 3.4 & 2.8 & 3.5 & 3.1 & 3.4 & 2.3 & 2.2 & 2.7 & 4.4 & 4.4 & 3.8 & 3.5 & 2.6 & 2.5 & 2.7 & 3.9 & 4.0 & 4.0 & 3.2 & 2.9 & 3.0 & 3.4 \\ & & 6 & 5.1 & 5.0 & 4.1 & 4.8 & 4.5 & 4.1 & 6.0 & 5.3 & 5.0 & 4.9 & 5.1 & 5.1 & 4.5 & 5.8 & 3.8 & 4.0 & 3.2 & 2.8 & 2.0 & 2.1 & 2.1 & 3.1 & 3.2 & 3.5 & 3.0 & 2.9 & 2.3 & 2.7 & 4.2 & 4.3 & 4.0 & 3.8 & 2.8 & 1.9 & 2.2 & 4.4 & 4.6 & 4.5 & 4.2 & 3.7 & 3.6 & 3.6 \\ & & 8 & 5.7 & 5.3 & 4.9 & 5.4 & 4.5 & 4.1 & 4.0 & 5.3 & 5.6 & 5.0 & 6.1 & 4.8 & 5.5 & 5.1 & 3.5 & 3.6 & 4.1 & 3.2 & 2.3 & 2.0 & 1.5 & 3.7 & 3.3 & 3.5 & 3.7 & 3.3 & 3.1 & 3.2 & 3.8 & 4.1 & 4.7 & 3.9 & 2.7 & 1.9 & 1.5 & 6.0 & 5.8 & 5.3 & 5.2 & 4.0 & 4.3 & 4.2 \\[0.3em] & 0.08 & 2 & 4.9 & 4.2 & 4.3 & 4.6 & 4.9 & 5.9 & 9.2 & 4.2 & 4.3 & 4.1 & 4.2 & 4.3 & 5.2 & 7.2 & 4.3 & 4.0 & 3.4 & 4.1 & 3.5 & 4.1 & 5.7 & 2.5 & 2.6 & 2.2 & 2.2 & 2.3 & 2.2 & 3.7 & 4.0 & 3.8 & 4.2 & 3.4 & 3.2 & 4.0 & 4.5 & 3.1 & 3.3 & 3.0 & 2.7 & 3.0 & 2.5 & 4.3 \\ & & 4 & 5.4 & 4.8 & 5.1 & 4.2 & 3.7 & 4.6 & 5.0 & 4.5 & 4.0 & 4.4 & 4.6 & 3.8 & 3.9 & 5.7 & 3.9 & 3.5 & 3.5 & 3.0 & 2.3 & 2.1 & 2.5 & 2.8 & 2.7 & 1.8 & 2.0 & 2.2 & 1.9 & 2.4 & 3.7 & 3.7 & 3.7 & 3.1 & 2.6 & 2.2 & 2.4 & 3.6 & 3.8 & 3.2 & 3.3 & 2.8 & 2.3 & 3.1 \\ & & 6 & 4.7 & 4.7 & 5.1 & 4.1 & 3.7 & 4.5 & 4.7 & 4.9 & 5.0 & 5.5 & 4.6 & 4.5 & 3.8 & 5.3 & 3.3 & 3.6 & 3.7 & 3.1 & 2.7 & 1.3 & 2.1 & 2.5 & 2.5 & 2.6 & 2.2 & 2.2 & 1.7 & 2.2 & 4.3 & 3.9 & 3.7 & 2.9 & 2.2 & 1.5 & 1.4 & 4.3 & 4.5 & 3.4 & 4.0 & 3.4 & 2.3 & 2.7 \\ & & 8 & 4.8 & 5.2 & 4.6 & 4.0 & 4.5 & 3.3 & 4.3 & 4.8 & 5.0 & 5.3 & 5.2 & 5.0 & 4.2 & 4.4 & 3.2 & 3.0 & 3.4 & 2.6 & 2.1 & 1.5 & 0.8 & 2.5 & 2.9 & 2.7 & 2.1 & 2.0 & 2.3 & 2.1 & 4.5 & 4.0 & 3.7 & 3.4 & 2.0 & 1.3 & 1.2 & 5.0 & 5.5 & 4.8 & 4.4 & 4.1 & 3.4 & 4.0 \\[0.3em] & 0.15 & 2 & 5.3 & 4.5 & 4.4 & 4.5 & 4.5 & 5.6 & 8.9 & 4.2 & 3.2 & 3.3 & 4.0 & 3.7 & 4.1 & 6.1 & 4.3 & 4.4 & 4.3 & 3.3 & 3.9 & 4.8 & 5.1 & 2.5 & 2.2 & 2.1 & 2.4 & 2.2 & 2.0 & 2.7 & 4.2 & 3.8 & 3.6 & 4.0 & 3.1 & 2.9 & 3.9 & 3.3 & 3.0 & 2.5 & 2.4 & 1.8 & 2.4 & 3.1 \\ & & 4 & 5.0 & 4.7 & 4.3 & 3.9 & 3.7 & 3.8 & 5.6 & 4.3 & 4.0 & 4.0 & 3.8 & 3.2 & 3.5 & 4.6 & 3.8 & 3.9 & 3.6 & 3.0 & 2.4 & 2.6 & 2.7 & 2.4 & 1.9 & 2.2 & 1.9 & 2.0 & 1.9 & 2.2 & 2.9 & 3.3 & 3.5 & 2.8 & 1.9 & 1.6 & 1.8 & 3.4 & 3.6 & 2.9 & 2.5 & 2.3 & 2.5 & 2.6 \\ & & 6 & 4.0 & 4.5 & 4.7 & 3.9 & 4.3 & 3.6 & 3.6 & 3.7 & 4.1 & 3.9 & 4.1 & 3.9 & 3.9 & 3.8 & 3.2 & 3.6 & 3.3 & 2.9 & 2.3 & 1.4 & 1.6 & 2.3 & 2.0 & 2.2 & 1.6 & 2.6 & 1.8 & 2.0 & 3.2 & 3.6 & 3.5 & 3.2 & 1.7 & 1.3 & 1.0 & 3.8 & 3.7 & 2.6 & 3.5 & 2.4 & 3.1 & 3.2 \\ & & 8 & 4.8 & 4.7 & 4.5 & 3.9 & 3.9 & 3.0 & 3.1 & 4.5 & 4.9 & 4.5 & 4.8 & 4.3 & 3.7 & 4.1 & 3.3 & 3.5 & 3.1 & 3.0 & 2.0 & 1.0 & 0.9 & 2.4 & 2.9 & 2.4 & 2.1 & 2.0 & 1.8 & 2.0 & 3.0 & 3.6 & 3.2 & 2.3 & 1.5 & 0.9 & 0.7 & 3.9 & 4.1 & 3.6 & 3.9 & 3.4 & 3.0 & 3.4 \\[0.3em] & 0.40 & 2 & 4.6 & 4.1 & 4.2 & 4.2 & 4.7 & 4.5 & 6.4 & 3.1 & 3.1 & 2.7 & 2.6 & 2.7 & 2.9 & 5.1 & 4.1 & 3.8 & 4.0 & 3.6 & 3.8 & 4.0 & 5.9 & 2.4 & 2.6 & 2.2 & 2.6 & 2.0 & 2.9 & 3.8 & 3.5 & 3.9 & 3.0 & 3.1 & 2.7 & 2.2 & 2.3 & 1.8 & 2.6 & 2.3 & 1.5 & 2.2 & 2.4 & 3.2 \\ & & 4 & 3.5 & 3.6 & 3.8 & 3.2 & 3.4 & 2.6 & 3.5 & 2.7 & 2.9 & 3.1 & 2.9 & 3.0 & 2.7 & 3.1 & 4.0 & 3.5 & 4.1 & 3.6 & 2.9 & 2.5 & 3.3 & 2.5 & 2.7 & 2.5 & 2.3 & 2.2 & 2.1 & 2.5 & 3.0 & 2.8 & 2.6 & 2.2 & 1.4 & 1.0 & 0.9 & 2.3 & 2.2 & 2.4 & 2.0 & 2.1 & 2.4 & 3.1 \\ & & 6 & 4.0 & 4.2 & 4.1 & 3.8 & 2.8 & 2.2 & 2.8 & 3.0 & 3.7 & 2.9 & 3.2 & 2.2 & 2.3 & 3.2 & 3.9 & 4.0 & 3.2 & 3.4 & 2.8 & 2.0 & 2.2 & 2.5 & 2.9 & 2.4 & 2.4 & 2.1 & 2.0 & 2.2 & 3.2 & 3.2 & 2.9 & 2.0 & 1.2 & 0.6 & 0.5 & 2.9 & 2.8 & 3.2 & 2.6 & 2.7 & 2.9 & 4.3 \\ & & 8 & 3.8 & 3.7 & 4.0 & 3.8 & 3.1 & 2.3 & 1.7 & 3.5 & 3.2 & 3.3 & 3.5 & 3.0 & 2.8 & 2.5 & 3.5 & 3.9 & 3.7 & 3.2 & 2.1 & 1.9 & 1.4 & 2.8 & 2.9 & 3.2 & 2.1 & 2.1 & 2.5 & 2.5 & 2.9 & 3.1 & 2.3 & 1.9 & 1.3 & 0.7 & 0.6 & 3.0 & 3.1 & 3.7 & 3.4 & 4.0 & 4.0 & 4.8 \\[0.3em] & 1.20 & 2 & 3.9 & 4.4 & 3.1 & 3.1 & 2.9 & 3.1 & 4.1 & 1.8 & 1.8 & 1.4 & 1.4 & 1.1 & 1.4 & 2.2 & 4.5 & 4.0 & 3.9 & 4.1 & 4.3 & 5.1 & 6.5 & 3.1 & 2.4 & 2.7 & 2.6 & 2.5 & 2.9 & 4.7 & 3.1 & 3.4 & 2.6 & 2.4 & 1.8 & 1.8 & 1.5 & 1.8 & 2.1 & 2.1 & 2.1 & 3.0 & 4.6 & 5.0 \\ & & 4 & 2.9 & 3.0 & 2.5 & 2.4 & 1.4 & 1.3 & 2.1 & 1.8 & 1.7 & 1.6 & 1.3 & 1.2 & 1.2 & 1.6 & 4.5 & 3.7 & 4.2 & 3.6 & 3.8 & 4.0 & 4.1 & 2.7 & 2.8 & 3.0 & 2.8 & 2.4 & 2.4 & 3.3 & 2.3 & 2.1 & 1.7 & 1.4 & 0.8 & 0.6 & 0.4 & 1.4 & 2.0 & 2.4 & 2.9 & 4.2 & 5.4 & 6.3 \\ & & 6 & 2.8 & 3.0 & 2.3 & 2.2 & 1.6 & 0.9 & 0.9 & 1.8 & 1.8 & 1.8 & 1.7 & 1.2 & 1.4 & 1.6 & 3.8 & 4.2 & 3.8 & 3.2 & 3.0 & 3.0 & 3.2 & 2.7 & 3.3 & 2.9 & 3.1 & 2.7 & 2.2 & 2.7 & 2.0 & 1.7 & 1.6 & 1.2 & 0.7 & 0.5 & 0.5 & 1.9 & 2.4 & 2.9 & 3.9 & 4.7 & 6.5 & 8.7 \\ & & 8 & 2.9 & 2.8 & 2.4 & 2.3 & 1.4 & 0.8 & 0.5 & 1.9 & 1.9 & 2.0 & 1.7 & 1.4 & 1.4 & 1.1 & 4.0 & 4.1 & 4.1 & 3.7 & 2.9 & 2.4 & 2.4 & 3.2 & 3.2 & 3.2 & 2.9 & 3.2 & 2.7 & 2.8 & 1.7 & 2.0 & 1.7 & 1.4 & 0.6 & 0.5 & 0.4 & 2.3 & 3.0 & 3.7 & 5.3 & 6.9 & 8.4 & 10.6 \\[0.3em] \hline 300 & 0.04 & 2 & 5.2 & 4.8 & 5.2 & 4.7 & 4.9 & 4.6 & 5.6 & 4.7 & 4.6 & 5.0 & 4.4 & 4.8 & 4.8 & 5.6 & 4.4 & 4.9 & 4.0 & 4.0 & 3.8 & 4.4 & 4.3 & 3.7 & 3.9 & 3.5 & 3.6 & 3.2 & 3.7 & 3.4 & 4.2 & 5.0 & 4.4 & 4.4 & 3.8 & 4.1 & 3.9 & 4.0 & 4.1 & 3.8 & 3.4 & 2.9 & 3.0 & 3.0 \\ & & 4 & 4.5 & 4.6 & 5.2 & 4.3 & 4.8 & 4.6 & 4.6 & 5.6 & 5.4 & 4.3 & 4.9 & 4.8 & 4.6 & 4.1 & 4.2 & 4.0 & 3.8 & 3.7 & 3.3 & 3.2 & 3.0 & 4.2 & 3.6 & 4.6 & 4.0 & 3.4 & 2.9 & 2.9 & 4.5 & 4.1 & 4.6 & 3.5 & 3.3 & 2.7 & 2.4 & 4.5 & 4.1 & 4.6 & 4.0 & 3.1 & 2.3 & 2.4 \\ & & 6 & 4.9 & 4.6 & 4.3 & 4.7 & 4.6 & 3.4 & 3.6 & 5.7 & 5.7 & 5.1 & 5.0 & 4.3 & 4.5 & 3.9 & 4.9 & 3.5 & 4.0 & 4.0 & 3.3 & 2.3 & 1.8 & 3.9 & 4.1 & 4.7 & 4.2 & 3.9 & 3.8 & 2.4 & 4.3 & 5.0 & 3.8 & 3.5 & 2.6 & 2.1 & 1.7 & 5.0 & 5.7 & 4.9 & 4.3 & 3.2 & 2.5 & 2.2 \\ & & 8 & 5.1 & 5.0 & 5.6 & 4.3 & 3.7 & 3.5 & 3.2 & 6.8 & 6.5 & 6.3 & 5.6 & 5.4 & 5.4 & 4.7 & 4.4 & 4.5 & 4.1 & 3.5 & 3.3 & 1.7 & 1.7 & 5.3 & 4.6 & 4.4 & 4.4 & 3.9 & 3.1 & 3.4 & 3.7 & 4.2 & 3.8 & 3.3 & 3.2 & 1.7 & 1.2 & 5.2 & 5.7 & 5.0 & 4.9 & 3.5 & 2.9 & 2.1 \\[0.3em] & 0.08 & 2 & 4.8 & 4.7 & 4.1 & 4.9 & 4.2 & 4.6 & 5.1 & 4.4 & 4.8 & 4.4 & 4.3 & 4.5 & 4.5 & 4.3 & 4.6 & 4.9 & 4.9 & 4.4 & 4.5 & 4.4 & 4.8 & 4.0 & 4.1 & 3.8 & 3.1 & 3.3 & 3.3 & 3.7 & 4.5 & 4.7 & 4.2 & 4.3 & 3.4 & 3.8 & 2.9 & 3.9 & 3.2 & 3.5 & 3.2 & 3.6 & 3.0 & 2.6 \\ & & 4 & 5.2 & 4.6 & 4.6 & 4.1 & 4.3 & 3.7 & 4.3 & 5.1 & 5.1 & 5.0 & 4.8 & 4.4 & 3.8 & 3.7 & 4.2 & 4.1 & 3.9 & 4.4 & 3.7 & 3.1 & 3.1 & 4.8 & 3.9 & 4.4 & 4.2 & 3.6 & 3.1 & 2.9 & 4.7 & 4.5 & 3.9 & 3.9 & 3.3 & 2.4 & 2.3 & 4.2 & 4.6 & 4.2 & 3.2 & 3.3 & 2.3 & 2.0 \\ & & 6 & 4.3 & 4.4 & 5.1 & 4.6 & 3.9 & 3.5 & 2.7 & 6.2 & 5.8 & 5.6 & 5.6 & 5.1 & 4.4 & 3.1 & 4.7 & 4.2 & 4.1 & 4.2 & 3.0 & 2.7 & 2.4 & 4.9 & 4.8 & 4.2 & 4.3 & 3.5 & 3.2 & 2.7 & 4.2 & 3.6 & 4.0 & 4.0 & 3.0 & 1.9 & 1.5 & 4.5 & 4.8 & 4.3 & 4.1 & 3.4 & 2.4 & 2.1 \\ & & 8 & 5.1 & 4.8 & 5.0 & 4.1 & 3.8 & 2.7 & 2.0 & 6.2 & 5.7 & 5.3 & 5.4 & 4.5 & 4.7 & 4.0 & 4.0 & 4.3 & 4.4 & 3.6 & 3.4 & 2.1 & 1.6 & 4.8 & 4.4 & 4.5 & 3.6 & 4.5 & 3.3 & 3.1 & 4.1 & 4.2 & 3.9 & 2.7 & 2.5 & 1.6 & 0.7 & 5.6 & 5.2 & 5.0 & 4.0 & 3.5 & 3.4 & 1.7 \\[0.3em] & 0.15 & 2 & 4.3 & 4.3 & 5.4 & 5.0 & 4.4 & 5.1 & 5.0 & 4.5 & 4.2 & 4.0 & 3.6 & 3.6 & 3.8 & 4.3 & 5.0 & 4.0 & 4.4 & 4.4 & 4.6 & 4.8 & 4.7 & 3.7 & 4.2 & 3.8 & 3.4 & 3.2 & 3.6 & 3.8 & 4.0 & 4.8 & 3.8 & 4.0 & 4.0 & 2.9 & 2.7 & 4.4 & 3.6 & 3.7 & 2.9 & 2.6 & 2.6 & 2.7 \\ & & 4 & 4.6 & 4.3 & 4.4 & 4.5 & 3.3 & 3.1 & 3.4 & 4.4 & 4.4 & 4.9 & 4.2 & 4.3 & 3.5 & 3.5 & 4.8 & 4.8 & 4.7 & 4.2 & 3.6 & 4.3 & 3.6 & 4.4 & 4.1 & 3.7 & 3.9 & 3.1 & 2.8 & 2.9 & 4.1 & 4.0 & 4.0 & 3.5 & 2.9 & 2.4 & 1.7 & 4.0 & 3.5 & 3.2 & 3.3 & 2.1 & 2.2 & 1.6 \\ & & 6 & 5.2 & 4.3 & 4.1 & 4.6 & 3.7 & 2.9 & 2.9 & 5.1 & 5.0 & 5.0 & 5.0 & 4.0 & 3.5 & 2.8 & 4.6 & 4.8 & 4.8 & 3.9 & 3.3 & 2.9 & 2.8 & 5.5 & 4.3 & 4.6 & 3.6 & 3.9 & 3.1 & 3.1 & 4.2 & 3.7 & 3.6 & 3.5 & 2.0 & 1.6 & 1.0 & 4.5 & 4.3 & 4.3 & 4.3 & 2.8 & 2.1 & 1.5 \\ & & 8 & 4.3 & 4.1 & 4.2 & 3.7 & 2.7 & 2.5 & 2.2 & 5.3 & 5.0 & 5.7 & 5.5 & 4.9 & 3.5 & 3.1 & 3.8 & 4.2 & 4.3 & 4.8 & 3.4 & 2.7 & 1.7 & 4.3 & 4.5 & 4.7 & 4.5 & 3.9 & 3.7 & 3.0 & 4.0 & 3.7 & 3.7 & 3.1 & 2.2 & 1.4 & 0.4 & 4.9 & 5.0 & 4.2 & 3.9 & 3.6 & 2.2 & 1.4 \\[0.3em] & 0.40 & 2 & 4.2 & 4.2 & 4.7 & 3.8 & 3.7 & 4.2 & 3.5 & 3.8 & 3.8 & 2.9 & 3.2 & 3.1 & 2.7 & 3.0 & 5.2 & 4.5 & 5.1 & 4.8 & 5.0 & 4.9 & 5.3 & 3.5 & 4.2 & 4.4 & 3.5 & 3.3 & 3.6 & 3.6 & 4.3 & 4.5 & 3.8 & 3.8 & 4.1 & 3.0 & 2.8 & 3.9 & 3.4 & 3.5 & 2.8 & 2.5 & 2.0 & 2.0 \\ & & 4 & 3.8 & 3.8 & 3.9 & 3.9 & 3.4 & 3.0 & 1.7 & 4.2 & 4.6 & 3.2 & 3.2 & 2.7 & 2.4 & 2.1 & 4.9 & 4.4 & 4.3 & 4.5 & 4.6 & 3.7 & 3.7 & 4.9 & 4.6 & 4.6 & 4.2 & 4.1 & 3.4 & 2.6 & 4.0 & 3.9 & 3.6 & 2.6 & 3.3 & 2.1 & 1.3 & 3.8 & 3.4 & 3.5 & 2.9 & 2.4 & 1.7 & 1.6 \\ & & 6 & 3.8 & 4.1 & 3.1 & 3.4 & 2.9 & 2.1 & 1.4 & 3.7 & 4.1 & 4.4 & 3.3 & 2.9 & 2.9 & 1.7 & 4.6 & 4.5 & 4.2 & 4.4 & 4.0 & 3.2 & 3.1 & 4.9 & 4.1 & 4.4 & 4.2 & 4.1 & 3.2 & 3.4 & 3.6 & 3.9 & 3.6 & 3.4 & 2.0 & 1.4 & 0.9 & 3.9 & 3.7 & 3.1 & 3.1 & 2.6 & 1.8 & 1.5 \\ & & 8 & 3.8 & 3.2 & 3.8 & 3.4 & 2.6 & 1.5 & 0.7 & 3.7 & 4.0 & 4.3 & 3.8 & 3.5 & 2.6 & 1.9 & 4.2 & 4.7 & 4.3 & 3.7 & 3.5 & 2.8 & 2.4 & 5.8 & 5.1 & 4.6 & 4.4 & 4.2 & 4.0 & 3.6 & 3.4 & 3.4 & 3.1 & 2.6 & 1.8 & 1.0 & 0.4 & 4.5 & 3.9 & 3.9 & 3.0 & 2.4 & 1.6 & 1.3 \\[0.3em] & 1.20 & 2 & 3.0 & 3.6 & 3.4 & 3.4 & 3.2 & 2.4 & 1.9 & 2.0 & 1.9 & 1.8 & 1.8 & 1.3 & 1.5 & 1.2 & 4.6 & 5.4 & 4.9 & 5.2 & 4.8 & 4.4 & 5.7 & 4.6 & 4.4 & 4.6 & 4.3 & 4.0 & 4.2 & 4.5 & 4.3 & 4.5 & 3.9 & 3.5 & 2.7 & 3.2 & 2.1 & 3.1 & 2.7 & 2.8 & 2.3 & 2.2 & 2.0 & 1.6 \\ & & 4 & 3.6 & 2.9 & 2.9 & 3.5 & 2.2 & 1.2 & 0.8 & 2.0 & 1.9 & 2.2 & 1.6 & 1.4 & 1.0 & 0.6 & 4.6 & 4.8 & 4.5 & 4.1 & 3.8 & 3.7 & 3.7 & 4.1 & 4.3 & 4.2 & 4.1 & 3.4 & 3.8 & 3.8 & 3.2 & 3.6 & 4.0 & 2.8 & 2.2 & 1.3 & 0.9 & 2.8 & 2.8 & 2.5 & 2.3 & 1.7 & 1.3 & 0.9 \\ & & 6 & 2.6 & 2.4 & 2.6 & 2.1 & 1.5 & 0.6 & 0.2 & 2.3 & 2.3 & 1.8 & 1.6 & 1.4 & 0.9 & 0.7 & 4.7 & 4.1 & 4.4 & 4.4 & 3.6 & 3.5 & 3.5 & 5.4 & 4.4 & 4.4 & 4.3 & 4.2 & 3.9 & 3.7 & 3.2 & 2.9 & 2.7 & 2.6 & 1.7 & 1.0 & 0.3 & 3.2 & 3.2 & 2.7 & 2.3 & 2.2 & 1.3 & 1.3 \\ & & 8 & 2.1 & 2.7 & 2.6 & 1.6 & 1.2 & 0.6 & 0.1 & 2.2 & 2.0 & 2.1 & 2.2 & 1.9 & 0.8 & 0.7 & 4.5 & 3.6 & 4.0 & 4.5 & 4.3 & 3.6 & 2.5 & 5.1 & 4.5 & 5.0 & 4.5 & 4.6 & 3.8 & 3.7 & 2.9 & 2.8 & 2.8 & 2.1 & 1.4 & 0.8 & 0.2 & 3.5 & 3.6 & 3.0 & 2.6 & 2.3 & 1.6 & 1.0 \\ \end{tabular} }

Empirical power

To study the empirical power of the proposed method, we consider the following models:

itemize[leftmargin=1.8cm] • First-order exponential autoregressive model: ${\mathbf x}_t = 0.15{\mathbf x}_{t-1} + \exp(-2{\mathbf x}_{t-1}^2) + \boldsymbol{\varepsilon}_t$ with $\boldsymbol{\varepsilon}_t\overset{{\rm i.i.d.}}{\sim}{\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Omega}_\varepsilon)$, where $\boldsymbol{\Omega}_\varepsilon=(\omega_{\varepsilon,kl})_{p\times p}$ with $\omega_{\varepsilon,kl}=I(k=l)+0.25I(k\neq l)$ for any $k,l\in[p]$. • The sum of a white noise and cosine of the first difference of an autoregressive process: ${\mathbf x}_t= \boldsymbol{\varepsilon}_t+0.8\cos(\mathbf z_t-\mathbf z_{t-1})$ with $\mathbf z_t=0.85\mathbf z_{t-1}+{\mathbf u}_t$, $\boldsymbol{\varepsilon}_t\overset{{\rm i.i.d.}}{\sim} {\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Omega}_\varepsilon)$ and ${\mathbf u}_t\overset{{\rm i.i.d.}}{\sim}\mathcal{N}(\boldsymbol{0},\boldsymbol{\Omega}_u)$, where $\boldsymbol{\Omega}_\varepsilon=(\omega_{\varepsilon,kl})_{p\times p}$ and $\boldsymbol{\Omega}_{u}=(\omega_{u,kl})_{p\times p}$ with $\omega_{\varepsilon,kl}=I(k=l)+0.3I(k\neq l)$ and $\omega_{u,kl}=0.7^{|k-l|}$ for any $k,l\in [p]$. • Threshold autoregressive model of order one: ${\mathbf x}_t=(x_{t,1},\ldots,x_{t,p})^{\scriptscriptstyle {\rm \top}}$ with $x_{t,j}=-0.45x_{t-1,j}I(x_{t-1,j}\\ \geqslant 1)+0.6x_{t-1,j}I(x_{t-1,j}< 1)+\varepsilon_{t,j}$ for each $j\in[p]$, where $\boldsymbol{\varepsilon}_t=(\varepsilon_{t,1},\ldots,\varepsilon_{t,p})^{{\scriptscriptstyle {\rm \top}}}\overset{{\rm i.i.d.}}{\sim}{\mathcal{N}}(\boldsymbol{0},{\bf I}_p)$.

Models 4--6 are the multivariate extensions of the univariate models considered in EV2006a (see Models 7--9 there). Table (ref) shows that for Models 4--6, the powers based on three different kernels are similar for the same map with the use of Bartlett kernel exhibiting slightly more power in most cases. When $n=100$ and for Model 4, using the linear and quadratic map leads to more power when $p/n\leqslant0.15$, but less power when $p/n>0.15$. This can be explained by the impact from the high dimension. The additional nonlinear serial dependence captured by the quadratic map is apparent when $p\leqslant 15$, but as the dimension $p$ increases to $120$, the signal related to nonlinear dependence is likely dominated by that related to linear dependence and possibly the noise, so using linear map alone yields more power. Similar phenomena occur for Models 5 and 6. As expected, when we increase the sample size $n$ from $100$ to $300$, we see the appreciation of the power as both linear and nonlinear serial dependence get strengthened at the sample level. Overall, the powers of our tests are quite encouraging for the three models, and all combinations of kernel and map under consideration.

By contrast, the three tests of HLZ2017 mostly fail to reject the martingale difference hypothesis for Models 4 and 5 in all settings. This is presumably due to the inability of their tests to capture nonlinear serial dependence. For Model 6, their tests exhibit great power, which is probably due to the fact that the model implies strong linear serial dependence although it is a nonlinear model per se. Indeed, the sample ACF at lag $1,2,3$ are 0.324, 0.120 and 0.046, respectively, based on our simulation. Again their tests cannot be implemented when $p$ is too large relative to $n$, as their ability of handling the high dimension is quite limited.

sidewaystable[htbp] \scriptsize \caption{Empirical power ($\%$) of the tests $T_{\rm QS}^l$, $T_{\rm PR}^l$, $T_{\rm BT}^l$, $T_{\rm QS}^q$, $T_{\rm PR}^q$, $T_{\rm BT}^q$, $Z_{\rm tr}$, $Z_{\rm det}$ and $Zd_{\rm tr}$ for Models 4--6 at the 5% nominal level. } \resizebox{!}{6.3cm}{ \begin{tabular}{ccc|ccccccccc| ccccccccc| ccccccccc} & & & \multicolumn{9}{c|}{Model 4} & \multicolumn{9}{c|}{Model 5} & \multicolumn{9}{c}{Model 6} \\[0.6em] $n$ & $p/n$ & $K$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ & $T_{\rm QS}^l$ & $T_{\rm PR}^l$ & $T_{\rm BT}^l$ & $T_{\rm QS}^q$ & $T_{\rm PR}^q$ & $T_{\rm BT}^q$ & $Z_{\rm tr}$ & $Z_{\rm det}$ & $Zd_{\rm tr}$ \\[0.5em] \hline 100 & 0.04 & 2 & 79.0 & 78.1 & 81.2 & 93.5 & 93.5 & 94.5 & 6.1 & 5.5 & 6.0 & 65.7 & 65.2 & 69.1 & 94.2 & 94.3 & 95.5 & 4.5 & 4.7 & 5.6 & 77.0 & 77.4 & 84.4 & 80.5 & 81.3 & 85.4 & 100 & 63.9 & 100 \\ & & 4 & 87.8 & 87.5 & 89.2 & 97.8 & 97.8 & 98.3 & 6.7 & 6.0 & 5.9 & 81.5 & 80.3 & 83.4 & 98.3 & 98.4 & 98.6 & 4.5 & 4.6 & 6.6 & 66.7 & 66.5 & 77.7 & 75.7 & 76.2 & 81.8 & 99.7 & 62.4 & 99.4 \\ & & 6 & 91.1 & 90.9 & 92.6 & 98.7 & 98.7 & 99.0 & 6.4 & 5.0 & 6.6 & 87.1 & 86.8 & 89.1 & 99.0 & 99.0 & 99.2 & 5.0 & 5.4 & 7.3 & 64.7 & 63.8 & 76.8 & 75.6 & 76.8 & 82.3 & 97.7 & 59.6 & 97.0 \\ & & 8 & 93.5 & 93.0 & 94.5 & 99.3 & 99.1 & 99.4 & 4.6 & 5.7 & 6.8 & 90.9 & 90.8 & 92.3 & 99.2 & 99.1 & 99.4 & 4.9 & 5.3 & 7.9 & 66.6 & 65.1 & 78.0 & 77.8 & 77.9 & 84.8 & 94.2 & 52.5 & 91.1 \\[0.3em] & 0.08 & 2 & 82.2 & 81.6 & 84.8 & 93.0 & 93.2 & 94.7 & 6.4 & 7.0 & 5.7 & 71.8 & 71.7 & 76.3 & 94.9 & 95.0 & 96.0 & 4.3 & 5.0 & 4.9 & 75.0 & 75.1 & 85.2 & 71.0 & 72.3 & 78.7 & 100 & 51.7 & 100 \\ & & 4 & 91.8 & 91.4 & 93.0 & 97.7 & 97.8 & 98.2 & 6.7 & 6.9 & 6.3 & 86.7 & 86.0 & 89.2 & 97.9 & 97.8 & 98.4 & 5.7 & 5.6 & 6.9 & 65.8 & 65.2 & 79.9 & 65.6 & 67.6 & 75.3 & 100 & 63.3 & 100 \\ & & 6 & 93.9 & 93.6 & 95.3 & 98.6 & 98.7 & 98.9 & 5.4 & 7.4 & 6.7 & 90.7 & 90.4 & 92.7 & 99.0 & 99.1 & 99.3 & 6.5 & 6.3 & 8.2 & 63.6 & 62.7 & 78.6 & 65.0 & 65.7 & 75.0 & 100 & 66.0 & 99.9 \\ & & 8 & 95.3 & 95.0 & 96.3 & 99.1 & 99.2 & 99.3 & 5.8 & 6.7 & 7.0 & 92.6 & 92.6 & 94.9 & 99.0 & 98.9 & 99.2 & 6.6 & 6.1 & 9.1 & 61.6 & 60.4 & 77.8 & 65.8 & 66.5 & 75.9 & 99.7 & 61.4 & 99.7 \\[0.3em] & 0.15 & 2 & 84.2 & 84.1 & 87.2 & 90.8 & 91.0 & 92.7 & NA & NA & 5.9 & 74.7 & 74.6 & 80.1 & 91.7 & 91.9 & 93.7 & NA & NA & 5.3 & 71.3 & 71.6 & 83.3 & 60.3 & 62.4 & 69.9 & NA & NA & 100 \\ & & 4 & 92.2 & 91.6 & 94.4 & 96.8 & 96.7 & 97.3 & NA & NA & 6.3 & 88.1 & 87.5 & 91.3 & 96.8 & 96.8 & 97.6 & NA & NA & 6.6 & 58.4 & 58.5 & 76.7 & 53.1 & 54.8 & 64.9 & NA & NA & 100 \\ & & 6 & 95.2 & 95.0 & 96.6 & 97.7 & 97.7 & 97.9 & NA & NA & 7.9 & 91.5 & 91.4 & 94.2 & 97.7 & 97.9 & 98.1 & NA & NA & 9.4 & 55.5 & 54.6 & 75.4 & 50.1 & 51.5 & 62.9 & NA & NA & 100 \\ & & 8 & 96.0 & 95.8 & 97.1 & 98.3 & 98.2 & 98.6 & NA & NA & 8.2 & 94.4 & 94.4 & 96.2 & 98.4 & 98.5 & 98.8 & NA & NA & 11.2 & 54.3 & 53.2 & 75.3 & 48.9 & 50.8 & 61.7 & NA & NA & 100 \\[0.3em] & 0.40 & 2 & 85.0 & 84.5 & 88.7 & 78.8 & 79.7 & 82.4 & NA & NA & 6.3 & 73.7 & 73.4 & 79.6 & 80.5 & 81.4 & 84.7 & NA & NA & 5.9 & 62.0 & 62.8 & 81.7 & 44.5 & 47.7 & 55.3 & NA & NA & 100 \\ & & 4 & 91.5 & 91.1 & 94.3 & 88.4 & 89.0 & 90.4 & NA & NA & 7.0 & 87.6 & 87.0 & 91.4 & 90.0 & 90.7 & 92.6 & NA & NA & 9.2 & 47.2 & 47.0 & 72.9 & 34.6 & 37.7 & 44.9 & NA & NA & 100 \\ & & 6 & 94.6 & 94.2 & 96.9 & 91.7 & 92.2 & 93.4 & NA & NA & 6.7 & 91.7 & 91.2 & 94.5 & 91.8 & 92.2 & 93.7 & NA & NA & 12.8 & 40.4 & 38.8 & 69.5 & 32.3 & 34.1 & 41.1 & NA & NA & 100 \\ & & 8 & 95.2 & 94.9 & 97.2 & 92.2 & 92.7 & 93.3 & NA & NA & 9.4 & 92.9 & 92.8 & 95.8 & 93.6 & 93.9 & 95.2 & NA & NA & 14.5 & 36.3 & 35.1 & 65.4 & 32.0 & 34.1 & 39.8 & NA & NA & 100 \\[0.3em] & 1.20 & 2 & 78.5 & 78.0 & 83.9 & 44.3 & 45.6 & 48.8 & NA & NA & NA & 67.6 & 68.3 & 77.9 & 52.5 & 55.0 & 58.8 & NA & NA & NA & 49.3 & 49.1 & 77.3 & 43.7 & 46.8 & 45.8 & NA & NA & NA \\ & & 4 & 87.6 & 86.9 & 91.8 & 57.0 & 58.6 & 61.8 & NA & NA & NA & 82.8 & 82.1 & 88.8 & 64.3 & 66.2 & 70.1 & NA & NA & NA & 30.0 & 29.4 & 60.8 & 46.9 & 49.4 & 43.5 & NA & NA & NA \\ & & 6 & 89.2 & 88.8 & 93.2 & 58.3 & 60.0 & 62.7 & NA & NA & NA & 87.1 & 86.4 & 92.2 & 68.6 & 71.0 & 73.6 & NA & NA & NA & 21.7 & 20.9 & 52.3 & 49.7 & 53.1 & 44.8 & NA & NA & NA \\ & & 8 & 90.8 & 90.3 & 94.5 & 62.6 & 64.2 & 66.9 & NA & NA & NA & 88.9 & 88.6 & 92.9 & 69.4 & 71.4 & 74.8 & NA & NA & NA & 17.5 & 16.7 & 44.4 & 53.9 & 56.2 & 47.9 & NA & NA & NA \\[0.3em] \hline 300 & 0.04 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 13.6 & 8.3 & 8.6 & 100 & 100 & 100 & 100 & 100 & 100 & 5.0 & 5.3 & 5.2 & 100 & 100 & 100 & 99.4 & 99.5 & 99.7 & 100 & 99.0 & 100 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 15.1 & 8.9 & 10.4 & 100 & 100 & 100 & 100 & 100 & 100 & 5.1 & 5.9 & 5.7 & 99.8 & 99.8 & 99.9 & 98.9 & 99.0 & 99.3 & 100 & 98.6 & 100 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 10.4 & 8.9 & 7.9 & 100 & 100 & 100 & 100 & 100 & 100 & 6.1 & 6.5 & 6.1 & 99.8 & 99.7 & 99.9 & 98.6 & 98.5 & 99.1 & 100 & 97.0 & 100 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 9.7 & 9.2 & 6.2 & 100 & 100 & 100 & 100 & 100 & 100 & 6.5 & 6.9 & 6.2 & 100 & 100 & 100 & 99.2 & 99.2 & 99.6 & 100 & 95.7 & 100 \\[0.3em] & 0.08 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 9.3 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 4.8 & 100 & 100 & 100 & 98.6 & 98.6 & 99.3 & NA & NA & 100 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 11.7 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 5.9 & 99.9 & 99.9 & 100 & 98.2 & 98.2 & 99.1 & NA & NA & 100 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 8.3 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 6.0 & 99.9 & 99.9 & 100 & 98.3 & 98.4 & 99.1 & NA & NA & 100 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 7.3 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 7.0 & 100 & 100 & 100 & 98.3 & 98.3 & 99.0 & NA & NA & 100 \\[0.3em] & 0.15 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 10.2 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 5.6 & 100 & 100 & 100 & 98.2 & 98.4 & 98.8 & NA & NA & 100 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 11.1 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 6.1 & 99.9 & 99.9 & 100 & 97.0 & 97.0 & 98.3 & NA & NA & 100 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 9.9 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 6.8 & 99.9 & 99.9 & 100 & 97.2 & 97.3 & 98.6 & NA & NA & 100 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 7.0 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & 7.3 & 99.9 & 100 & 100 & 96.8 & 97.0 & 98.6 & NA & NA & 100 \\[0.3em] & 0.40 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 96.4 & 96.4 & 98.2 & NA & NA & NA \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 99.9 & 100 & 94.3 & 94.4 & 97.4 & NA & NA & NA \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 99.9 & 99.9 & 100 & 92.0 & 92.1 & 96.9 & NA & NA & NA \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 92.5 & 92.3 & 96.9 & NA & NA & NA \\[0.3em] & 1.20 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 92.9 & 92.7 & 97.0 & NA & NA & NA \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 82.9 & 83.2 & 93.0 & NA & NA & NA \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 76.0 & 76.1 & 89.6 & NA & NA & NA \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 100 & 100 & 100 & NA & NA & NA & 100 & 100 & 100 & 69.3 & 69.3 & 87.1 & NA & NA & NA \\ \end{tabular} }
sidewaystable[htbp] \scriptsize \caption{Empirical powers ($\%$) of the tests $T_{\rm BT}^l$ and $T_{\rm BT}^q$ for Models 4--6 at the 5% nominal level, where $c$ represents the constant which is multiplied by Andrews' bandwidth. } \resizebox{!}{5.2cm}{ \begin{tabular}{ccc|ccccccc|ccccccc|ccccccc|ccccccc|ccccccc|ccccccc} & & & \multicolumn{7}{c|}{Model 4 with $T_{\rm BT}^l$} & \multicolumn{7}{c|}{Model 4 with $T_{\rm BT}^q$} & \multicolumn{7}{c|}{Model 5 with $T_{\rm BT}^l$} & \multicolumn{7}{c|}{Model 5 with $T_{\rm BT}^q$} & \multicolumn{7}{c|}{Model 6 with $T_{\rm BT}^l$} & \multicolumn{7}{c}{Model 6 with $T_{\rm BT}^q$} \\[0.4em] \hline & & & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c|}{$c$} & \multicolumn{7}{c}{$c$} \\[0.3em] $n$ & $p/n$ & $K$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ & $2^{-3}$ & $2^{-2}$ & $2^{-1}$ & $2^0$ & $2^1$ & $2^2$ & $2^3$ \\[0.4em] \hline 100 & 0.04 & 2 & 81.2 & 81.8 & 80.5 & 78.8 & 76.6 & 73.2 & 72.1 & 95.7 & 95.9 & 95.6 & 93.5 & 93.8 & 92.8 & 92.4 & 73.2 & 73.7 & 72.0 & 69.0 & 64.3 & 61.5 & 61.2 & 96.0 & 96.3 & 96.0 & 95.8 & 95.0 & 94.0 & 93.9 & 99.5 & 98.2 & 93.2 & 84.3 & 79.1 & 79.9 & 88.1 & 99.6 & 96.7 & 90.6 & 84.4 & 84.7 & 87.3 & 95.6 \\ & & 4 & 89.9 & 89.8 & 89.3 & 88.6 & 86.7 & 80.6 & 78.3 & 99.1 & 98.8 & 98.9 & 98.0 & 97.7 & 96.6 & 97.0 & 85.9 & 85.6 & 85.9 & 83.3 & 79.6 & 72.9 & 69.7 & 99.0 & 98.9 & 98.8 & 98.8 & 97.9 & 97.5 & 96.7 & 98.5 & 97.7 & 91.6 & 78.6 & 68.2 & 68.0 & 78.2 & 99.4 & 96.9 & 91.2 & 82.8 & 78.8 & 84.0 & 93.1 \\ & & 6 & 93.2 & 93.2 & 94.0 & 91.3 & 89.8 & 85.1 & 80.6 & 99.3 & 99.2 & 99.4 & 99.0 & 99.0 & 98.2 & 98.2 & 90.2 & 90.4 & 90.5 & 89.3 & 85.9 & 79.9 & 75.9 & 99.0 & 99.5 & 99.1 & 99.0 & 98.8 & 98.6 & 98.2 & 98.7 & 97.6 & 91.8 & 78.5 & 65.0 & 64.2 & 72.5 & 99.7 & 97.7 & 91.0 & 84.3 & 79.7 & 82.6 & 92.0 \\ & & 8 & 94.4 & 94.7 & 94.4 & 93.6 & 91.2 & 88.3 & 82.7 & 99.6 & 99.6 & 99.2 & 99.4 & 99.1 & 99.1 & 98.8 & 93.1 & 93.4 & 92.5 & 91.8 & 89.1 & 84.7 & 78.3 & 99.5 & 99.2 & 99.4 & 99.5 & 99.4 & 98.9 & 98.9 & 98.6 & 97.7 & 91.2 & 78.9 & 65.4 & 61.6 & 70.3 & 99.5 & 98.0 & 91.8 & 85.1 & 80.3 & 83.2 & 92.4 \\[0.3em] & 0.08 & 2 & 87.0 & 86.7 & 86.6 & 85.0 & 80.9 & 76.6 & 72.7 & 95.7 & 95.7 & 95.2 & 94.9 & 94.2 & 92.6 & 92.4 & 80.6 & 80.8 & 78.9 & 76.6 & 69.9 & 65.3 & 63.2 & 96.0 & 96.4 & 95.7 & 95.3 & 93.9 & 93.5 & 93.3 & 99.9 & 99.3 & 94.6 & 85.1 & 77.4 & 77.6 & 88.7 & 99.7 & 95.3 & 86.8 & 78.4 & 76.5 & 84.0 & 94.3 \\ & & 4 & 93.7 & 93.5 & 93.8 & 92.6 & 89.8 & 83.9 & 77.3 & 98.4 & 98.9 & 98.7 & 98.2 & 97.5 & 97.1 & 96.4 & 90.4 & 90.7 & 91.1 & 88.9 & 84.1 & 76.9 & 70.5 & 98.9 & 98.4 & 98.6 & 98.7 & 97.6 & 96.8 & 96.3 & 99.9 & 99.2 & 94.0 & 78.3 & 66.9 & 66.9 & 79.4 & 99.3 & 96.1 & 85.7 & 74.2 & 71.9 & 79.8 & 92.6 \\ & & 6 & 96.7 & 96.2 & 96.1 & 94.6 & 93.5 & 87.2 & 80.4 & 99.2 & 99.2 & 99.3 & 99.0 & 98.5 & 98.2 & 97.9 & 93.9 & 94.3 & 93.8 & 92.6 & 89.9 & 82.0 & 74.6 & 99.1 & 99.2 & 99.3 & 99.0 & 98.6 & 98.1 & 98.1 & 99.7 & 99.3 & 94.0 & 78.5 & 62.2 & 59.3 & 73.7 & 99.3 & 95.8 & 85.9 & 76.4 & 70.0 & 77.7 & 91.8 \\ & & 8 & 97.1 & 96.9 & 97.2 & 96.2 & 93.9 & 90.2 & 83.5 & 99.4 & 99.4 & 99.1 & 99.2 & 98.6 & 98.9 & 98.4 & 95.9 & 95.8 & 95.6 & 95.1 & 92.1 & 86.6 & 79.1 & 99.5 & 99.5 & 99.4 & 99.3 & 99.0 & 99.0 & 98.4 & 99.8 & 99.5 & 94.8 & 78.5 & 61.2 & 56.5 & 70.6 & 99.1 & 96.3 & 87.9 & 76.0 & 69.7 & 76.2 & 90.5 \\[0.3em] & 0.15 & 2 & 88.4 & 88.5 & 89.5 & 87.0 & 82.2 & 75.4 & 69.6 & 94.4 & 93.6 & 93.7 & 92.8 & 90.5 & 90.3 & 90.3 & 84.7 & 83.8 & 83.9 & 79.0 & 72.5 & 65.9 & 59.6 & 95.1 & 95.4 & 94.4 & 93.2 & 92.1 & 91.1 & 90.5 & 100 & 99.8 & 95.4 & 83.1 & 76.1 & 76.3 & 87.1 & 99.0 & 93.9 & 81.1 & 71.2 & 70.1 & 79.7 & 93.2 \\ & & 4 & 95.6 & 95.4 & 94.7 & 94.6 & 91.4 & 83.6 & 74.9 & 97.9 & 97.7 & 97.4 & 97.3 & 97.1 & 95.8 & 94.7 & 93.0 & 92.6 & 92.5 & 91.2 & 86.1 & 75.5 & 69.4 & 97.9 & 97.9 & 98.1 & 97.3 & 96.9 & 95.3 & 94.9 & 99.9 & 99.5 & 94.8 & 76.5 & 61.1 & 60.5 & 77.3 & 98.3 & 93.0 & 78.6 & 63.0 & 59.7 & 71.7 & 90.7 \\ & & 6 & 97.2 & 96.8 & 97.2 & 96.9 & 94.2 & 87.8 & 76.9 & 98.3 & 98.4 & 98.3 & 98.0 & 98.2 & 97.0 & 95.7 & 95.6 & 95.6 & 95.5 & 93.8 & 90.7 & 81.4 & 74.8 & 98.9 & 98.6 & 98.5 & 98.2 & 97.8 & 96.6 & 96.8 & 99.9 & 99.4 & 94.6 & 75.3 & 58.6 & 53.0 & 69.7 & 98.1 & 93.3 & 78.7 & 63.0 & 58.3 & 68.4 & 89.5 \\ & & 8 & 97.8 & 98.0 & 98.3 & 97.0 & 94.9 & 89.5 & 80.8 & 98.6 & 98.6 & 99.0 & 98.7 & 98.3 & 97.1 & 96.4 & 96.9 & 96.9 & 96.6 & 95.7 & 93.1 & 85.4 & 77.3 & 98.8 & 98.8 & 98.8 & 98.6 & 98.5 & 97.7 & 97.2 & 99.9 & 99.5 & 94.9 & 78.2 & 55.2 & 48.3 & 66.7 & 97.9 & 93.6 & 79.2 & 65.0 & 56.9 & 65.7 & 87.5 \\[0.3em] & 0.40 & 2 & 89.7 & 90.9 & 89.5 & 88.2 & 81.8 & 71.2 & 63.0 & 82.9 & 84.6 & 82.3 & 81.4 & 79.2 & 76.2 & 78.4 & 85.7 & 85.3 & 85.3 & 81.4 & 72.6 & 61.7 & 53.4 & 87.1 & 87.3 & 86.9 & 85.2 & 81.8 & 79.7 & 81.6 & 99.9 & 99.5 & 96.2 & 82.6 & 67.3 & 67.3 & 83.4 & 95.6 & 87.2 & 69.5 & 55.1 & 55.0 & 68.6 & 90.4 \\ & & 4 & 95.6 & 96.3 & 95.7 & 93.9 & 88.7 & 79.3 & 66.2 & 90.4 & 90.2 & 90.7 & 89.6 & 88.0 & 84.7 & 84.0 & 93.2 & 94.2 & 93.1 & 91.2 & 84.8 & 73.1 & 59.2 & 93.6 & 93.4 & 93.2 & 91.4 & 90.4 & 87.7 & 86.9 & 99.9 & 99.3 & 94.8 & 73.1 & 49.7 & 47.0 & 68.7 & 92.6 & 84.6 & 63.1 & 44.1 & 43.0 & 56.7 & 86.5 \\ & & 6 & 97.2 & 97.8 & 97.0 & 96.1 & 92.3 & 82.3 & 68.2 & 92.4 & 92.6 & 92.0 & 92.1 & 91.7 & 88.6 & 87.6 & 95.7 & 96.5 & 96.3 & 93.5 & 88.9 & 76.9 & 64.9 & 94.8 & 94.6 & 94.6 & 93.5 & 93.0 & 89.7 & 89.1 & 99.8 & 99.4 & 95.3 & 69.5 & 41.9 & 36.9 & 60.5 & 89.4 & 80.8 & 58.7 & 40.1 & 41.6 & 52.0 & 83.8 \\ & & 8 & 97.8 & 97.8 & 97.9 & 96.7 & 93.0 & 85.7 & 71.0 & 93.3 & 93.1 & 93.7 & 93.1 & 92.4 & 90.7 & 88.7 & 97.5 & 97.4 & 97.0 & 95.6 & 92.3 & 81.1 & 67.4 & 94.9 & 95.1 & 95.2 & 94.1 & 93.6 & 91.3 & 90.2 & 99.8 & 99.4 & 95.0 & 66.2 & 38.8 & 33.1 & 54.8 & 87.6 & 79.8 & 56.9 & 39.6 & 38.1 & 51.9 & 83.1 \\[0.3em] & 1.20 & 2 & 86.7 & 86.2 & 88.1 & 83.9 & 76.3 & 62.3 & 49.4 & 51.2 & 49.4 & 49.8 & 48.7 & 45.3 & 43.6 & 46.9 & 84.2 & 84.1 & 84.0 & 78.5 & 65.1 & 52.3 & 41.8 & 64.0 & 63.2 & 63.0 & 58.6 & 55.2 & 51.8 & 58.8 & 99.7 & 99.3 & 97.0 & 77.6 & 54.1 & 54.0 & 75.4 & 78.9 & 67.6 & 51.2 & 46.8 & 50.3 & 64.6 & 89.0 \\ & & 4 & 93.6 & 93.6 & 93.0 & 90.4 & 84.4 & 69.1 & 51.9 & 60.3 & 60.9 & 60.7 & 58.4 & 55.2 & 51.9 & 52.5 & 91.4 & 91.5 & 91.2 & 88.3 & 79.5 & 59.5 & 45.3 & 72.4 & 71.0 & 71.7 & 69.7 & 65.4 & 60.1 & 61.8 & 99.1 & 99.0 & 95.2 & 63.0 & 30.7 & 28.2 & 53.1 & 64.1 & 56.3 & 41.5 & 42.9 & 49.9 & 59.9 & 85.0 \\ & & 6 & 95.3 & 96.1 & 96.0 & 93.9 & 87.1 & 72.2 & 53.4 & 63.6 & 63.9 & 63.1 & 63.7 & 60.1 & 56.4 & 54.5 & 94.5 & 94.6 & 94.5 & 91.8 & 84.5 & 67.7 & 50.3 & 75.2 & 74.7 & 74.1 & 73.7 & 69.4 & 64.4 & 66.4 & 99.1 & 98.0 & 93.0 & 53.2 & 23.1 & 20.5 & 43.1 & 55.7 & 49.7 & 40.3 & 45.1 & 54.0 & 64.0 & 84.4 \\ & & 8 & 96.4 & 96.0 & 96.3 & 94.7 & 89.3 & 75.2 & 55.3 & 63.8 & 65.7 & 65.5 & 65.8 & 61.0 & 58.6 & 57.2 & 96.1 & 95.5 & 95.4 & 93.3 & 86.7 & 73.2 & 54.2 & 75.6 & 74.6 & 74.8 & 74.8 & 70.8 & 68.1 & 67.5 & 98.3 & 96.8 & 92.8 & 48.7 & 20.8 & 18.3 & 40.4 & 50.8 & 46.6 & 41.2 & 48.1 & 57.5 & 68.9 & 85.1 \\[0.4em] \hline 300 & 0.04 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.9 & 99.9 & 100 & 99.9 & 99.8 & 99.6 & 99.5 & 99.4 & 99.7 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.7 & 99.5 & 100 & 100 & 99.8 & 99.5 & 99.1 & 99.2 & 99.4 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.8 & 99.6 & 99.2 & 100 & 99.9 & 99.8 & 99.4 & 99.0 & 99.1 & 99.5 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.6 & 99.3 & 100 & 100 & 99.8 & 99.5 & 99.4 & 99.2 & 99.0 \\[0.3em] & 0.08 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.9 & 99.9 & 100 & 99.9 & 99.7 & 99.3 & 99.0 & 99.0 & 99.5 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.6 & 99.3 & 100 & 99.9 & 99.5 & 98.9 & 98.9 & 98.5 & 99.1 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.7 & 99.3 & 100 & 100 & 99.5 & 98.9 & 98.5 & 98.4 & 99.1 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.6 & 99.2 & 100 & 99.8 & 99.7 & 99.3 & 98.5 & 98.5 & 99.1 \\[0.3em] & 0.15 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 100 & 100 & 100 & 99.3 & 99.1 & 98.8 & 99.0 & 99.5 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.8 & 99.8 & 99.5 & 100 & 99.8 & 99.5 & 98.7 & 97.8 & 97.7 & 98.8 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.7 & 99.1 & 100 & 100 & 99.4 & 98.6 & 97.8 & 98.1 & 98.5 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.7 & 99.2 & 100 & 100 & 99.4 & 98.7 & 98.1 & 97.3 & 98.5 \\[0.3em] & 0.40 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 100 & 99.8 & 99.3 & 97.9 & 97.3 & 97.6 & 98.7 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.7 & 99.4 & 99.8 & 99.8 & 98.9 & 97.5 & 95.3 & 94.9 & 96.6 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.7 & 98.8 & 99.8 & 99.8 & 98.5 & 97.3 & 94.2 & 93.3 & 95.5 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.7 & 98.4 & 99.8 & 99.7 & 98.9 & 96.6 & 92.7 & 91.5 & 94.5 \\[0.3em] & 1.20 & 2 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.9 & 99.7 & 99.5 & 98.9 & 96.9 & 95.2 & 94.3 & 95.8 \\ & & 4 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.8 & 98.7 & 98.5 & 98.3 & 97.4 & 92.9 & 87.0 & 84.6 & 88.8 \\ & & 6 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 99.5 & 97.4 & 98.0 & 97.5 & 96.1 & 89.5 & 79.0 & 76.0 & 82.5 \\ & & 8 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 100 & 98.8 & 95.5 & 96.6 & 96.0 & 94.5 & 86.5 & 72.4 & 68.5 & 77.6 \\ \end{tabular} }

Power curve

In this subsection, we perturb Models 1--3 so that the new sequence is not a MDS and present power curves. For given constant $a\in\{0,0.5,1,1.5,2,2.5\}$, the model settings are as follows:

itemize[leftmargin=1.8cm] • Let ${\mathbf x}_t$ follow Model 1 and $\mathbf y_t={\mathbf x}_t+a\exp(-2{\mathbf x}_{t-1}^2)$. • Let ${\mathbf x}_t$ follow Model 2 and $\mathbf y_t={\mathbf x}_t+a\cos(\boldsymbol\varepsilon_{t-1}\circ\boldsymbol{\sigma}_{t-1})$, where $\boldsymbol \varepsilon_{t-1}$ and $\boldsymbol\sigma_{t-1}$ are specified in Model 2. • Let ${\mathbf x}_t$ follow Model 3 and $\mathbf y_t={\mathbf x}_t+a\log({\mathbf x}_{t-2}^2)$.

We aim to test whether $\{\mathbf y_t\}_{t\in\mathbb{Z}}$ defined in Models 1'--3' is a MDS. When $a=0$, $\mathbf y_t={\mathbf x}_t$ and Models 1'--3' become Models 1--3, respectively, which follow the null hypothesis. Figures (ref)--(ref) display the empirical sizes and powers of our proposed tests ($T_{\rm BT}^l, T_{\rm BT}^q$) and HLZ2017's test ($Zd_{\rm tr}$) when the sample size $n=100$. Notice that $Zd_{\rm tr}$ is feasible when $p<n$. Thus when $p/n=1.2$, there is no power curve for $Zd_{\rm tr}$. As seen from Figure (ref), our tests and HLZ2017's test control the empirical sizes well under the null hypothesis with $a=0$ and the empirical powers increase for larger values of the distance parameter $a$. But our tests outperform HLZ2017's test especially for large $K$. In Figure (ref), HLZ2017's test almost cannot detect the alternative hypotheses, but our tests still work well. This is presumably due to the inability of their test to capture nonlinear serial dependence. Based on Figure (ref), similar phenomenon is observed that the empirical powers increase as the distance $a$ grows. Somewhat counter-intuitively, the empirical powers of $Zd_{\rm tr}$ decrease when $a$ increases from $2$ to $2.5$, which means the power is non-monotonic. In addition, comparing the results of our tests for two maps, we find that the test based on linear and quadratic map is more powerful than the test only based on linear map for the three models. This should not be surprising. Since the alternatives in the three models are nonlinear transformations, the linear and quadratic map can capture both linear and nonlinear dependence. Generally speaking, both of the two maps perform well in the three models.

figure[figure omitted — 265 chars of source]
figure[figure omitted — 263 chars of source]
figure[figure omitted — 266 chars of source]

Real data analysis

In this section, we apply our proposed tests to a real dataset, which collects weekly closing prices from 17 September 2004 to 26 December 2008 for 394 stocks. The returns of the stocks are obtained by the log difference of the data. And the sample size $n$ for the returns is 223. These stocks can be classified into 9 major sectors, which consist of materials (22 stocks), real estate (25 stocks), utilities (26 stocks), consumer staples (30 stocks), healthcare (55 stocks), industrials (56 stocks), financials (58 stocks), IT (60 stocks), and consumer discretionary (62 stocks). Here we examine the validity of the martingale difference hypothesis within each sector and for all stocks using our tests and the ones proposed in HLZ2017. Note that neither $Z_{{\rm tr}}$ nor $Z_{{\rm det}}$ is applicable here, since $p<\sqrt{n}$ is violated for each sector. Hence we only present the results of $Zd_{\rm tr}$ for each sector, as it is not usable when we apply to all stock returns. Denote by ${\mathbf x}_t$ the returns of these stocks at time $t$. Financial theory usually assumes the stock prices follow geometric Brownian Motion which implies $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ under the efficient markets hypothesis. We can propose the test statistic $T_{{\rm mean}}=|n^{-1/2}\sum_{t=1}^n{\mathbf x}_t|_\infty$ for the null hypothesis $H_0:\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$. Using the method given in Section 4.1 of CCW2020 with three kernels (QS, PR, BT) to estimate the associated long-run covariance matrix, the associated p-values for such null hypothesis are 0.759, 0.749 and 0.753, respectively, which means there is no strong evidence against the zero-mean assumption of ${\mathbf x}_t$ in our real data.

Table (ref) reports the p-values of $Zd_{\rm tr}$ and our tests with assuming $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ and without assuming $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$. It appears that there is no strong evidence against the martingale difference hypothesis based on all tests, except for a marginally significant p-value of $Zd_{\rm tr}$ when $K=2$ for the sector of consumer staples. Generally speaking, the martingale difference hypothesis is expected to hold for the weekly returns data, so in a sense both our tests and $Zd_{\rm tr}$ help confirming this property. For the same map, the use of different kernels do not seem to affect the p-values much, indicating the insensitivity of our results with respect to the kernel. For this particular dataset, the use of linear and quadratic maps also produces p-values that are not far away from the use of linear maps alone, for most sectors. The p-values corresponding to $Zd_{\rm tr}$ seem to monotonically decrease as $K$ goes down from $8$ to $2$ for all sectors, an interesting phenomenon worthy of some theoretical investigation. In addition, the results of our tests with assuming $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ and without assuming $\mathbb{E}({\mathbf x}_t)=\boldsymbol{0}$ are quite similar, which is consistent with the aforementioned conclusion that $\mathbb{E}({\mathbf x}_t)$ is not significantly different from zero. Overall, our tests are preferred to the ones proposed in HLZ2017 due to the fact that they can be used regardless of whether the dimension $p$ exceeds the sample size $n$.

table[table omitted — 6,940 chars of source]

Discussion

In this paper, we propose a new martingale difference test that captures nonlinear serial dependence and works in the high-dimensional environment, as motivated by the increasing availability of high-dimensional nonlinear time series from economics and finance. Under mild moment and weak temporal dependence assumptions, we establish the validity of Gaussian approximation and provide a simulation-based approach for critical values. In addition to its built-in capability of accommodating both low and high dimensions, our test also has a number of appealing features such as being robust to conditional moments of unknown forms and strong/weak cross-series dependence. From our numerical simulations and a real data analysis, we observe quite encouraging finite sample performance. Therefore we feel confident to recommend its use by the practitioners when there is a need to assess the martingale difference hypothesis for econometric/financial time series of moderate or high dimension.

In the literature, testing quantile/directional predictability has been studied for low-dimensional time series; see HLOW2016. It would be also interesting to extend their test to the high-dimensional setting. A sound data-driven bandwidth choice in our simulation-based approach for generating the critical values merits additional research, especially from a testing-optimal viewpoint. We leave these topics for future investigation.

Technical proofs

In this section, we provide the detailed proofs for all theoretical results stated in the paper, and also introduce necessary lemmas and propositions with proofs. Throughout this section, we use $C$ to denote a generic positive finite constant that does not depend on $(p,d,n,K)$ and may be different in different uses. For two sequences of positive numbers $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ or $b_n\gtrsim a_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n\leqslant c_0$ for some positive constant $c_0$. We write $a_n\asymp b_n$ if $a_n\lesssim b_n$ and $b_n\lesssim a_n$ hold simultaneously. We write $a_n\ll b_n$ or $b_n\gg a_n $ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$. For a countable set $\mathcal{F}$, we use $|\mathcal{F}|$ to denote the cardinality of $\mathcal{F}$.

Write $\mathbf u := ( u_1,\ldots, u_{Kpd} )^{{\scriptscriptstyle {\rm \top}}} =(\hat\boldsymbol{\gamma}_1^{{\scriptscriptstyle {\rm \top}}},\ldots,\hat\boldsymbol{\gamma}_{K}^{{\scriptscriptstyle {\rm \top}}})^{{\scriptscriptstyle {\rm \top}}}$ with $\hat{\boldsymbol{\gamma}}_j=(n-j)^{-1}\sum_{t=1}^{n-j} {\rm vec}\{\boldsymbol{\phi}({\mathbf x}_{t}){\mathbf x}_{t+j}^{{\scriptscriptstyle {\rm \top}}}\}$ for any $j\in[K]$. Let $\tilde n=n-K$. Recall $\boldsymbol \eta_t=([{\rm vec} \{\boldsymbol{\phi}({\mathbf x}_t){\mathbf x}_{t+1}^{{\scriptscriptstyle {\rm \top}}}\}]^{\scriptscriptstyle {\rm \top}},\ldots, [{\rm vec} \{ \boldsymbol{\phi}({\mathbf x}_t){\mathbf x}_{t+K}^{{\scriptscriptstyle {\rm \top}}}\}]^{\scriptscriptstyle {\rm \top}} )^{{\scriptscriptstyle {\rm \top}}}$. Since $\{{\mathbf x}_t\}$ is an $\alpha$-mixing process satisfying Condition (ref), we know the newly defined process $\{\boldsymbol \eta_t\}$ is also $\alpha$-mixing with the $\alpha$-mixing coefficients $\{\tilde{\alpha}_K(k)\}_{k\geqslant1}$ satisfying

align[align omitted — 102 chars of source]

where the positive constants $\tau_2$, $C_3$ and $C_4$ are specified in Condition 2. Write $\bar \boldsymbol \eta :=(\bar{\eta}_1,\ldots,\bar{\eta}_{Kpd})^{\scriptscriptstyle {\rm \top}}= \tilde n^{-1}\sum_{t=1}^{\tilde n}\boldsymbol \eta_t$. For each $j\in[K]$, define $Z_j =n\max_{\ell\in\mathcal{L}_j} u_{\ell}^2 $ and $\tilde Z_j = \tilde n \max_{\ell\in\mathcal{L}_j} \bar\eta_\ell^2$ with $\mathcal{L}_j:=\{(j-1)pd+1,\ldots,jpd\}$. Then the test statistic can be written as $ T_n=n\sum_{j=1}^{K}|\hat\boldsymbol{\gamma}_j|_{\infty}^2= \sum_{j=1}^K Z_j$. Furthermore, we let $ \tilde T_n := \sum_{j=1}^K \tilde Z_j$.

A key proposition

Let $\{\mathbf z_t\}_{t=1}^{n}$ be a $d_{z}$-dimensional dependent sequence with $\mathbb{E}(\mathbf z_t)=\boldsymbol 0$ for any $t\in[n]$. Define $\mathbf s_{n,z}=n^{-1/2}\sum_{t=1}^n \mathbf z_t$ and $\boldsymbol{\Xi}=\mbox{Var}(n^{-1/2}\sum_{t=1}^n \mathbf z_t)$. Write $\mathbf z_t=(z_{t,1},\ldots,z_{t,d_z})^{\scriptscriptstyle {\rm \top}}$. We assume $\{\mathbf z_t\}_{t=1}^n$ satisfy the following three assumptions:

itemize• There exist universal constants $b_1>1$, $b_2>0$ and $r_1 \in (0,1]$ such that $ \sup_{t\in[n]}\sup_{j\in[d_{z}]}\mathbb{P}(|z_{t,j}|>u) \leqslant b_1\exp(-b_2 u^{r_1} )$ for any $u>0$. • There exist universal constants $a_1>1$, $a_2>0$ and $r_2\in(0,1]$ such that the $\alpha$-mixing coefficients of the sequence $\{\mathbf z_t\}_{t=1}^n$, denoted by $\{\alpha_z(k)\}_{k\geqslant1}$, satisfying $ \alpha_z(k) \leqslant a_1 \exp(-a_2|k-m|_{+}^{r_2})$ for any $k\geqslant1$ and some $m=m(n)>0$, where $m=o(n)$ may diverge with $n$. • There exists a universal constant $c>0$ such that $\mathbb{E}(|n^{-1/2}\sum_{t=1}^{n} z_{t,j}|^{2}) \geqslant c$ for any $j\in[d_{z}]$.

Let $\mathbf s_{n,y}\sim \mathcal{N}(\boldsymbol{0},\boldsymbol{\Xi})$ be independent of $\mathcal{Z}_n=\{\mathbf z_1,\ldots,\mathbf z_n\}$. Define

align[align omitted — 236 chars of source]

CCW2020 gives an upper bound for $\varrho_n$ when $m$ is a fixed constant. Proposition (ref) presents a more general result that allows $m$ to diverge with $n$, whose proof is presented in the supplementary material.

propositionAssume $d_{z}\geqslant n^{\varpi}$ for some sufficiently small constant $\varpi>0$. Under {\rm AS1--AS3}, it holds that \begin{align*} \varrho_n\lesssim \frac{m^{1/3}(\log d_z)^{2/3}}{n^{1/9}}\{m^{1/6}(\log d_z)^{1/2}+m^{1/3}+(\log d_z)^{1/(3r_2)}\} \end{align*} provided that $\log d_{z}\ll\min\{m^{3r/(6+2r)}n^{7r/(18+6r)}, m^{-3r_1/(6+2r_1)}n^{7r_1/(18+6r_1)}, n^{r_2/(9-3r_2)}\}$ with $m\lesssim n^{1/9}(\log n)^{1/3}$, where $r=r_1r_2/(r_1+r_2)$ and $r_1$ and $r_2$ are specified in {\rm AS1} and {\rm AS2}, respectively.

\newtheorem{lemma}{Lemma} \setcounter{lemma}{0}

Proposition (ref) requires that $m$ involved in Assumption AS2 cannot diverge faster than $n^{1/9}(\log n)^{1/3}$. The proof of Proposition 3 is based on the widely used “large-and-small-blocks" technique in time series analysis. The key step for the proof of Proposition (ref) is to establish the associated Gaussian approximation result for the partial sum over the large blocks, see Lemma (ref) in the supplementary material. The restrictions on $\log d_z$ given in Proposition (ref) are derived from the conditions of Lemma (ref) with suitable selections of the lengths of large and small blocks. In the proofs of Propositions (ref) and (ref), and Theorem (ref), we need the following lemma whose proof is given in the supplementary material.

lemmaUnder {\rm AS1--AS3}, it holds that \begin{align} \max_{0\leqslant a\leqslant n-q}\max_{j\in[d_z]}\mathbb{P}\bigg(\max_{k\in[q]} \bigg|\sum_{t=a+1}^{a+k} z_{t,j}\bigg| \geqslant x\bigg) \lesssim & \exp(-Cq^{-1}m^{-1}x^2) + qx^{-1}\exp(-Cx^r) \nonumber\\ & +qx^{-1}\exp(-Cm^{-r_1} x^{r_1}) \end{align} for any $x>0$ and $m\leqslant q\leqslant n$, where $r=r_1r_2/(r_1+r_2)$.

Proof of Proposition 1

Recall $ T_n=\sum_{j=1}^K Z_j$ and $ \tilde T_n := \sum_{j=1}^K \tilde Z_j$. To construct Proposition 1, we need the following lemma whose proof is given in the supplementary material.

lemmaAssume Conditions {\rm 1--3} hold. Let $\tau=\tau_1\tau_2/(\tau_1+\tau_2)$. If $\log(Kpd)=o(n^{\tau/2})$ and $K^{\tau_1}\log(Kpd)=o(n^{\tau_1/2})$, then \begin{align*} |T_n-\tilde T_n| \lesssim \frac{K^{3/2}\{\log(Kpd)\}^{1/2}}{{n}^{1/2}}\max[\{\log(Kpd)\}^{1/\tau}, K\{\log(Kpd)\}^{1/\tau_1}] \end{align*} with probability at least $1-C(Kpd)^{-1}$ under $H_0$.

Recall $\bar\boldsymbol \eta=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\boldsymbol \eta_t$ and $G_K = \sum_{j=1}^K \max_{\ell\in\mathcal{L}_j} |g_\ell|^2$ with $\mathbf g=(g_1,\ldots,g_{Kpd})^{{\scriptscriptstyle {\rm \top}}} \sim {\mathcal{N}}(\boldsymbol{0}, \boldsymbol{\Sigma}_{n,K})$ where $\boldsymbol{\Sigma}_{n,K}=\tilde{n}\mathbb{E}\{(\bar\boldsymbol \eta-\boldsymbol \mu)(\bar\boldsymbol \eta-\boldsymbol \mu)^{{\scriptscriptstyle {\rm \top}}} \}$ and $\boldsymbol \mu =\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\mathbb{E}(\boldsymbol \eta_t)$. Under $H_0$, we have $\boldsymbol \mu=\boldsymbol{0}$. Thus $\boldsymbol{\Sigma}_{n,K}=\tilde{n}\mathbb{E}(\bar\boldsymbol \eta\bar\boldsymbol \eta^{{\scriptscriptstyle {\rm \top}}})$. Define $\mathbf v := (v_1,\ldots, v_{Kpd})^{{\scriptscriptstyle {\rm \top}}}= \tilde n^{1/2} \bar\boldsymbol \eta$. Our proof includes two steps: (i) using Proposition (ref) to show $\sup_{x>0}|\mathbb{P}(\tilde T_n\leqslant x) - \mathbb{P}(G_K\leqslant x) |=o(1)$, and (ii) using Lemma (ref) to show $\sup_{x>0}|\mathbb{P}(T_n\leqslant x) - \mathbb{P}(G_K\leqslant x) |=o(1)$.

Step 1. For any $j_1,\ldots,j_K\in[pd]$ and $x>0$, let $\mathcal A_{j_1,\ldots,j_K}(x)=\{{\mathbf b} \in\mathbb{R}^{Kpd}: {\mathbf b}_{S_{j_1,\ldots,j_K}}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}_{S_{j_1,\ldots,j_K}}\leqslant x\}$ with $S_{j_1,\ldots,j_K}=\{j_1,j_2+pd,\ldots,j_K+(K-1)pd\}$. Define $\mathcal A(x;K) = \bigcap_{j_1=1}^{pd} \cdots \bigcap_{j_K=1}^{pd} \mathcal A_{j_1,\ldots,j_K}(x)$. We then have $\{\tilde T_n\leqslant x\}=\{\mathbf v\in\mathcal A(x;K)\}$ and $\{G_K\leqslant x\}=\{\mathbf g\in\mathcal A(x;K)\}$. Note that the set $\mathcal A_{j_1,\ldots,j_K}(x)$ is convex that only depends on the components in $S_{j_1,\ldots,j_K}$. For a generic integer $q\geqslant2$, denote by $\mathbb{S}^{q-1}$ the $q$-dimensional unit sphere. We can reformulate $\mathcal A_{j_1,\ldots,j_K}(x)$ as follows:

align*[align* omitted — 275 chars of source]

Define $\mathcal{F} =\bigcup_{j_1=1}^{pd}\cdots\bigcup_{j_K=1}^{pd}\{{\mathbf a}\in\mathbb{S}^{Kpd-1}:{\mathbf a}_{S_{j_1,\ldots,j_K}}\in\mathbb{S}^{K-1}\}$. Then $ \mathcal A(x;K)=\bigcap_{{\mathbf a}\in\mathcal{F}}\{{\mathbf b}\in\mathbb{R}^{Kpd}: {\mathbf a}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}\leqslant\sqrt{x}\}$. For the unit sphere $\mathbb{S}^{K-1}$ equipped with $|\cdot|_2$, it is well-known that its $\epsilon$-covering number $N_{\mathbb{S}^{K-1},\epsilon}$ satisfies $\epsilon^{-K} \leqslant N_{\mathbb{S}^{K-1},\epsilon}\leqslant (1+2\epsilon^{-1})^K$, see Lemma 5.2 of Vershynin2012. Let $\mathcal S_\epsilon$ be an $\epsilon$-net of $\mathbb{S}^{K-1}$ with cardinality $N_{\mathbb{S}^{K-1},\epsilon}$. Without loss of generality, we assume $\mathcal S_\epsilon \subset \mathbb{S}^{K-1}$. Then $\tilde{\mathcal S}_{\epsilon}^{(j_1,...,j_K)}:=\{{\mathbf a}\in\mathbb{S}^{Kpd-1}: {\mathbf a}_{S_{j_1,\ldots,j_K}}\in\mathcal S_\epsilon \}$ provides an $\epsilon$-net of $\{{\mathbf a}\in\mathbb{S}^{Kpd-1}:{\mathbf a}_{S_{j_1,\ldots,j_K}}\in\mathbb{S}^{K-1}\}$ for any given $(j_1,\ldots,j_K)\in[pd]^K$, and $|\tilde{\mathcal{S}}_{\epsilon}^{(j_1,\ldots,j_K)}|=N_{\mathbb{S}^{K-1},\epsilon}$. Furthermore, we know $\mathcal{F}_\epsilon=\bigcup_{j_1=1}^{pd}\cdots\bigcup_{j_K=1}^{pd} \tilde{\mathcal S_\epsilon}^{(j_1,\ldots,j_K)}\subset\mathcal{F}$ is an $\epsilon$-net of $\mathcal{F}$ with $|\mathcal{F}_\epsilon|$ satisfying $ \epsilon^{-K} \leqslant |\mathcal{F}_\epsilon| \leqslant \{(2+\epsilon)\epsilon^{-1}pd\}^K$. Recall $ \mathcal A(x;K)=\bigcap_{{\mathbf a}\in\mathcal{F}}\{{\mathbf b}\in\mathbb{R}^{Kpd}:{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}\leqslant\sqrt{x}\}$. Define $A_1(x) = \bigcap_{{\mathbf a}\in\mathcal{F}_\epsilon} \{{\mathbf b}\in\mathbb{R}^{Kpd}:{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}\leqslant(1-\epsilon)\sqrt{x}\} $ and $A_2(x) = \bigcap_{{\mathbf a}\in\mathcal{F}_\epsilon} \{{\mathbf b}\in\mathbb{R}^{Kpd}:{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} {\mathbf b}\leqslant\sqrt{x}\}$. We can show that $ A_1(x) \subset \mathcal A(x;K) \subset A_2(x)$. Define

align*[align* omitted — 283 chars of source]

It then holds that

align*[align* omitted — 365 chars of source]

Analogously, we also have $\mathbb{P}\{\mathbf v\in\mathcal A(x;K)\} \geqslant \mathbb{P}\{\mathbf g\in\mathcal A(x;K)\} - \rho_{1,g}(x) - \rho_{2,g}(x)$. Hence, we have

align[align omitted — 156 chars of source]

We set $\epsilon=n^{-1}$ throughout the following arguments. Then $|\mathcal{F}_\epsilon|\geqslant n^K$. Due to $\tau_2\in(0,1]$, it holds that $K\lesssim (\log|\mathcal{F}_\epsilon|)^{1/\tau_2}$. Note that $K\lesssim n^{1/9}(\log n)^{1/3}$. By Proposition (ref) with $m=K$, $d_z\lesssim (npd)^K$ and $(r_1,r_2)=(\tau_1,\tau_2)$, we have

align*[align* omitted — 471 chars of source]

provided that $\log(npd)\ll\min\{K^{(\tau-6)/(6+2\tau)}n^{7\tau/(18+6\tau)}, K^{-(6+5\tau_1)/(6+2\tau_1)}n^{7\tau_1/(18+6\tau_1)}, K^{-1}n^{\tau_2/(9-3\tau_2)}\}$. To make $\sup_{x>0}\rho_{1,g}(x)=o(1)$, we need to require $\log(npd)\ll \min\{n^{2/21}K^{-10/7}, n^{\tau_2/(3+6\tau_2)}K^{-(1+3\tau_2)/(1+2\tau_2)}\}$. Notice that $ \rho_{2,g}(x) = \mathbb{P}\{(1-\epsilon)\sqrt{x} < \max_{{\mathbf a}\in\mathcal{F}_\epsilon}{\mathbf a}^{{\scriptscriptstyle {\rm \top}}}\mathbf g \leqslant \sqrt{x}\}$. If $x\leqslant K^3\{\log(npd)\}^{3}$, by Nazarov's inequality (Lemma A.1, CCK2017), we have $ \rho_{2,g}(x) \leqslant C\epsilon\sqrt{x\log |\mathcal{F}_\epsilon|} \lesssim n^{-1}K^2\{\log(npd)\}^2$. If $x>K^3\{\log(npd)\}^3$, by Markov inequality, we have

align*[align* omitted — 392 chars of source]

where the last step is based on Lemma 7.4 in FSZ2018. Hence, $\sup_{x>0}\rho_{2,g}(x)=o(1)$ if $\log(npd)\ll \min\{n^{2/21}K^{-10/7}, n^{\tau_2/(3+6\tau_2)}K^{-(1+3\tau_2)/(1+2\tau_2)}\}$. Due to $|\mathbb{P}(\tilde T_n\leqslant x) - \mathbb{P}(G_K\leqslant x) | =|\mathbb{P}\{\mathbf v\in\mathcal A(x;K)\} - \mathbb{P}\{\mathbf g\in\mathcal A(x;K)\}|$, (ref) implies $$ \sup_{x>0}|\mathbb{P}(\tilde T_n\leqslant x) - \mathbb{P}(G_K\leqslant x) |= o(1)$$ provided that $\log(npd)\ll\min\{K^{(\tau-6)/(6+2\tau)}n^{7\tau/(18+6\tau)}, K^{-(6+5\tau_1)/(6+2\tau_1)}n^{7\tau_1/(18+6\tau_1)}, K^{-1}n^{\tau_2/(9-3\tau_2)}, \\ K^{-(1+3\tau_2)/(1+2\tau_2)}n^{\tau_2/(3+6\tau_2)}, K^{-10/7}n^{2/21}\}$.

Step 2. For any $\zeta>0$, we have

align[align omitted — 309 chars of source]

Note that $K=o(n)$. Selecting $\zeta=CK^{3/2}\{\log(npd)\}^{1/2}n^{-1/2}\max[\{\log(npd)\}^{1/\tau}, K\{\log(npd)\}^{1/\tau_1}]$ for some sufficiently large constant $C>0$, Lemma (ref) yields $ \mathbb{P}( |T_n - \tilde T_n| > \zeta)= o(1)$. In the sequel, we will consider $\mathbb{P}(x-\zeta<G_K\leqslant x+\zeta)$ under the scenarios $x\leqslant \zeta$ and $x>\zeta$, respectively. Notice that $(1,0,\ldots,0)^{\scriptscriptstyle {\rm \top}}\in\mathcal{F}$ and $(-1,0,\ldots,0)^{\scriptscriptstyle {\rm \top}}\in\mathcal{F}$. Recall that $\mathbf g=(g_1,\ldots,g_{Kpd})^{\scriptscriptstyle {\rm \top}}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_{n,K})$ and $\{G_K\leqslant x\}=\{\max_{{\mathbf a}\in\mathcal{F}}{\mathbf a}^{\scriptscriptstyle {\rm \top}}\mathbf g\leqslant\sqrt{x}\}$ for any $x>0$. Then we have

align[align omitted — 455 chars of source]

where the last step is due to the anti-concentration inequality of normal random variable. For any $x>\zeta$, it holds that

align*[align* omitted — 1,225 chars of source]

Recall $|\mathcal{F}_\epsilon| \leqslant \{(2+\epsilon)\epsilon^{-1}pd\}^K$ with $\epsilon=n^{-1}$. By Nazarov's inequality, we have $ \sup_{x>\zeta}\mathbb{P}\{(1-\epsilon)(\sqrt{x}-\sqrt{\zeta}) < \max_{{\mathbf a}\in\mathcal{F}_{\epsilon}}{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} \mathbf g \leqslant (1-\epsilon)\sqrt{x}\} \lesssim \sqrt{\zeta K\log(npd)}$ and $\sup_{x>\zeta}\mathbb{P}(\sqrt{x} < \max_{{\mathbf a}\in\mathcal{F}_{\epsilon}}{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} \mathbf g \leqslant \sqrt{x} + \sqrt{\zeta})\lesssim \sqrt{\zeta K\log(npd)}$. Due to $ \mathbb{P}\{(1-\epsilon)\sqrt{x} < \max_{{\mathbf a}\in\mathcal{F}_{\epsilon}}{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} \mathbf g \leqslant \sqrt{x} + \sqrt{\zeta}\} =\rho_{2,g}(x) + \mathbb{P}(\sqrt{x} < \max_{{\mathbf a}\in\mathcal{F}_{\epsilon}}{\mathbf a}^{{\scriptscriptstyle {\rm \top}}} \mathbf g \leqslant \sqrt{x} + \sqrt{\zeta})$, together with (ref), we have

align*[align* omitted — 159 chars of source]

If $\log(npd)\ll \min\{K^{-5\tau/(3\tau+2)}n^{\tau/(3\tau+2)},K^{-7\tau_1/(3\tau_1+2)}n^{\tau_1/(3\tau_1+2)}\}$, then $\zeta K\log(npd)=o(1)$. By (ref), to make $ \sup_{x>0}|\mathbb{P}( T_n\leqslant x) - \mathbb{P}(G_K\leqslant x)|=o(1)$, we need to require $K\lesssim n^{1/9}(\log n)^{1/3}$ and

align*[align* omitted — 404 chars of source]

Due to $\log(npd)\rightarrow\infty$ as $n\rightarrow\infty$, $K$ should satisfy the restriction $ K \ll n^{f_1(\tau_1,\tau_2)}$ with $f_1(\tau_1,\tau_2)$ specified in (ref). If $K=O(n^\delta)$ for some constant $0\leqslant\delta<f_1(\tau_1,\tau_2)$, there exists a constant $c>0$ depending on $(\tau_1,\tau_2,\delta)$ such that $ \sup_{x>0}|\mathbb{P}( T_n\leqslant x) - \mathbb{P}(G_K\leqslant x)|=o(1)$ provided that $\log(pd)\ll n^c$. $\hfill\Box$

Proof of Proposition 2

Write $\boldsymbol \mu=(\mu_1,\ldots,\mu_{Kpd})^{\scriptscriptstyle {\rm \top}} =\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\mathbb{E}(\boldsymbol \eta_t)$. Define $\boldsymbol{\Sigma}_{n,K}^* = \sum_{j=-\tilde n+1}^{\tilde n-1} \mathcal K(j/b_n) \mathbf H_j$, where $\mathbf H_j=\tilde{n}^{-1} \sum_{t=j+1}^{\tilde n}\mathbb{E}\{(\boldsymbol \eta_t-\boldsymbol \mu)(\boldsymbol \eta_{t-j}-\boldsymbol \mu)^{{\scriptscriptstyle {\rm \top}}}\}$ if $j\geqslant0$ and $\mathbf H_j=\tilde{n}^{-1} \sum_{t=-j+1}^{\tilde n}\mathbb{E}\{(\boldsymbol \eta_{t+j}-\boldsymbol \mu)(\boldsymbol \eta_t-\boldsymbol \mu)^{{\scriptscriptstyle {\rm \top}}}\}$ if $j<0$. By the triangle inequality, we have

align*[align* omitted — 240 chars of source]

Let $\tau_*=(\tau_1\tau_2)/(\tau_1+2\tau_2)$. As we will show later in Sections (ref) and (ref), $|\boldsymbol{\Sigma}_{n,K}^*-\boldsymbol{\Sigma}_{n,K}|_\infty \lesssim n^{-\rho}K^{2}$, and

align*[align* omitted — 429 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$. Therefore, $ K^3\{\log(npd)\}^2|\widehat\boldsymbol{\Sigma}_{n,K} - \boldsymbol{\Sigma}_{n,K}|_\infty =o_{\rm p}(1)$ provided that $0<\rho < (\vartheta-1)/(3\vartheta-2)$ and

align*[align* omitted — 361 chars of source]

Due to $\log(npd)\rightarrow\infty$ as $n\rightarrow\infty$, $K$ should satisfy the restriction $ K \ll n^{f_2(\rho,\vartheta)}$ with $f_2(\rho,\vartheta)$ specified in (ref). If $K=O(n^{\delta})$ for some constant $0\leqslant\delta<f_2(\rho,\vartheta)$, there exists a constant $c>0$ depending on $(\tau_1,\tau_2,\rho,\vartheta,\delta)$ such that $ K^3\{\log(npd)\}^2|\widehat\boldsymbol{\Sigma}_{n,K} - \boldsymbol{\Sigma}_{n,K}|_\infty=o_{{\rm p}}(1)$ provided that $\log(pd)\ll n^c$. $\hfill\Box$

Convergence rate of $|\widehat{\boldsymbol{\Sigma}}_{n,K}-\boldsymbol{\Sigma}_{n,K}^*|_\infty$.

Without lose of generality, we can assume $\boldsymbol \mu=\boldsymbol{0}$. Recall that $\widehat{\boldsymbol{\Sigma}}_{n,K} = \sum_{j=-\tilde n+1}^{\tilde n-1} \mathcal K(j/b_n) \widehat{\mathbf H}_j$, where $\widehat \mathbf H_j=\tilde{n}^{-1}\sum_{t=j+1}^{\tilde{n}}(\boldsymbol \eta_t-\bar{\boldsymbol \eta})(\boldsymbol \eta_{t-j}-\bar{\boldsymbol \eta})^{{\scriptscriptstyle {\rm \top}}}$ if $j\geqslant0$, $\widehat \mathbf H_j=\tilde{n}^{-1}\sum_{t=-j+1}^{\tilde{n}}({\boldsymbol \eta}_{t+j}-\bar{\boldsymbol \eta})({\boldsymbol \eta}_{t}-\bar{\boldsymbol \eta})^{{\scriptscriptstyle {\rm \top}}}$ otherwise, $\tilde{n}=n-K$ and $\bar{{\boldsymbol \eta}}=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}} {\boldsymbol \eta}_t$. By the triangle inequality, it holds that

align*[align* omitted — 1,127 chars of source]

In the sequel, we will specify the convergence rate of ${\rm I}$, ${\rm II}$, ${\rm III}$ and ${\rm IV}$ respectively. Recall $\boldsymbol \eta_t=(\eta_{t,1},\ldots,\eta_{t,Kpd})^{\scriptscriptstyle {\rm \top}}$.

Convergence rate of ${\rm I}$. Given $\ell_1,\ell_2\in[Kpd]$, we define $\psi_{t,j}=\eta_{t+j,\ell_1}\eta_{t,\ell_2} -\mathbb{E}(\eta_{t+j,\ell_1}\eta_{t,\ell_2})$. For any $M=o(n)\rightarrow\infty$ satisfying $M\gtrsim K$ and $b_n=o(M)$, we have

align[align omitted — 655 chars of source]

for any $x>0$. Lemma 2 of CTW2013 yields $ \max_{0\leqslant j\leqslant \tilde{n}-1}\max_{t\in[\tilde{n}-j]}\mathbb{P}(|\psi_{t,j}|>x)\leqslant C\exp(-Cx^{\tau_1/2})$ for any $x>0$. By Condition 4 and $b_n\asymp n^\rho$ for some $\rho\in(0,1)$, we have $\sum_{j=M+1}^{\tilde{n}-1}\mathcal{K}(j/b_n)\lesssim \sum_{j=M+1}^{\tilde{n}-1}(j/b_n)^{-\vartheta} \lesssim n^{\rho\vartheta} M^{1-\vartheta}$. Analogous to Lemma 4 of CQYZ2018, we can show that

align[align omitted — 666 chars of source]

for any $x>0$. Write $D_n=\sum_{j=0}^M|\mathcal{K}(j/b_n)|$ and $\tau_*=(\tau_1\tau_2)/(\tau_1+2\tau_2)$. It is easy to see $D_n\lesssim b_n\asymp n^\rho$. For each given $j$, we observe that $\{\psi_{t,j}\}$ is also an $\alpha$-mixing sequence and its $\alpha$-mixing coefficients $\tilde{\alpha}_{\psi_{t,j}}(k)\leqslant \tilde{\alpha}_K(|k-j|_+)\leqslant C_3\exp(-C_4|k-j-K|_+^{\tau_2})$, where $\tilde{\alpha}_K(\cdot)$ is the $\alpha$-mixing coefficients of the process $\{\boldsymbol \eta_t\}$ defined in (ref). By Bonferroni inequality and Lemma (ref) with $q=\tilde{n}-j$, $m=j+K$, $r_1=\tau_1/2$, $r_2=\tau_2$ and $r=\tau_*$, we have

align*[align* omitted — 531 chars of source]

for any $x>0$. Together with (ref) and (ref), it holds that

align*[align* omitted — 663 chars of source]

for any $x>0$, which implies that

align[align omitted — 363 chars of source]

To make ${\rm I}$ converge as fast as possible, we need to specify the optimal $M$ in (ref). If $ \log(npd) \leqslant n^{(1-\rho)(\vartheta-1)\tau_1/\{\vartheta(4-\tau_1)\}}$, with selecting $ M\asymp n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)}$, we have

align*[align* omitted — 249 chars of source]

If $ \log(npd) > n^{(1-\rho)(\vartheta-1)\tau_1/\{\vartheta(4-\tau_1)\}}$, with selecting $ M \asymp n^{(1-\rho+\rho\vartheta)/\vartheta}$, we have

align*[align* omitted — 199 chars of source]

Therefore, we can conclude that

align*[align* omitted — 357 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$.

Convergence rates of ${\rm II}$ and ${\rm III}$. Given $\ell_1,\ell_2\in[Kpd]$, write

align*[align* omitted — 202 chars of source]

By Bonferroni inequality and the triangle inequality, it holds that

align*[align* omitted — 609 chars of source]

for any $x>0$. Note that $\sum_{j=0}^{\tilde{n}-1}|\mathcal{K}(j/b_n)|\lesssim b_n\asymp n^\rho $. By Bonferroni inequality, the triangle inequality and Lemma (ref), we have

align*[align* omitted — 733 chars of source]

for any $x>0$, where $\tau=\tau_1\tau_2/(\tau_1+\tau_2)$. Condition 1 yields that $\sup_{t\in[\tilde{n}]}\sup_{\ell\in[Kpd]}\mathbb{E}(|\eta_{t,\ell}|)\leqslant C$. Analogously, it holds that

align*[align* omitted — 234 chars of source]

for any $x>0$. Therefore, by Bonferroni inequality, we have

align*[align* omitted — 556 chars of source]

for any $x>0$, which implies that

align*[align* omitted — 438 chars of source]

Note that $M\gtrsim K$, $\tau_1\in(0,1]$ and $\tau_*<\tau$ in (ref) and $K=o(n)$. Then

align*[align* omitted — 358 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$. Similarly, we also have

align*[align* omitted — 359 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$.

Convergence rate of ${\rm IV}$. Given $\ell_1,\ell_2\in[Kpd]$, write

align*[align* omitted — 190 chars of source]

By Bonferroni inequality and the triangle inequality, it holds that

align*[align* omitted — 207 chars of source]

for any $x>0$. Identical to the arguments for deriving the upper bound of ${\rm II}_{1,\ell_1,\ell_2}(x)$, we know the same upper bound also holds for $\mathbb{P}\{{\rm IV}(\ell_1,\ell_2)>x\}$. Hence, we have

align*[align* omitted — 358 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$.

Therefore, we can conclude that

align*[align* omitted — 522 chars of source]

Identically, we can also show

align*[align* omitted — 477 chars of source]

Hence, we have

align*[align* omitted — 429 chars of source]

provided that $K\lesssim n^{(2\rho\vartheta+1-2\rho)/(2\vartheta-1)}\{\log(npd)\}^{(4-\tau_1)/(2\tau_1\vartheta-\tau_1)} \wedge n^{(1-\rho+\rho\vartheta)/\vartheta}$. $\hfill\Box$

Convergence rate of $|\boldsymbol{\Sigma}_{n,K}^*-\boldsymbol{\Sigma}_{n,K}|_\infty$.

Note that $\boldsymbol{\Sigma}_{n,K} = \tilde{n}\mathbb{E}\{(\bar\boldsymbol \eta-\boldsymbol \mu)(\bar\boldsymbol \eta-\boldsymbol \mu)^{\scriptscriptstyle {\rm \top}})$, $\mathbf H_j=\tilde{n}^{-1} \sum_{t=j+1}^{\tilde n}\mathbb{E}\{(\boldsymbol \eta_t-\boldsymbol \mu)(\boldsymbol \eta_{t-j}-\boldsymbol \mu)^{{\scriptscriptstyle {\rm \top}}}\}$ if $j\geqslant0$ and $\mathbf H_j=\tilde{n}^{-1} \sum_{t=-j+1}^{\tilde n}\mathbb{E}\{(\boldsymbol \eta_{t+j}-\boldsymbol \mu)(\boldsymbol \eta_t-\boldsymbol \mu)^{{\scriptscriptstyle {\rm \top}}}\}$ if $j<0$, where $\bar\boldsymbol \eta=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\boldsymbol \eta_t$, $\boldsymbol \mu=\tilde{n}^{-1}\sum_{t=1}^{\tilde{n}}\mathbb{E}(\boldsymbol \eta_t)$ and $\boldsymbol \eta_t=(\eta_{t,1},\ldots,\eta_{t,Kpd})^{\scriptscriptstyle {\rm \top}}$. We write $\boldsymbol{\Sigma}_{n,K}=\{\sigma_{n,K}(\ell_1,\ell_2)\}_{(Kpd)\times(Kpd)}$, $\mathbf H_j=\{H_j(\ell_1,\ell_2)\}_{(Kpd)\times(Kpd)}$ and $\mathring\eta_{t,\ell}=\eta_{t,\ell}-\mathbb{E}(\eta_{t,\ell})$. For any $\ell_1,\ell_2\in[Kpd]$, it holds that

align*[align* omitted — 763 chars of source]

By Davydov's inequality, $|H_{j}(\ell_1,\ell_2)| \leqslant \tilde n^{-1}\sum_{t=j+1}^{\tilde n}|\mathbb{E}(\mathring\eta_{t,\ell_1}\mathring\eta_{t-j,\ell_2})| \lesssim \tilde n^{-1}(\tilde n-j)\exp(-C|j-K|_+^{\tau_2})$ for any $j\geqslant 1$. This bound also holds for $|H_{-j}(\ell_1,\ell_2)|$ with $j\geqslant 1$. Observe that $ \boldsymbol{\Sigma}_{n,K}^* := \{\sigma_{n,K}^*(\ell_1,\ell_2)\}_{(Kpd)\times(Kpd)}=\sum_{j=-\tilde n+1}^{\tilde n-1} \mathcal K(j/b_n) \mathbf H_j$ and $\mathcal K(\cdot)$ is symmetric with $\mathcal K(0)=1$. By the triangle inequality and Condition 4,

align*[align* omitted — 479 chars of source]

Thus $|\boldsymbol{\Sigma}_{n,K}^*-\boldsymbol{\Sigma}_{n,K}|_\infty \lesssim b_n^{-1}K^{2}$. $\hfill\Box$

Proof of Theorem 1

Recall $G_K = \sum_{j=1}^{K} \max_{\ell\in\mathcal{L}_j} |g_\ell|^2$ and $\hat{G}_K = \sum_{j=1}^{K} \max_{\ell\in\mathcal{L}_j} |\hat{g}_\ell|^2$ with $ \mathbf g=({g}_1,\ldots,{g}_{Kpd})^{\scriptscriptstyle {\rm \top}} \sim {\mathcal{N}}(\boldsymbol{0},\boldsymbol{\Sigma}_{n,K})$ and $ \hat{\mathbf g}=(\hat{g}_1,\ldots,\hat{g}_{Kpd})^{\scriptscriptstyle {\rm \top}} \sim {\mathcal{N}}(\boldsymbol{0},\widehat\boldsymbol{\Sigma}_{n,K})$. As shown in Proposition 1, $\sup_{x>0}|\mathbb{P}(T_n\leqslant x)-\mathbb{P}(G_K\leqslant x)|=o(1)$. Write $\mathcal{X}_n=\{{\mathbf x}_1,\ldots,{\mathbf x}_n\}$. To construct Theorem 1, it suffices to show $\sup_{x>0} |\mathbb{P}(G_K\leqslant x) - \mathbb{P}(\hat G_K\leqslant x\,|\,\mathcal{X}_n) | =o(1)$. Recall $\rho_{2,g}(x)=|\mathbb{P}\{\mathbf g\in A_2(x)\} - \mathbb{P}\{\mathbf g\in A_1(x)\}|$ for $A_1(x)$ and $A_2(x)$ defined in Section (ref). Here we also define $$ \rho_{3,g}(x) := |\mathbb{P}\{\mathbf g\in A_1(x)\} - \mathbb{P}\{\hat\mathbf g\in A_1(x)\,|\,\mathcal{X}_n\}|\vee|\mathbb{P}\{\mathbf g\in A_2(x)\} - \mathbb{P}\{\hat\mathbf g\in A_2(x)\,|\,\mathcal{X}_n\}|\,.$$ Identical to the result $\{G_K\leqslant x\}=\{\mathbf g\in\mathcal{A}(x;K)\}$ stated in Section (ref), we also have $\{\hat{G}_K\leqslant x\}=\{\hat{\mathbf g}\in\mathcal{A}(x;K)\}$ for any $x>0$, where $\mathcal{A}(x;K)$ is defined in Section (ref). Then it holds that

align*[align* omitted — 527 chars of source]

for any $x>0$. Similarly, we can obtain the reverse inequality. Notice that we have shown in Section (ref) that $\sup_{x>0}\rho_{2,g}(x)=o(1)$. Therefore,

align*[align* omitted — 199 chars of source]

By Lemma 13 of CCW2020, it holds that

align*[align* omitted — 556 chars of source]

with $\Delta_n=\max_{{\mathbf a}_1,{\mathbf a}_2\in\mathcal{F}} | {\mathbf a}_1^{{\scriptscriptstyle {\rm \top}}} (\boldsymbol{\Sigma}_{n,K}-\widehat \boldsymbol{\Sigma}_{n,K}) {\mathbf a}_2|$, where $\mathcal{F}$ is defined in Section (ref). Recall $|{\mathbf a}|_0\leqslant K$ and $|{\mathbf a}|_2=1$ for any ${\mathbf a}\in\mathcal{F}$. Thus, $ |{\mathbf a}_1^{{\scriptscriptstyle {\rm \top}}}(\boldsymbol{\Sigma}_{n,K}-\widehat\boldsymbol{\Sigma}_{n,K}){\mathbf a}_2| \leqslant |{\mathbf a}_1|_1 |{\mathbf a}_2|_1 |\boldsymbol{\Sigma}_{n,K}-\widehat\boldsymbol{\Sigma}_{n,K}|_\infty \leqslant K|\boldsymbol{\Sigma}_{n,K}-\widehat\boldsymbol{\Sigma}_{n,K}|_\infty$. Then we have $\sup_{x>0} |\mathbb{P}\{\mathbf g\in A_1(x)\} - \mathbb{P}\{\hat\mathbf g\in A_1(x)\,|\,\mathcal{X}_n\}|\lesssim K|\boldsymbol{\Sigma}_{n,K}-\widehat\boldsymbol{\Sigma}_{n,K}|_\infty^{1/3}\{\log(npd)\}^{2/3}$. Analogously, we also have $\sup_{x>0} |\mathbb{P}\{\mathbf g\in A_2(x)\} - \mathbb{P}\{\hat\mathbf g\in A_2(x)\,|\,\mathcal{X}_n\}|\lesssim K|\boldsymbol{\Sigma}_{n,K}-\widehat\boldsymbol{\Sigma}_{n,K}|_\infty^{1/3}\{\log(npd)\}^{2/3}$. Hence,

align[align omitted — 215 chars of source]

By Proposition 2, we complete the proof. $\hfill\Box$

Proof of Theorem 2

Recall that $\mathcal{X}_n=\{{\mathbf x}_1,\ldots,{\mathbf x}_n\}$ and $\hat{G}_K = \sum_{j=1}^{K} \max_{\ell\in\mathcal{L}_j} |\hat{g}_\ell|^2$ with $\hat{\mathbf g}=(\hat{g}_1,\ldots,\hat{g}_{Kpd})^{\scriptscriptstyle {\rm \top}}$. By Bonferroni inequality, we have

align*[align* omitted — 325 chars of source]

for any $x>0$. Since $\hat\mathbf g\sim {\mathcal{N}}(\boldsymbol{0},\widehat{\boldsymbol{\Sigma}}_{n,K})$ with $\widehat{\boldsymbol{\Sigma}}_{n,K}=\{\hat{\sigma}_{n,K}(\ell_1,\ell_2)\}_{Kpd\times Kpd}$, then $$ \mathbb{E}\bigg(\max_{\ell\in\mathcal{L}_j}|\hat{g}_\ell|\,\bigg|\,\mathcal{X}_n\bigg) \leqslant \big[1+\{2\log(pd)\}^{-1}\big]\{2\log(pd)\}^{1/2}\max_{\ell\in\mathcal{L}_j}\{\hat{\sigma}_{n,K}(\ell,\ell)\}^{1/2}$$ for any $j\in[K]$. Recall $\boldsymbol{\Sigma}_{n,K}=\{{\sigma}_{n,K}(\ell_1,\ell_2)\}_{Kpd\times Kpd}$ and $\varrho=\max_{\ell\in[Kpd]}\sigma_{n,K}(\ell,\ell)$. Define an event $$\mathcal{E}_0(\nu)=\bigg\{\max_{\ell\in[Kpd]}\bigg|\frac{\hat{\sigma}_{n,K}(\ell,\ell)}{\sigma_{n,K}(\ell,\ell)} -1\bigg| \leqslant \nu \bigg\}\,,$$ where $\nu>0$ and $\nu\asymp \{K\log(pd)\}^{-1}$. As shown in Proposition 2, $ \max_{\ell\in[Kpd]}|\hat\sigma_{n,K}(\ell,\ell)-\sigma_{n,K}(\ell,\ell)| =o_{\rm p}[K^{-3}\{\log(npd)\}^{-2}]=o_{\rm p}(\nu)$. From Condition 3, we have $\min_{\ell\in[Kpd]}\sigma_{n,K}(\ell,\ell)\geqslant C$, where $C$ is a positive constant. It holds that

align*[align* omitted — 265 chars of source]

Thus $\mathbb{P}\{\mathcal{E}_0(\nu)^c\} \to 0$ as $n\rightarrow\infty$. Restricted on $\mathcal{E}_0(\nu)$, it holds that

align*[align* omitted — 205 chars of source]

By Borell inequality for Gaussian process, it holds that

align*[align* omitted — 307 chars of source]

for any $x>0$. Let $x_*= K(1+\nu)\varrho([1+\{2\log(pd)\}^{-1}]\{2\log(pd)\}^{1/2} + \{2\log(4K/\alpha)\}^{1/2} )^2$. Restricted on $\mathcal{E}_0(\nu)$, we have

align*[align* omitted — 239 chars of source]

which yields that

align*[align* omitted — 915 chars of source]

Since $\mathbb{P}\{\mathcal{E}_0(\nu)^{\rm c}\,|\,\mathcal{X}_n\}=o_{\rm p}(1)$, then $\mathbb{P}\{\mathcal{E}_0(\nu)^{\rm c}\,|\,\mathcal{X}_n\}\leqslant \alpha/4$ with probability approaching one. Hence, $\mathbb{P}(\hat{G}_K>x_*\,|\,\mathcal{X}_n)\leqslant 5\alpha/6$ with probability approaching one. Following the definition of $\hat{\rm cv}_\alpha$, it holds with probability approaching one that

align[align omitted — 131 chars of source]

with $\lambda(K,p,d,\alpha)=\{2\log(pd)\}^{1/2} +\{2\log(4K/\alpha)\}^{1/2}$.

We next specify the lower bound of $T_n$. Recall that $T_n=n\sum_{j=1}^K|\hat{\boldsymbol{\gamma}}_j|_\infty^2= \sum_{j=1}^{K} \max_{\ell\in\mathcal{L}_j} (n^{1/2}u_\ell)^2$, where $\mathbf u=(u_1,\ldots,u_{Kpd})^{\scriptscriptstyle {\rm \top}}=(\hat\boldsymbol{\gamma}_1^{\scriptscriptstyle {\rm \top}},\ldots,\hat\boldsymbol{\gamma}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ with $\hat{\boldsymbol{\gamma}}_j=(n-j)^{-1}\sum_{t=1}^{n-j}{\rm vec}\{\boldsymbol{\phi}({\mathbf x}_t){\mathbf x}_{t+j}^{\scriptscriptstyle {\rm \top}}\}$. Let $\tilde\mathbf u=(\tilde{u}_{1},\ldots,\tilde{u}_{Kpd})^{\scriptscriptstyle {\rm \top}}=(\boldsymbol{\gamma}_1^{\scriptscriptstyle {\rm \top}},\ldots,\boldsymbol{\gamma}_K^{\scriptscriptstyle {\rm \top}})^{\scriptscriptstyle {\rm \top}}$ with $\boldsymbol{\gamma}_j=(n-j)^{-1}\sum_{t=1}^{n-j}\mathbb{E}[{\rm vec}\{\boldsymbol{\phi}({\mathbf x}_t){\mathbf x}_{t+j}^{\scriptscriptstyle {\rm \top}}\}]$. Define $\ell_j^*=\arg\max_{\ell\in\mathcal{L}_j}|\tilde{u}_\ell|$ for $j\in[K]$. By Cauchy-Schwarz inequality, it holds that

align*[align* omitted — 671 chars of source]

According to the definition of $\mathbf u$ and $\tilde{\mathbf u}$, we have $n^{1/2}(u_{\ell_j^*}-\tilde{u}_{\ell_j^*})=n^{1/2}(n-j)^{-1}\sum_{t=1}^{n-j}[\phi_{l_1^*}({\mathbf x}_t)x_{t+j,l_2^*}-\mathbb{E}\{\phi_{l_1^*}({\mathbf x}_t)x_{t+j,l_2^*}\}]$ for some $l_1^*\in[d]$ and $l_2^*\in[p]$. Note that $K\ll n^{1/7}$. By Bonferroni inequality and Lemma (ref), it holds that for any $x>0$

align*[align* omitted — 535 chars of source]

with $\tau=\tau_1\tau_2/(\tau_1+\tau_2)$, which implies that $n\sum_{j=1}^{K}(u_{\ell_j^*}-\tilde{u}_{\ell_j^*})^2=O_{\rm p}(K^2\log K )$. Choose $u>0$ such that $(1+\nu)^{1/2}[1+\{2\log(pd)\}^{-1}+u]=1+\epsilon_n$ for some $\epsilon_n>0$. Due to $\sum_{j=1}^{K}\tilde{u}_{\ell_j^*}^2 \geqslant n^{-1}K\varrho\lambda^2(K,p,d,\alpha) (1+\epsilon_n)^2$ and $n\sum_{j=1}^{K}(u_{\ell_j^*}-\tilde{u}_{\ell_j^*})^2=O_{\rm p}(K^2\log K)$, by (ref), it holds with probability approaching one that

align*[align* omitted — 647 chars of source]

Notice that $\epsilon_n\rightarrow 0$ and $\varrho\lambda^2(K,p,d,\alpha)K^{-1}(\log K)^{-1}\epsilon_n^2\rightarrow\infty$. Then it holds that

align*[align* omitted — 152 chars of source]

which implies that $u\asymp \epsilon_n$. It yields that $K\varrho\lambda^2(K,p,d,\alpha)u \gg K^{3/2}(\log K)^{1/2}\varrho^{1/2}\lambda(K,p,d,\alpha)$ and $K\varrho\lambda^2(K,p,d,\alpha)u\rightarrow\infty$. Therefore, we have $ T_n-\hat{\rm cv}_\alpha > K\varrho\lambda^2(K,p,d,\alpha)u$ with probability approaching one. Hence, $ \mathbb{P}_{H_1}(T_n>\hat{\rm cv}_\alpha)\to 1$ as $n\rightarrow\infty$. $\hfill\Box$

\singlespacing

thebibliography{xx} \harvarditem{Andrews}{1991}{Andrews1991} Andrews, D. W. K. (1991). Heteroskedasticity and autocorrelation consistent covariance matrix estimation. {\em Econometrica}, {\bf 59}, 817--858. \harvarditem{Bierens}{1984}{Bierens1984} Bierens, H. J. (1984). Model specification testing of time series regressions. {\em Journal of Econometrics}, {\bf 26}, 323--353. \harvarditem{Bierens}{1990}{Bierens1990} Bierens, H. J. (1990). A consistent conditional moment test of functional form. {\em Econometrica}, {\bf 58}, 1443--1458. \harvarditem{Bierens and Ploberger}{1997}{BP1997} Bierens, H. J. and Ploberger, W. (1997). Asymptotic theory of integrated conditional moment tests. {\em Econometrica}, {\bf 65}, 1129--1151. \harvarditem[Boussama et al.]{Boussama, Fuchs and Stelzer}{2011}{Boussama_etal:2011} Boussama, F., Fuchs, F. and Stelzer, R. (2011). Stationary and geometric ergodicity of BEKK multivariate GARCH models. {\em Stochastic Processes and their Applications}, {\bf 121}, 2331--2360. \harvarditem{Box and Pierce}{1970}{BP1970} Box, G. E. P. and Pierce, D. A. (1970). Distribution of residual autocorrelations in autoregressive-integrated moving average time series models. {\em Journal of the American Statistical Association}, {\bf 65}, 1509--1526. \harvarditem[Cai et al.]{Cai, Liu and Xia}{2014}{CLX2013} Cai, T. T., Liu, W. and Xia, Y. (2014). Two-sample test of high dimensional means under dependence. {\em Journal of the Royal Statistical Society Series B}, {\bf 76}, 349--372. \harvarditem[Chang et al.(2015)]{Chang, Chen and Chen}{2015}{CCC2015} Chang, J., Chen, S. X. and Chen, X. (2015). High dimensional generalized empirical likelihood for moment restrictions with dependent data. {\em Journal of Econometrics}, {\bf 185}, 283--304. \harvarditem[Chang et al.(2021a)]{Chang, Chen, Tang and Wu}{2021}{CCTW2021} Chang, J., Chen, S. X., Tang, C. Y. and Wu, T. T. (2021a). High-dimensional empirical likelihood inference. {\it Biometrika}, {\bf108}, 127--147. \harvarditem[Chang et al.(2021b)]{Chang, Chen and Wu}{2021}{CCW2020} Chang, J., Chen, X. and Wu, M. (2021b). Central limit theorems for high dimensional dependent data. {\em arXiv:2104.12929}. \harvarditem[Chang et al.(2018a)]{Chang, Guo and Yao}{2018}{CGY2018} Chang, J., Guo, B. and Yao, Q. (2018a). Principal component analysis for second-order stationary vector time series. {\it The Annals of Statistics}, {\bf 46}, 2094-2124. \harvarditem[Chang et al.(2018b)]{Chang, Qiu, Yao and Zou}{2018}{CQYZ2018} Chang, J., Qiu, Y., Yao, Q. and Zou, T. (2018b). Confidence regions for entries of a large precision matrix. {\em Journal of Econometrics}, {\bf 206}, 57--82. \harvarditem[Chang et al.(2018c)]{Chang, Tang and Wu}{2018}{CTW2018} Chang, J., Tang, C. Y. and Wu, T. T. (2018c). A new scope of penalized empirical likelihood with high-dimensional estimating equations. {\it The Annals of Statistics}, {\bf46}, 3185--3216. \harvarditem[Chang et al.(2013)]{Chang, Tang and Wu}{2013}{CTW2013} Chang, J., Tang, C. Y. and Wu, Y. (2013). Marginal empirical likelihood and sure independence feature screening. {\em The Annals of Statistics}, {\bf 41}, 2123--2148. \harvarditem[Chang et al.(2017a)]{Chang, Yao and Zhou}{2017}{CYZ2017} Chang, J., Yao, Q. and Zhou, W. (2017a). Testing for high-dimensional white noise using maximum cross correlations. {\em Biometrika}, {\bf 104}, 111--127. \harvarditem[Chang et al.(2017b)]{Chang, Zheng, Zhou and Zhou}{2017}{CZZZ2017} Chang, J., Zheng, C., Zhou, W.-X. and Zhou, W. (2017b). Simulation-based hypothesis testing of high dimensional means under covariance heterogeneity. {\em Biometrics}, {\bf 73}, 1300--1310. \harvarditem[Chang et al.(2017c)]{Chang, Zhou, Zhou and Wang}{2017}{CZZW2017} Chang, J., Zhou, W., Zhou, W.-X. and Wang, L. (2017c). Comparing large covariance matrices under weak conditions on the dependence structure and its application to gene clustering. {\em Biometrics}, {\bf 73}, 31--41. \harvarditem{Chen and Deo}{2006}{CD2006} Chen, W. W. and Deo, R. S. (2006). The variance ratio statistic at large horizons. {\em Econometric Theory}, {\bf 22}, 206--234. \harvarditem{Chen}{2018}{Chen2018} Chen, X. (2018). Gaussian and bootstrap approximations for high-dimensional u-statistics and their applications. {\em The Annals of Statistics}, {\bf 46}, 642--678. \harvarditem{Chen and Kato}{2019}{CK2019} Chen, X. and Kato, K. (2019). Randomized incomplete u-statistics in high dimensions. {\em The Annals of Statistics}, {\bf 47}, 3127--3156. \harvarditem[Chernozhukov et al.]{Chernozhukov, Chetverikov and Kato}{2013}{CCK2013} Chernozhukov, V., Chetverikov, D. and Kato, K. (2013). Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors. {\em The Annals of Statistics}, {\bf 41}, 2786--2819. \harvarditem[Chernozhukov et al.]{Chernozhukov, Chetverikov and Kato}{2017}{CCK2017} Chernozhukov, V., Chetverikov, D. and Kato, K. (2017). Central limit theorems and bootstrap in high dimensions. {\em The Annals of Probability}, {\bf 45}, 2309--2352. \harvarditem[Chernozhukov et al.(2019)]{Chernozhukov et al.}{{2019}}{CCK2019} Chernozhukov, V., Chetverikov, D. and Kato, K. (2019). Inference on causal and structural parameters using many moment inequalities. {\em Review of Economic Studies}, {\bf 86}, 1867--1900. \harvarditem[Chernozhukov et al.(2022a)]{Chernozhukov, Chetverikov, Kato and Koike}{2022}{CCKK2019} Chernozhukov, V., Chetverikov, D., Kato, K. and Koike, Y. (2022a). Improved central limit theorem and bootstrap approximations in high dimensions. {\em The Annals of Statistics}, in press. \harvarditem[Chernozhukov et al. (2022b)]{Chernozhukov, Chetverikov and Koike}{2022}{CCK2020} Chernozhukov, V., Chetverikov, D. and Koike, Y. (2022b). Nearly optimal central limit theorem and bootstrap approximations in high dimensions. {\em The Annals of Applied Probability}, in press. \harvarditem{Cochrane}{2005}{Cochrane2005} Cochrane, J. H. (2005). {\em {Asset Pricing}}. Princeton University Press. \harvarditem{de Jong}{1996}{DeJong1996} de Jong, R. M. (1996). The Bierens tests under data dependence. {\em Journal of Econometrics}, {\bf 72}, 1--32. \harvarditem{Deng and Zhang}{2020}{DZ2020} Deng, H. and Zhang, C.-H. (2020). Beyond gaussian approximation: Bootstrap for maxima of sums of independent random vectors. {\em The Annals of Statistics}, {\bf 48}, 3643--3671. \harvarditem{Deo}{2000}{Deo2000} Deo, R. S. (2000). Spectral tests of the martingale hypothesis under conditional heteroskedasticity. {\em Journal of Econometrics}, {\bf 99}, 291--315. \harvarditem{Dom\'inguez and Lobato}{2003}{DL2003} Dom\'inguez, M. A. and Lobato, I. N. (2003). A consistent test for the martingale difference hypothesis. {\em Econometric Reviews}, {\bf 22}, 351--377. \harvarditem{Durlauf}{1991}{Durlauf1991} Durlauf, S. N. (1991). Spectral-based test for the martingale hypothesis. {\em Journal of Econometrics}, {\bf 50}, 355--376. \harvarditem{Escanciano and Lobato}{2009}{EL2009} Escanciano, J. C. and Lobato, I. N. (2009). Testing the martingale hypothesis. In Mills, T. C. and Patterson, K., Eds., {\em Palgrave Handbook of Econometrics}. Palgrave Macmillan, London. \harvarditem{Escanciano and Velasco}{2006}{EV2006a} Escanciano, J. C. and Velasco, C. (2006). Generalized spectral tests for the martingale difference hypothesis. {\em Journal of Econometrics}, {\bf 134}, 151--185. \harvarditem{Fama}{1970}{Fama1970} Fama, E. F. (1970). Efficient capital markets: A review of theory and empirical work. {\em Journal of Finance}, {\bf 25}, 383--417. \harvarditem{Fama}{1991}{Fama1991} Fama, E. F. (1991). Efficient capital markets: {II}, {\em Journal of Finance}, {\bf 46}, 1575--1617. \harvarditem{Fama}{2013}{Fama2013} Fama, E. F. (2013). Two pillars of asset pricing, {\em Nobel Prize Lecture}. \harvarditem[Fan et al.]{Fan, Shao and Zhou}{2018}{FSZ2018} Fan, J., Shao, Q.-M. and Zhou, W.-X. (2018). Are discoveries spurious? Distribution of maximum spurious correlations and their applications. {\em The Annals of Statistics}, {\bf 46}, 989--1017. \harvarditem{Fang and Koike}{2021}{FK2020} Fang, X. and Koike, Y. (2021). High-dimensional central limit theorems by Stein's method. {\em The Annals of Applied Probability}, {\bf 31}, 1660--1686. \harvarditem{Hafner and Preminger}{2009}{Hafner:Preminger:2009} Hafner, C. M. and Preminger, A. (2009). On asymptotic theory for multivariate GARCH models. {\em Journal of Multivariate Analysis}, {\bf 100}, 2044--2054. \harvarditem{Hall}{1978}{Hall1978} Hall, R. E. (1978). Stochastic implications of the life cycle-permanent income hypothesis: Theory and evidence. {\em Journal of Political Economy}, {\bf 86}, 971--987. \harvarditem[Han et al.]{Han, Linton, Oka and Whang}{2016}{HLOW2016} Han, H., Linton, O., Oka, T. and Whang, Y.-J. (2016). The cross-quantilogram: measuring quantile dependence and testing directional predictability between time series. {\em Journal of Econometrics}, {\bf 193}, 251--270. \harvarditem[Hong et al.]{Hong, Linton and Zhang}{2017}{HLZ2017} Hong, S., Linton, O. and Zhang, H. (2017). An investigation into multivariate variance ratio statistics and their application to stock market predictability. {\em Journal of Financial Econometrics}, {\bf 15}, 173--222. \harvarditem{Hong}{1996}{Hong1996} Hong, Y. (1996). Consistent testing for serial correlation of unknown form. {\em Econometrica}, {\bf 64}, 837--864. \harvarditem{Hong}{1999}{Hong1999} Hong, Y. (1999). Hypothesis testing in time series via the empirical characteristic function: A generalized spectral density approach. {\em Journal of the American Statistical Association}, {\bf 94}, 1201--1220. \harvarditem{Hong and Lee}{2003}{HL2003} Hong, Y. and Lee, T.-H. (2003). Inference on predictability of foreign exchange rate changes via generalized spectrum and nonlinear time series models. {\em Review of Economics and Statistics}, {\bf 85}, 1048--1062. \harvarditem{Hong and Lee}{2005}{HL2005} Hong, Y. and Lee, Y.-J. (2005). Generalized spectral tests for conditional mean models in time series with conditional heteroskedasticity of unknown form. {\em Review of Economic Studies}, {\bf 72}, 499--541. \harvarditem{Koul and Stute}{1999}{KS1999} Koul, H. L. and Stute, W. (1999). Nonparametric model checks for time series. {\em The Annals of Statistics}, {\bf 27}, 204--236. \harvarditem[Kuchibhotla et al.]{Kuchibhotla, Mukherjee and Banerjee}{2021}{KMB2020} Kuchibhotla, A. K., Mukherjee, S. and Banerjee, D. (2021). High-dimensional {CLT}: Improvements, non-uniform extensions and large deviations. {\em Bernoulli}, {\bf 27}, 192--217. \harvarditem{LeRoy}{1989}{LeRoy1989} LeRoy, S. F. (1989). Efficient capital markets and martingales. {\em Journal of Economic Literature}, {\bf 27}, 1583--1621. \harvarditem{Ljung and Box}{1978}{LB1978} Ljung, G. M. and Box, G. E. P. (1978). On a measure of lack of fit in time series models. {\em Biometrika}, {\bf 65}, 297--303. \harvarditem{Lo}{1997}{Lo1997} Lo, A. W. (1997). {\em Market efficiency: Stock market behaviour in theory and practice, vols. I and II}, Edward Elgar. \harvarditem{Lo and MacKinlay}{1988}{LM1988} Lo, A. W. and MacKinlay, A. C. (1988). Stock market prices do not follow random walks: Evidence from a simple specification test. {\em The Review of Financial Studies}, {\bf 1}, 41--66. \harvarditem[Lobato et al.]{Lobato, Nankervis and Savin}{2001}{LNS2001} Lobato, I., Nankervis, J. C. and Savin, N. E. (2001). Testing for autocorrelation using a modified Box-Pierce q test. {\em International Economic Review}, {\bf 42}, 187--205. \harvarditem{Newey and West}{1987}{NW1987} Newey, W. and West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. {\em Econometrica}, {\bf 55}, 703--708. \harvarditem{Park and Whang}{2005}{PW2005} Park, J. Y. and Whang, Y.-J. (2005). A test of the martingale hypothesis. {\em Studies in Nonlinear Dynamics and Econometrics}, {\bf 9}, article 2. \harvarditem{Phillips and Jin}{2014}{PJ2014} Phillips, P. C. B. and Jin, S. (2014). Testing the martingale hypothesis. {\it Journal of Business $\&$ Economic Statistics}, {\bf 32}, 537--554. \harvarditem{Poterba and Summers}{1988}{PS1988} Poterba, J. M. and Summers, L. H. (1988). Mean reversion in stock prices: Evidence and implications. {\em Journal of Financial Economics}, {\bf 22}, 27--59. \harvarditem[Shao(2011a)]{Shao}{2011a}{Shao2011a} Shao, X. (2011a). A bootstrap-assisted spectral test of white noise under unknown dependence. {\em Journal of Econometrics}, {\bf 162}, 213--224. \harvarditem[Shao(2011b)]{Shao}{2011b}{Shao2011b} Shao, X. (2011b). Testing for white noise under unknown dependence and its applications to goodness-of-fit for time series models. {\em Econometric Theory}, {\bf 27}, 312--343. \harvarditem{Stute}{1997}{Stute1997} Stute, W. (1997). Nonparametric model checks for regression. {\em The Annals of Statistics}, {\bf 25}, 613--641. \harvarditem{Vershynin}{2012}{Vershynin2012} Vershynin, R. (2012). Inroduction to the non-asymptotic analysis of random matrices. In Eldar, Y. C. and Kutyniok, G., Eds., {\em Compressed Sensing: Theory and Applications}. Cambridge University Press. \harvarditem[Wong et al.]{Wong, Li and Tewari}{2020}{Wong_etal:2020} Wong, K. C., Li, Z. and Tewari, A. (2020). Lasso guarantees for $\beta$-mixing heavy-tailed time series. {\em The Annals of Statistics}, {\bf 48}, 1124--1142. \harvarditem{Wu}{2005}{Wu2005} Wu, W.-B. (2005). Nonlinear system theory: Another look at dependence. {\em Proceedings of the National Academy of Sciences USA}, {\bf 102}, 14150--14154. \harvarditem{Yu and Chen}{2021}{YC2020} Yu, M. and Chen, X. (2021). Finite sample change point inference and identification for high dimensional mean vectors. {\em Journal of the Royal Statistical Society Series B}, {\bf 83}, 247--270. \harvarditem{Zhang and Wu}{2017}{ZW2017} Zhang, D. and Wu, W.-B. (2017). Gaussian approximation for high dimensional time series. {\em The Annals of Statistics}, {\bf 45}, 1895--1919. \harvarditem{Zhang and Cheng}{2018}{ZC2018} Zhang, X. and Cheng, G. (2018). Gaussian approximation for high dimensional vector under physical dependence. {\em Bernoulli}, {\bf 24}, 2640--2675.

\def\spacingset#1{ {#1}} \spacingset{1.6}

\if11 {

center[center omitted — 175 chars of source]

} \fi

\if01 {

center[center omitted — 125 chars of source]

} \fi

\setcounter{equation}{0}

\onehalfspacing