EconBase
← Back to paper

Max-sum tests for cross-sectional dependence of high-demensional panel data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

175,514 characters · 30 sections · 59 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Max-Sum tests for cross-sectional dependence of high-dimensional panel data

frontmatter\runtitle{Max-sum test for cross-sectional dependence} \begin{aug} , , \and \runauthor{L. Feng et al.} \address{Address of Long Feng and Binghui Liu\\ Key Laboratory of Applied Statistics of MOE $\&$ School of Mathematics and Statistics\\ Northeast Normal University\\ \phantom{E-mail:\ }\printead*{e1,e3}} \address{Address of Tiefeng Jiang\\ School of Statistics\\ 313 Ford Hall\\ 224 Church Street SE\\ Minneapolis, MN55455\\ \printead{e2}} \address{Address of Wei Xiong\\ School of Statistics\\ University of International Business and Economics\\ \printead{e4}} \end{aug} \begin{abstract} We consider a testing problem for cross-sectional dependence for high-dimensional panel data, where the number of cross-sectional units is potentially much larger than the number of observations. The cross-sectional dependence is described through a linear regression model. We study three tests named the sum test, the max test and the max-sum test, where the latter two are new. The sum test is initially proposed by Breusch and Pagan (1980). We design the max and sum tests for sparse and non-sparse residuals in the linear regressions, respectively. And the max-sum test is devised to compromise both situations on the residuals. Indeed, our simulation shows that the max-sum test outperforms the previous two tests. This makes the max-sum test very useful in practice where sparsity or not for a set of data is usually vague. Towards the theoretical analysis of the three tests, we have settled two conjectures regarding the sum of squares of sample correlation coefficients asked by Pesaran (2004 and 2008). In addition, we establish the asymptotic theory for maxima of sample correlations coefficients appeared in the linear regression model for panel data, which is also the first successful attempt to our knowledge. To study the max-sum test, we create a novel method to show asymptotic independence between maxima and sums of dependent random variables. We expect the method itself is useful for other problems of this nature. Finally, an extensive simulation study as well as a case study are carried out. They demonstrate advantages of our proposed methods in terms of both empirical powers and robustness for residuals regardless of sparsity or not. \end{abstract} \begin{keyword} \kwd{high-dimensional data} \kwd{panel data models} \kwd{hypothesis tests} \kwd{cross-sectional dependence} \kwd{asymptotic normality} \kwd{extreme-value distribution} \kwd{asymptotic independence} \kwd{max-sum test}. \end{keyword}

Introduction

In this paper we will study the cross-sectional dependence for the following linear regression model for panel data

align[align omitted — 73 chars of source]

for $i=1,\cdots,N$ and $t=1,\cdots,T$, where $i$ represents households, individuals, firms, etc., and $t$ represents time. In the literature of panel data, the index $i$ stands for {\it sections}. For each section $i$, the corresponding model is a standard multiple linear regression model, where $y_{it}\in \mathbb{R}$ is the dependent variable and $x_{it}\in \mathbb{R}^p$ is the regressor with slope parameter $\beta_i\in \mathbb{R}^p$. The first coordinate of $x_{it}$ is one if there is an intercept in the linear regression model (ref). The value of $\beta_i$ may vary across $i$. In (ref), we assume $\{\epsilon_{it};\, 1\leq t\leq T\}$ are independent and identically distributed (i.i.d.) for each section $i$. However, across sections the random errors may be dependent, that is, $\{\epsilon_{it};\, 1\leq i\leq N\}$ may be dependent for some $t.$ Such dependence is referred to as cross-sectional dependence. The objective of this paper is to test if there exists cross-sectional dependence by using a few of new methods. Before stating our results, we will introduce some background next.

In statistics and econometrics, panel data or longitudinal data are multi-dimensional data involving measurements over time, which contain observations of various phenomena over multiple time periods for the same unit, for instance, a household or a firm. In the study of panel data models, the cross-sectional dependence is an important concept, described as the interaction between cross-sectional units, which could arise from the behavioral interaction between units.

Stephan Stephan1934 argues that “in dealing with social data, we know that by virtue of their very social character, persons, groups and their characteristics are interrelated and not independent. " However, to make theoretical study easier, experts assume cross-sectional independence in various model setups HsiaoP2012, Pesaran2004. If data across individuals are dependent, inferences under the assumption of cross-sectional independence would be inaccurate and misleading; see HsiaoP2012, pesaran2015testing and the literature therein. To this end, testing the existence of cross-sectional dependence is an important task, which has attracted more attention in recent years, see, for instance, Chudik2013, moscone2009a, pesaran2015testing, Pesaran2015,ST12.

Perhaps the most widely known test for cross-sectional independence is the Lagrange Multiplier (LM) statistic proposed by Breusch and Pagan breusch1980the in 1980 (Google records 5353 citations currently). Their test statistic is the sum of squares of sample correlation coefficients between the residuals from the ordinary least square (OLS). Precisely, for each $i$, let $\hat{\beta}_i$ be the standard estimator of $\beta_i$ in the linear regression for observations $\{(y_{it},x_{it});\, t=1,\cdots, T\}$ and the quantity $\hat{\bold{\epsilon}}_{it}=y_{it}-x_{it}'\hat{\beta}_i$ denotes the residual. For each $i,j=1,\cdots,N$, define the sample correlation $\hat{\rho}_{ij}$ by

align[align omitted — 216 chars of source]

Breusch and Pagan breusch1980the propose the Lagrange multiplier test statistic defined by

align[align omitted — 73 chars of source]

To get the rejection region, we need to figure out the limiting distribution of $S_N$ as $N$ goes to infinity. Under the null hypothesis that there is no cross-sectional dependence, that is, $\{\epsilon_{it};\, 1\leq i\leq N, 1\leq t \leq T\}$ from (ref) are independent, the asymptotic distribution of $S_N$ is understood when the cross-sectional dimension $N$ is fixed and the time dimension $T$ goes to infinity. In fact, assuming that $\epsilon_{it}$'s are normally distributed, Breusch and Pagan breusch1980the show that, for fixed $N$,

align[align omitted — 45 chars of source]

in distribution as $T\to\infty$, where $d=N(N-1)/2.$ If $N$ is relatively large, the above chi-square approximation is not accurate Pesaran2004. A natural amendment is approximating $\chi^2(d)$ by the standard normal distribution: $(\chi^2(d)-d)/\sqrt{2d}$ goes to the standard normal distribution as $N$ goes to infinity. However, as both $N$ and $T$ are very large, taking limit by sending $T \to\infty$ followed by sending $N\to\infty$ is not legitimate mathematically, and the approximation may not be accurate statistically (our Remark (ref) shows such an example). For this consideration Pesaran Pesaran2004 and Pesaran {\it et al.} pesaran2008a provide two versions of normalization of $S_N$ and conjecture that both versions satisfy the central limit theorem (CLT); some of the insights why the CLTs hold can be seen, for example, from ST12 and Pesaran2015. In this paper we prove the two conjectures in Theorems (ref) and (ref). This enables us to carry out the test for cross-sectional dependence through $S_N$ in (ref). In the future, when $S_N$ is used to be a test statistic, we call it the sum test.

On the other hand, when data are sparse, experts in recent years realize that a better test than sum statistics is the maximum of sample correlation coefficients. This is confirmed in, for example, Cai2014; see also caiLiu2011Adaptive, cai2013two-sample and caiZhang2016. With this philosophy in mind, to test the cross-sectional dependence when the residuals $\hat{\bold{\epsilon}}_{it}$ are sparse, we propose statistic

align[align omitted — 68 chars of source]

where $\hat{\rho}_{ij}$ is defined as in (ref). Later, when $L_N$ is used to be a test statistic, we refer it to as the max test. Its limiting distribution is obtained as both $N$ and $T$ go to infinity under various moment conditions (Theorems (ref), (ref) and (ref)). The corresponding rejection region based on the test statistic $L_N$ is given after Theorem (ref).

In practice it is hard to tell or differentiate if a set of data is sparse. We then combine the sum test $S_N$ and the max test $L_N$ to propose another test $C_N$, which is the minimum of the $p$-values corresponding to the tests based on $S_N$ and $L_N$. We prove in Theorem (ref) that, under normalization, $S_N$ and $L_N$ are asymptotically independent as both $N$ and $T$ go to infinity. Hence the limiting distribution of $C_N$ is identified. In further discussions, when $C_N$ is used to be a test statistic, we name it the max-sum test. From simulation we see this test, taking care of both sparsity and non-sparsity cases, is better than the sum test and the max test. The tool of deriving asymptotic independence between the sum and the maximum of random variables is new to our knowledge. It seems a universal method to handle asymptotic independence between random variables of this nature.

To sum up, to test cross-sectional dependence for panel data models, in this paper we study three types of tests, {\it i.e.}, the sum test, the max test and the max-sum test. To carry the test, we have solved two open problems about the CLTs for the sum of squares of residuals; the limiting distributions of the maxima of residuals are systematically studied; a new method of studying asymptotic independence is created to develop part of the above theory successfully.

The proposed tests

Problem description

Review model (ref) that $y_{it}=x_{it}'\beta_i+\epsilon_{it}$ for $i=1,\cdots,N$ and $t=1,\cdots,T$, where $i$ indexes the cross-sectional units and $t$ indexes the observations. In this model, $y_{it}\in \mathbb{R}$ is the dependent variable, and $x_{it}\in \mathbb{R}^p$ is the non-random, exogenous regressor with slope parameter $\beta_i\in \mathbb{R}^p$ that are allowed to vary across $i$. We assume $\{\epsilon_{it};\, 1\leq t\leq T\}$ are i.i.d. real-valued random variables for each section $i$. However, across sections the random errors may be dependent, that is, $\{\epsilon_{it};\, 1\leq i\leq N\}$ may be dependent for some $t.$ Such dependence is called cross-sectional dependence. Set

align[align omitted — 163 chars of source]

for $i=1,2,\cdots, N.$ Then $\bold{x}_i$ is a $T\times p$ matrix; both $\bold{y}_i$ and $\bold{\epsilon}_{i}$ are $T$-dimensional vectors. Throughout the paper we assume that the $T$ entries of $\epsilon_i$ are i.i.d. with mean zero for each $i$. Recalling (ref), the cross-sectional independence is the same as saying that

align[align omitted — 136 chars of source]

In general, although sometimes we assume $\bold{\epsilon}_{1}$ has the normal distribution, we do not need the exact distribution of $\bold{\epsilon}_{1}$ but rather its moments.

Test statistics

First, we list some notations used in the rest of the paper. Reviewing (ref), for each $i=1,\cdots,N$, let

align[align omitted — 172 chars of source]

where $\bold{I}_T$ is the $T\times T$ identity matrix and $\bold{P}_i$ is a $T\times T$ projection matrix with $\bold{P}_i^2=\bold{P}_i$ and the rank of $\bold{P}_i$ is $T-p$. For each $i,j=1,\cdots,N$, let $\hat{\rho}_{ij}$ denote the sample correlation coefficient computed by the Ordinary Least Squares (OLS) residuals $(\hat{\bold{\epsilon}}_{i1}, \cdots, \hat{\bold{\epsilon}}_{iT})^T$ and $(\hat{\bold{\epsilon}}_{j1}, \cdots, \hat{\bold{\epsilon}}_{jT})^T$ where $\hat{\bold{\epsilon}}_{it}=y_{it}-x_{it}'\hat{\beta}_i$ for each $i$ and $t$. Under model (ref), it is easy to see that

eqnarray*[eqnarray* omitted — 111 chars of source]

for each $i$. Thus, by (ref),

align[align omitted — 352 chars of source]

In this paper, to test the null hypothesis (ref), we will study three types of tests as follows:

align[align omitted — 230 chars of source]

respectively, where

align[align omitted — 224 chars of source]

Here, $F(y)= \exp(-e^{-y/2}/\sqrt{8\pi})$ is the extreme-value distribution function of type I, also called the Gumble distribution in literature, and $\Phi(y)$ is the distribution function of $N(0,1)$.

For the sum test in (ref), we will establish that, under $H_0$ in (ref), $(S_N-\mu_N)/N$ converges weakly to the standard normal distribution when both $N$ and $T$ go to infinity with a certain restriction (Theorem (ref)), hence a level-$\alpha$ test will be performed through rejecting $H_0$ when $(S_N-\mu_N)/N$ is larger than the $1-\alpha$ quantile $z_{\alpha}= \Phi^{-1}(1-\alpha)$ of the standard normal distribution.

For the max test in (ref), under $H_0$, we will establish that $TL_N^2-4\log N+\log\log N$ has an asymptotic extreme-value distribution as both $N$ and $T$ go to infinity (Theorems (ref), (ref) and (ref)). We do not impose normality assumptions but rather moment conditions. Recall $F(y)$ is defined below (ref). A level-$\alpha$ test will then be performed by rejecting $H_0$ when $TL_N^2-4\log N+\log\log N$ is larger than the $1-\alpha$ quantile $q_{\alpha}=-\log(8\pi)-2\log\log(1-\alpha)^{-1}$ of $F(y)$.

Furthermore, for the max-sum test in (ref), its asymptotic distribution under $H_0$ is constructed based on the asymptotic independence between $(S_N-\mu_N)/N$ and $TL_N^2-4\log N+\log\log N$ as both $N$ and $T$ go to infinity (Theorem (ref) and Corollary (ref)). So a level-$\alpha$ test will be performed through rejecting $H_0$ when $C_N<1-\sqrt{1-\alpha}$.

Contributions

In this paper, for the panel data model (ref) we study the cross-sectional dependence. The asymptotic distributions of three test statistics based on residuals are established. As application, three hypothesis tests are accomplished. A real data analysis by using our results is provided. We will now further elaborate below.

In the theoretical part, we have solved two open problems on the sum of squares of residuals conjectured by economists (Pesaran2004, pesaran2008a; see also Pesaran2015, ST12). We have developed an extreme-value theory for the maximum of residuals. Further, a new method is developed to show the sum and the maximum are asymptotically independent. There are not many results in literature to show asymptotic independence between sums of and maxima of random variables. Close references are BBQ98, Xu2016. Our method, being different from earlier literature, provides a general and novel tool for showing asymptotic independence between sums of and maxima of random variables.

In application, we propose three tests on the cross-sectional dependence for high-dimensional panel data: the sum test, the max test and the max-sum test. The max test is the first high-dimensional max test for cross-sectional dependence in panel data models, which is good for sparse residuals while existing test statistics of sum types tend to fail. The sum test is useful for non-sparse residuals, which is clearly demonstrated by simulation in, for example, Pesaran2004, Pesaran2015, pesaran2008a, ST12. We are able to derive the limiting distribution of the sums in this paper.

Furthermore, the max-sum test is constructed based on the asymptotic independence between the max and the sum statistics aforementioned. It is the first max-sum test for studying cross-sectional dependence for high-dimensional panel data. The advantage is that the test works well for both sparse and non-sparse residuals. Comparing the pros and cons of the max test and the sum test, the max-sum test definitely overcomes both disadvantages. Our simulations reveal this fact clearly; see Figure (ref) and its interpretation at the last part of Section (ref). The max-sum test is particularly useful considering it is hard to quantify or determine in practice whether a data set is sparse or not.

Theoretical results

We now present the main theoretical results based on the three types of tests in the order of the sum test, the max test and the max-sum test. Their proofs are presented in Section (ref).

The limiting distribution for the sum test

Recall that the sum test described in (ref) is a classical one for testing cross-sectional dependence in panel data models. However, the asymptotic theory has not been established yet. Pesaran from Pesaran2004, pesaran2008a conjectures that $S_N$ satisfies the central limit theorem. Some insights on this aspect are given, for example, in ST12 and Pesaran2015. In the following we will present our solution to the problem as well as another one in which the details are given below. The following assumption will be needed throughout the paper. Recall a random variable $V$ is said to be continuous if $P(V=v)=0$ for every $v\in \mathbb{R}.$

align[align omitted — 421 chars of source]

If the $T$ entries of $\bold{\epsilon}_{i}$ are i.i.d. continuous random variables, by using a conditional argument, we then trivially have $P(\bold{a}'\bold{\epsilon}_{i}=0)=0$ for any $\bold{a}\in \mathbb{R}^T\backslash\{\bold{0}\}$. This implies that $P(\bold{M}\bold{\epsilon}_{i}= \bold{0})=0$ for any $l\times T$ matrix $\bold{M}\ne 0$ and $l\geq 1$. So $\hat{\rho}_{ij}$ in (ref) is well-defined if $T>p$ because $\mbox{rank}(\bold{P}_i)=T-p$; see the explanation below (ref).

Now we present our solutions to Pesaran's conjectures from Pesaran2004, pesaran2008a as follows. For mathematical rigor, we assume the parameter $T$ depends on $N$. The notation $N_k(\mu, \bold{\Sigma})$ stands for the $k$-dimensional multivariate normal distribution with mean vector $\mu$ and covariance matrix $\bold{\Sigma}$. Although the linear regression model in (ref) requires $p\geq 1$, that is, there are at least one regressors, the following theorem applies to the case that

align[align omitted — 100 chars of source]

see (ref).

theoremAssume $p\geq 0$ is fixed and $T/\sqrt{N}\to \infty$ as $N\to \infty$. Let assumption (ref) hold with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. Let $S_N$ be as in (ref) and $\mu_N$ be as in (ref), respectively. Then $(S_N-\mu_N)/N$ converges to $N(0, 1)$ in distribution as $N\to\infty$.

From the assumption, it is allowed that the number of cross-sectional units $N$ is much larger than the number of observations $T$, for example, $N$ is of order $T^{3/2}$. We apply the framework of the Lindeberg-Feller martingale CLT to study $S_N$. Although the method is simple and is easy to follow, the technical steps are very involved due to the complex nature of the sample correlation coefficients $\hat{\rho}_{ij}$ in (ref). Many computations focus on the conditional means, variances and higher moments.

Considering a possible better convergence rate than the CLT given in Theorem (ref), pesaran2008a revises the statistic $S_N$ and proposes a new one as follows.

align[align omitted — 134 chars of source]

where

align[align omitted — 296 chars of source]

Pesaran {\it et al}. pesaran2008a conjecture that $Q_N$ also satisfies the CLT. We confirm it in the next theorem.

theoremAssume the setting in Theorem (ref). Then $Q_N$ converges to $N(0, 1)$ in distribution as $N\to \infty$.

Our simulation in Figure (ref) shows that the effects of the two approximations in Theorems (ref) and (ref) are too close to be distinguishable. Theorem (ref) allows us to perform a level-$\alpha$ test by rejecting the null hypothesis from (ref) when $(S_N-\mu_N)/N$ is larger than the $1-\alpha$ quantile $z_{\alpha}= \Phi^{-1}(1-\alpha)$ of $N(0,1)$. Theorem (ref) establishes the central limit theorem of $Q_N$ under the same null hypothesis. This provides a theoretical guarantee for $Q_N$ that has been used in econometrics; see, for example, Chudik2013, moscone2009a, pesaran2015testing.

Now we make some comments.

remarkLet us see how heuristically the mean $\mu_N$ and the standard deviation $N$ in Theorem (ref) are figured out. In fact, $\mu_N$ is computed by using Lemma (ref)(i). The variance is calculated via Lemma (ref)(iv) by noticing $\frac{1}{T}\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]\to 1$ and $\frac{1}{T^2}[\mbox{tr}(\bold{P}_i\bold{P}_j)]^2\to 1$ as $N\to\infty$ and by regarding $\{\hat{\rho}_{ij}^2,\, 1\leq i<j \leq N\}$ as independent random variables although they are weakly correlated.
remarkTake $p=0$ in Theorem (ref). From (ref), $\bold{P}_i=\bold{I}_T$ for all $1\leq i \leq N$ and hence $\mbox{tr}(\bold{P}_i\bold{P}_j)=T$. Theorem (ref) then says that, assuming $T/\sqrt{N}\to \infty$, \begin{align} \frac{S_N}{N}-\frac{1}{2}N\to N\Big(-\frac{1}{2}, 1\Big) \end{align} in distribution as $N\to\infty.$ This is the trivial case that no linear regression is involved. And $\hat{\rho}_{ij} =\frac{\bold{\epsilon}_i'\bold{\epsilon}_j} {\|\bold{\epsilon}_i\|\cdot\|\bold{\epsilon}_j\|}$ from (ref), where $\bold{\epsilon}_i$ and $\bold{\epsilon}_j$ are independent Gaussian vectors with distributions $N_T(\bold{0}, \sigma_i^2\bold{I})$ and $N_T(\bold{0}, \sigma_j^2\bold{I})$, respectively. Suppose $N/T\to \gamma\in (0, \infty).$ Rewrite (ref) by using the Slutsky lemma to see \begin{align} \sum_{1\le i<j\le N}\hat{\rho}_{ij}^2-\frac{N(N-1)}{2T} \to N(0, \gamma^2) \end{align} as $N\to \infty$. This recovers the result by Schott Schott2005. A quick reminder is that our assumption “\,$T/\sqrt{N}\to \infty$" is less stringent than “$N/T\to \gamma\in (0, \infty)$". A further discussion about Schott's work is continued in the next remark.
remarkAssume $V\sim N_d(\bold{\mu}, \bold{\Sigma})$, that is, $V$ follows a $d$-dimensional multivariate normal distribution with mean vector $\bold{\mu}$ and covariance matrix $\bold{\Sigma}$. Assume a random sample of size $n$ is given. Under the assumption $d/n\to c>0$, Schott Schott2005 studies the null hypothesis that the $d$-entries of $V$ are independent, that is, $\bold{\Sigma}$ is diagonal. Comparing this with model (ref), his model corresponds to (ref). His result is stated in (ref). Jiang Jiang2019 investigates the same testing problem through the likelihood ratio test and obtains the CLT for a big class of alternative hypothesis. However, neither derivations of the above two CLTs help the proofs of Theorems (ref) and (ref) in this paper.

The scenario in the following remark is not practical. We consider it purely for mathematical purposes. They serve for further discussions.

remarkAssume $\bold{x}_1=\cdots=\bold{x}_N$. Then $\bold{P}_1=\cdots =\bold{P}_N$ and $\mbox{tr}(\bold{P}_i\bold{P}_j)=\mbox{tr}(\bold{P}_1)=T-p.$ Hence, \begin{eqnarray*} \mu_N=\frac{T}{T-p}\cdot \frac{N(N-1)}{2}=N\cdot \Big[\frac{TN}{2(T-p)}-\frac{1}{2}+o(1)\Big]. \end{eqnarray*} Theorem (ref) says that \begin{eqnarray*} \frac{S_N}{N}-\frac{TN}{2(T-p)}\to N\Big(-\frac{1}{2}, 1\Big) \end{eqnarray*} in distribution as $N\to\infty$. In particular, if $\frac{N}{T}\to c\in (0, \infty)$, then $\frac{TN}{2(T-p)}=\frac{N}{2}+\frac{1}{2}c p+o(1)$. Hence \begin{align} \frac{S_N}{N}-\frac{1}{2}N\to N\Big(\frac{c p-1}{2}, 1\Big). \end{align} A point for this extreme example is that, as these $\bold{x}_i$ are highly correlated, the CLT in (ref) is indeed different from the trivial CLT in (ref). Interestingly, the next example is completely different from this one.
remarkAssume $T=Np$. In this case, $\frac{N}{T}=\frac{1}{p}.$ Construct \begin{eqnarray*} \bold{x}_1=(\bold{I}_p, \bold{0}_{p\times (T-p)})',\ \bold{x}_2=(\bold{0}_{p\times p}, \bold{I}_p, \bold{0}_{p\times (T-2p)})',\ \cdots, \bold{x}_N=(\bold{0}_{p\times (T-p)}, \bold{I}_{p})'. \end{eqnarray*} They are $T\times p$ matrices. Then $\bold{x}_i'\bold{x}_i=\bold{I}_p$ for each $1\leq i\leq N$, and hence \begin{eqnarray*} \bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'= \begin{pmatrix} \bold{0}_{p\times p} & \cdots &\bold{0}_{p\times p} & \cdots & \bold{0}_{p\times p}\\ \vdots & \vdots & \vdots & \vdots\\ \bold{0}_{p\times p} & \cdots & \bold{I}_p & \cdots &\bold{0}_{p\times p}\\ \vdots & \vdots & \vdots & \vdots\\ \bold{0}_{p\times p} & \cdots &\bold{0}_{p\times p}& \cdots & \bold{0}_{p\times p} \end{pmatrix}, \end{eqnarray*} where each $\bold{0}_{p\times p}$ is a $p\times p$ submatrix with all entries equal to zero. In other words, we may regard $\bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'$ as an $N\times N$ matrix with each entry being a block of $p\times p$ matrix, and the only non-zero entry is the $(i, i)$-entry $\bold{I}_p$. Then $\bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'\cdot \bold{x}_j(\bold{x}_j'\bold{x}_j)^{-1}\bold{x}_j'=\bold{0}_{T\times T}$ for $i\ne j$. It follows from the definition of $\bold{P}_i$ in (ref) that $\mbox{tr}(\bold{P}_i\bold{P}_j)=T-2p$. Thus, \begin{eqnarray*} \mu_N=\frac{N(N-1)}{2}\cdot \frac{T(T-2p)}{(T-p)^2}=N\cdot \Big[\frac{N-1}{2}+o(1)\Big] \end{eqnarray*} by the assumption $T=Np$. Then \begin{eqnarray*} \frac{S_N}{N}-\frac{1}{2}N\to N\Big(-\frac{1}{2}, 1\Big) \end{eqnarray*} in distribution as $N\to\infty.$ The essence for this example demonstrates that the projection matrices $\bold{P}_i$ are orthogonal to each other contrary to the highly correlated case in Remark (ref). We see the CLT here is more like the one in the trivial case from Remark (ref) but is different from that in Remark (ref).
remarkReview (ref) that $S_N\to \chi^2(d)$ as $T \to\infty$ while $N$ is fixed, where $d=\frac{1}{2}N(N-1)$. By using the approximation $(\chi^2(d)-d)/\sqrt{2d}\to N(0,1)$ as $d\to\infty$, we see that, if taking limit above were legitimate, we would have \begin{align*} \frac{S_N-\frac{1}{2}N(N-1)}{\sqrt{N(N-1)}}\to N(0, 1). \end{align*} By the Slutsky lemma, this entails \begin{align*} \frac{S_N}{N}-\frac{1}{2}N\to N\Big(-\frac{1}{2}, 1\Big) \end{align*} in distribution. It is interesting to see this weak convergence in Remarks (ref) and (ref), but not in Remark (ref). In fact, for big data with the feature that two or more parameters are large, to study a statistic of interest, it is not always valid to send parameters to infinity one by one; see such examples in, for instance, JiangQi2015, JiangYang2013, ZhengBaiYao15.

The limiting distribution for the max test

Recall model (ref) and notations in (ref). As in Section (ref), we assume that $\bold{x}_i\bold{x}_i'$ is invertible for each $1\leq i\leq N$ and the quantity $T$ depends on $N$. From (ref) and assumption (ref), we know $\{\hat{\rho}_{ij}:1\leq i,j \leq N\}$ are invariant of $\sigma^2_1,\cdots,\sigma^2_N$, so we are able to assume, without loss of generality, $\sigma_1=\cdots=\sigma_N\equiv 1$. Review $\epsilon_{11}$ in assumption (ref) and $L_N$ in (ref). As explained in (ref), we will also consider the case $p=0$. The main results in this section are presented as follows.

theoremAssume $p\geq 0$ is fixed and $\lim_{N\to\infty}T/N=c\in (0, \infty)$. Let $\bold{\epsilon}_{1},\cdots, \bold{\epsilon}_{N}$ be i.i.d. and assumption (ref) hold with $E|\epsilon_{11}|^{\tau}<\infty$ for some $\tau>8$. Then, as $N\to\infty$, $TL_N^2-4\log N +\log\log N$ converges weakly to the distribution function $F(y)= \exp(-e^{-y/2}/\sqrt{8\pi})$, $y \in \mathbb{R}$.
theoremAssume $p\geq 0$ is fixed and $\log N=o(T^{1/5})$ as $N\to \infty$. Let $\bold{\epsilon}_{1},\cdots, \bold{\epsilon}_{N}$ be i.i.d. and assumption (ref) hold with $Ee^{\omega|\epsilon_{11}|}<\infty$ for some $\omega>0.$ Then, as $N\to\infty$, $TL_N^2-4\log N +\log\log N$ converges weakly to the distribution function $F(y)= \exp(-e^{-y/2}/\sqrt{8\pi})$, $y \in \mathbb{R}$.

We say $\xi$ is a {\it subgaussian} random variable if there exists $\sigma>0$ such that $Ee^{t\xi}\leq e^{\sigma^2t^2/2}$ for all $t\in \mathbb{R}.$ By the Markov inequality, it is easy to see $P(|\xi|\geq x)\leq 2e^{-x^2/(2\sigma^2)}$ for all $x>0$. As a consequence, $Ee^{\theta \xi^2} < \infty$ for all $\theta < \frac{1}{2\sigma^2}$. Obviously, bounded random variables and Gaussian random variables are all subgaussian random variables.

theoremAssume $p\geq 0$ is fixed and $\log N=o(T^{1/3})$ as $N\to \infty$. Let $\bold{\epsilon}_{1},\cdots, \bold{\epsilon}_{N}$ be i.i.d. and assumption (ref) hold with $\epsilon_{11}$ being a subgaussian random variable. Then, as $N\to\infty$, $TL_N^2-4\log N +\log\log N$ converges weakly to the distribution function $F(y)= \exp(-e^{-y/2}/\sqrt{8\pi})$, $y \in \mathbb{R}$.

The strategy of the proofs of Theorems (ref)-(ref) is to approximate $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ for any $p\geq 0$ by $L_N$ for the case $p=0$, in which the limiting behavior is understood in CJF13. We have to show the difference between the two versions of $L_N$ is negligible.

Theorems (ref)-(ref) indicate that we get the same asymptotic distribution of the max statistic $TL_N^2-4\log N +\log\log N$ under different moment assumptions. This allows us to have a flexibility to work on different pairs of $(N, T)$. Under null hypothesis (ref) and the assumptions imposed in the above three theorems, we conclude that a level-$\alpha$ test by rejecting the null hypothesis when $TL_N^2-4\log N+\log\log N$ is larger than the $1-\alpha$ quantile $q_{\alpha}= -\log(8\pi)-2\log\log(1-\alpha)^{-1}$ of $F(y)$.

The limiting distribution for the max-sum test

Review the accounts before the statement of Theorem (ref). We have the following conclusion on asymptotic independence.

theoremLet $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Under the same assumptions as in Theorem (ref), we have that $(S_N-\mu_N)/N$ and $TL_N^2-4\log N +\log\log N$ are asymptotically independent as $N\to \infty$.

By employing a new trick we prove the asymptotic independence between the maximum and the sum of random variables in Theorem (ref). This method is expected to be used in many of such type of problems. In fact, there are few literature to prove asymptotic independence between sums of and maxima of random variables. Some close references are BBQ98, Xu2016. The method here is new. It gives a novel tool to establish asymptotic independence between sums of and maxima of random variables.

To understand the idea quickly, we start with the set-up (ref). The first observation is that the maximum of many random variables, seemingly a global property, can be understood from their local property, that is, the maxima of subsets of random variables with fixed sizes. This step is done through the inclusion-exclusion formula. The second observation is that any such subset of random variables and the sum are independent with very high probability. Consequently the probability of the intersection of the events related to local maxima and the sum can be written as the product of two individual probabilities. Then we use the inclusion-exclusion formula one more time to get the product of two individual probabilities up to negligible errors. Details are shown at the beginning of Section (ref).

An immediate application is given below. By Theorems (ref) and (ref), we know that

align[align omitted — 182 chars of source]

Let $\Phi(x)$ be the distribution function of $N(0,1).$ Trivially, both $F(y)$ and $\Phi(x)$ are continuous functions. Set $P_{S_N}=1-\Phi\{(S_N-\mu_N)/N\}$ and $P_{L_N}=1-F(TL_N^2-4\log N+\log\log N)$. By Theorem (ref), (ref) and (ref), we see that $P_{L_N}$ and $P_{S_N}$ are asymptotically independent and each limit is $U[0, 1]$, the uniform distribution over $[0, 1].$ So the following holds easily.

coroSet $C_N=\min\{P_{S_N}, P_{L_N}\}$. Assume the setting in Theorem (ref). Then $C_N$ converges to $W:=\min\{U,V\}$ in distribution as $N\to\infty$, where $U$ and $V$ are i.i.d. random variables with distribution $U[0, 1].$ The distribution function of $W$ is given by $G(w)=2w-w^2$ for $w \in [0, 1]$.

According to Corollary (ref), the proposed max-sum test in (ref) allows us to perform a level-$\alpha$ test by rejecting the null hypothesis (ref) if $C_N<1-\sqrt{1-\alpha}$.

Simulation studies

We now conduct simulations to compare the finite sample performance of the tests studied in this paper and another test in literature. The tests we have worked on in this paper are based on $S_N$, $L_N$, $C_N$, $Q_N$ in (ref), (ref), (ref), (ref), respectively. The other one is based on $CD$ from Pesaran2004 defined by

align[align omitted — 97 chars of source]

Notice $Q_N$ here is the same as the notation LM$_{\textrm{adj}}$ from pesaran2008a. In the following we will explain our simulation designs and state our simulation findings.

Simulation designs

We consider the data generating process used in pesaran2008a, which is specified as

align*[align* omitted — 74 chars of source]

for $i=1, \cdots, N$ and $t=1, \cdots, T$. Comparing the notations in model (ref), we have $x_{it}=(1, x_{2it}, \cdots, x_{pit})^T\in \mathbb{R}^{p}$ and $\beta_i=(\alpha_i, \beta_{2i}, \cdots, \beta_{pi})\in \mathbb{R}^{p}.$

Now we independently generate $\alpha_i\sim N(0,1)$ and $\beta_{li}\sim N(1,0.04)$. The covariates are generated by

align*[align* omitted — 43 chars of source]

for $i=1,\cdots, N,~t=-50, -49, \cdots, T$ and $l=2,\cdots, p$ with $x_{li,-51}=0$, where $v_{lit}\sim N(0,\zeta_{li}^2/(1-0.6^2))$ and $\zeta_{li}^2\sim \chi^2_6/6$. In this case, $\zeta_{li}^2$'s are independently sampled first, then $v_{lit}$'s are independently generated by conditioning on the values of $\zeta_{li}^2$.

Now we generate $\epsilon_{it}$'s under null hypothesis (ref). Let $\epsilon_{it}=\sigma_i w_{it}$, where $w_{it}$'s are generated from three different distributions: (i) $N(0,1)$; (ii) $t_6/\sqrt{6/4}$; (iii) $(\chi^2_5-5)/\sqrt{10}$. Here $t_d$ is the $t$-distribution of degree $d$ and $\chi^2_d$ is the chi-square distribution of degree $d.$ The normalization in (ii) and (iii) is such that each new random variable has mean zero and variance one. Let $\sigma_i^2\sim \frac{p}{2}\chi^2_2$, as in the dynamic setup of pesaran2008a.

We turn to produce data under the alternative hypothesis. Let $\bold{\eta}_t:=(\eta_{1t},\cdots,\eta_{Nt})'$ be generated from the above three different distributions under the null hypothesis. Set $\bold{\epsilon}_{.t}=(\epsilon_{1t},\cdots,\epsilon_{Nt})'={\bm\Sigma}^{1/2}\bold{\eta}_t$. Please differentiate the notation $\bold{\epsilon}_{.t}$ here and $\bold{\epsilon}_{i}$ in (ref). We consider the following two cases of the covariance matrix ${\bm\Sigma}={\bf D}^{1/2}{\bf R}{\bf D}^{1/2}$ with ${\bf D}=\mathrm {diag}\{\sigma^2_1,\cdots,\sigma^2_N\}$.

itemize• Non-sparse case. Randomly select a subset $A\subset\{1,\cdots,N\}$ with cardinality $N^{0.5}$. Let ${\bf R}=(\rho_{ij})_{N\times N}$ be a symmetric matrix with $\rho_{ij}=1$ if $i=j$. For $i< j$, define $\rho_{ij}=0$ if $i\not \in A$ or $j\not \in A$, and $\rho_{ij}$ has the uniform distribution over $(\sqrt{3\log N/T}, \sqrt{5\log N/T})$ if $i\in A$ and $j \in A$. • Sparse case. Randomly select a subset $A\subset\{1,\cdots,N\}$ with cardinality $N^{0.3}$. Let ${\bf R}=(\rho_{ij})_{N\times N}$ be a symmetric matrix with $\rho_{ij}=1$ if $i=j$. For $i< j$, define $\rho_{ij}=0$ if $i\not \in A$ or $j\not \in A$, and $\rho_{ij}$ has the uniform distribution over $(\sqrt{8\log N/T}, \sqrt{10\log N/T})$ if $i\in A$ and $j \in A$.

To ensure that the covariance matrix ${\bm\Sigma}={\bf D}^{1/2}{\bf R}{\bf D}^{1/2}$ is positive definite, we replace the correlation matrix ${\bf R}$ with ${\bf R}+\lambda {\bf I}_N$, where $\lambda:=|\lambda_{\min}({\bf R})|+0.05$ and $\lambda_{\min}({\bf R})$ is the minimum eigenvalue of ${\bf R}$. Then, we consider two choices of the sample size $T=50,100$, and three choices of the dimension $N=50,100,200$.

Simulation results

We now present simulation results on the tests of $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. All the conclusions are based on 1,000 replications. The empirical sizes and powers of these tests in non-sparse and sparse cases are summarized in Tables (ref) to (ref). The power curves are plotted in Figure (ref). We next analyze them in detail.

Table (ref) indicates that all methods have empirical sizes not much larger than 5%. Here, the max test $L_N$ and the max-sum test $C_N$ tend to have smaller empirical sizes than the remaining ones, especially as $T$ is relatively small. This is not very surprising because it is common for maximum methods designed for raw data models; see, for example, Liu2008.

Tables (ref) and (ref) show the information of empirical powers in both non-sparse and sparse cases. Review the sum test $S_N$ and the sum-based test $Q_N$ are originally proposed in breusch1980the and pesaran2008a, respectively. The two are well studied in this paper. Tables (ref) and (ref) show that $S_N$ and $Q_N$ perform best in non-sparse cases in terms of empirical powers, but very poorly in sparse cases. On the contrary, the proposed max test $L_N$ performs the best in sparse cases, but very poorly in dense cases. Interestingly, it can be seen from Figure (ref) that the empirical power performance of our proposed max-sum test $C_N$ is always very close to the optimal one among all of the tests, regardless of the local alternative being sparse or not. This shows a very appealing property for the test $C_N$ which compromises the tests for residuals in both sparse and non-sparse cases. In fact, it is hard to tell in reality if residuals are sparse or not.

table[table omitted — 2,020 chars of source]
table[table omitted — 2,308 chars of source]
table[table omitted — 2,285 chars of source]

Figure (ref) shows the changes of the powers of all the tests as the degree of sparsity changes. Now we explain the procedure to generate the empirical power curves in Figure (ref). In fact, the horizonal direction in the plot is $n$, the degree of sparsity to be defined; the vertical direction represents powers. Specifically, the simulation is designed as follows. Review the general simulation design in Section (ref). We take $T=50$, $N=200$, $p=2$, $n=2,\cdots,16$; $w_{it}$ are generated from normal distributions; a subset $A\subset\{1,\cdots,N\}$ is randomly selected with cardinality $n$; ${\bf R}=(\rho_{ij})_{1\le i,j\le N}$, where $\rho_{ij}=1$ if $i=j$; for $i\neq j$, set $\rho_{ij}=0$ if $i\not \in A$ or $j\not \in A$, and $\rho_{ij}$ has the uniform distribution over $(\sqrt{8(\log n)^{-1}\log N/T}, \sqrt{10(\log n)^{-1}\log N/T})$ if $i\in A$ and $j \in A$. So a larger $n$ means a lower level of sparsity.

Figure (ref) indicates that the empirical power of the max-sum test $C_N$ is always very close to the maximum power of all tests for all $n$. By contrast, the empirical power curves of the remaining methods are monotone, {\it i.e.}, the empirical powers of both $S_N$ and $Q_N$ generally increase with the decrease of sparsity. On the contrary, the empirical power of the max test increases with the increase of sparsity. However, every test excluding the max-sum test $C_N$, favors either the sparse case or the non-sparse case, not both cases simultaneously.

figure[figure omitted — 256 chars of source]

Application

In this section, we apply the five tests to the securities in the Standard $\&$ Poor (S$\&$P) 500 index of large cap U.S. equity market. As seen earlier, they are $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. This demonstrates the practical usefulness of the proposed tests. The S$\&$P 500 index is primarily intended as a leading indicator of U.S. equities. The composition of this index is monitored by Standard and Poor to ensure the widest possible overall market representation while reducing the index turnover to a minimum. In this section, we consider 374 securities that have been included in the S$\&$P 500 index during the whole period from January 2005 to November 2018.

In particular, the panel data on the safe rate of return, and the market factors are obtained from Ken French's data library web page. The one-month US treasury bill rate is chosen as the risk-free rate ($r_{ft}$), the value-weighted return on all NYSE, AMEX, and NASDAQ stocks from CRSP is used as a proxy for the market return ($r_{mt}$), the average return on the three small portfolios minus the average return on the three big portfolios ($SMB_t$), and the average return on two value portfolios minus the average return on two growth portfolios ($HML_t$). SMB and HML are based on the stocks listed on the NYSE, AMEX and NASDAQ. All data are measured in percent per month. During January 2005 to November 2018, a total of 163 consecutive observations are obtained.

The Fama-French three-factor model Fama1993Common is given as follows: \[ y_{it}=r_{it}-r_{ft}=\beta_{0i}+\beta_{1i} (r_{mt}-r_{ft})+\beta_{2i}SMB_t+\beta_{3i} HML_t+\epsilon_{it} \] for each $1\leq i\leq N$ and $1\leq t\leq T$ with $N=374$. We are interested in the following null hypothesis:

align*[align* omitted — 78 chars of source]

That is, we are testing that the 374 variables are independent.

Now we evaluate the performance of the five tests in Section (ref), that is, $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. We randomly sample $T=15,25,35$ observations from the 163 monthly returns. At each value of $T$, the experiment is repeated 1000 times. It is trivial to see

eqnarray*[eqnarray* omitted — 125 chars of source]

Similarly,

eqnarray*[eqnarray* omitted — 98 chars of source]

This says that, although there is a dependency when sample $15$ numbers from a total of $163$ numbers for $1000$ times, comparing to $10^{15}$, the number of repeats $1000$ is still reasonable. The same also applies to the cases $T=25$ and $T=35$.

The results are summarized in Table (ref). It suggests that all tests except the max test always reject the null hypothesis of cross-sectional independence. So this indicates the definite cross-sectional dependence among stock returns under the three-factor model by Fama-French. In particular, the max test rejects the null hypothesis when $T$ grows to 35, but never reject it when $T$ reduces to 15. To understand this phenomenon, we point out a well known fact that there may exist a large number of underlying dependencies between stocks in the same industry or relevant industries. This leads us to believe that this is indeed a non-sparse case in which the sum and max-sum tests are more valid.

Concluding remarks

In this paper we study three tests: the sum test, the max test and the max-sum test, where the latter two are new ones. Two conjectures on the sum test have been settled. A new method to show asymptotic independence between the maximum and the sum of squares of a given set of random variables is established. Now we make some comments.

1. Under the Gaussian assumption, we obtain the CLTs for $S_N$ in Theorems (ref) and (ref). However, the Gaussian assumption is not needed in the study on the maxima of sample correlations in Theorems (ref), (ref) and (ref). One question is whether the Gaussian assumption can be removed from Theorems (ref) and (ref). Our proofs rely on the framework in Lemma (ref) where the normal assumption is essential. Another question is about the restriction between $N$ and $T$ in Theorems (ref) and (ref). Can the assumption “$T/\sqrt{N}\to \infty$" be relaxed? What is the behavior of $S_N$ for other regimes of relationship between $N$ and $T$?

2. The linear regression in (ref) is one of many panel data models; see, for example, the book length treatment in Baltagi2013, Hsiao2014, Pesaran2015, Wooldridge2010, among others. Some of other models can be studied similarly for the properties we have pursued in this paper. We leave them as a future work to our authors.

3. A new way is established to show the asymptotic independence between the sum of and the maximum of a set of random variables. The detail of the method is elaborated at the beginning of Section (ref). We expect this method will also work for other set of random variables of similar feature.

4. For the sum $S_N=\sum_{1\le i<j\le N}T\hat{\rho}_{ij}^2$, Theorem (ref) states that the central limit theorem of $S_N$ involves with projection matrices $\bold{P}_i$ defined via data; see (ref). However, interestingly enough, as shown in Theorems (ref), (ref) and (ref), the behavior of the maximum $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ does not depend on $\bold{P}_i$. Only parameters $N$ and $T$ participate in the limiting process.

5. As seen in Section (ref), a simulation study is carried out for tests based on $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. It shows that the max-sum test is always very close to the maximum power of all tests for both sparse and non-sparse residuals. By contrast, the empirical power curves of other methods favor only for one of the two types of residuals. In fact, in practice, it is hard to differentiate if a set of numbers is sparse or not. This implies the max-sum test is also desirable for other statistical models as long as three things are known: the central limit theorem holds for the sum of a set of random variables; the maximum of the set of random variables is the Gumbel distribution asymptotically; the sum and the maximum are asymptotically independent.

table[table omitted — 440 chars of source]

Proof

There are three subsections in this part. In each subsection we first accumulate some first hand or second hand of understanding before the proofs of main theorems are presented. Considering many proofs are involved, we postpone some of them in Appendix. They are interesting in their own right.

In this paper we use the following notation. For a sequence of random variables $\{U_N;\, N \geq 1\}$ and a sequence of constants $\{a_N;\, N\geq 1\}$, the notation $U_N=o_p(a_N)$ means that $U_N/a_N \to 0$ in probability as $N\to\infty;$ we write $U_N=O_p(a_N)$ if $\{U_N/a_N;\, N\geq 1\}$ is stochastically bounded, that is, $\lim_{A\to \infty}\limsup_{N\to \infty}P(|U_N/a_N|\geq A)=0.$ In particular, if $U_N=O_p(a_N)$ then $U_N=o_p(a_Nb_N)$ for any sequence of numbers $\{b_N;\, N\geq 1\}$ with $\lim_{N\to\infty}b_N=\infty.$ We write $a_N\sim b_N$ if $\lim_{N\to\infty}\frac{a_N}{b_N}=1$ for any two sequence of numbers $\{a_N;\, N\geq 1\}$ and $\{b_N;\, N\geq 1\}$.

The proofs of Theorems (ref) and (ref)

The proof of Theorem (ref) is lengthy. The main tool is the Lindeberg-Feller central limit theorem for martingales. Automatically many computations of conditional means and variances as well as higher moments are needed for sample correlation coefficients $\hat{\rho}_{ij}$. They are non-trivial. To make the proof organized, we decide to put key steps in a few of sections. This may best facilitate the understanding of readers.

Prelude 1: technical lemmas towards proofs of Theorems (ref) and (ref)

The proofs of the results in this section will be presented in Section (ref).

lemmaLet $\xi$ be a random variable with $E\xi=a$. Let $\tau\geq 2$ be given. The following holds. (i) If $a=0$, then \begin{eqnarray*} E[ |\xi^2-E\xi^2|^{\tau} ] \leq 32^{\tau}\cdot E(|\xi|^{2\tau}). \end{eqnarray*} (ii) If $a\ne 0$, then \begin{eqnarray*} && E[ |\xi^2-E\xi^2|^{\tau} ]\\ &\leq & 16^{\tau}\cdot \Big[|a|^{-\tau}\cdot {\rm Var}(\xi)^{\tau}+\sqrt{E(|\xi-a|^{2\tau})}\Big]\\ &&\cdot \Big[|a|^{\tau}+|a|^{-\tau}\cdot {\rm Var}(\xi)^{\tau}+\sqrt{E(|\xi-a|^{2\tau})}\Big]. \end{eqnarray*}

The following is the Marcinkiewicz-Zygmund inequality; see, e.g., p. 386 and p. 387 from CT88.

lemmaLet $m\geq 1$ and $\{\xi_i;\, 1\leq i \leq m\}$ be independent random variables with $E\xi_i=0$ for each $i$ and $\sup_{1\leq i \leq m}E(|\xi_i|^\tau)<\infty$ for some $\tau\geq 2$. Then there exists a constant $K_{\tau}>0$ depending on $\tau$ only such that \begin{align} E(|\xi_1+\cdots +\xi_m|^{\tau})\leq & K_{\tau}\cdot E\big[\big(\xi_1^2+\cdots +\xi_m^2\big)^{\tau/2}\big] \\ \leq & K_{\tau}\cdot m^{(\tau/2)-1}\big(E|\xi_1|^{\tau}+\cdots +E|\xi_m|^{\tau}\big). \end{align}

Prelude 2: mixing moments on random variables uniformly distributed on spheres

In this subsection we develop some identities and inequalities regarding moments of random vectors with the uniform distribution on high-dimensional unit spheres. The tools and methods are of independent interest. The proof of Lemma (ref) is given in this section to show the main idea and starting point. The remaining proofs of other lemmas will be presented in Section (ref).

Review the setting above (ref) and notation $\bold{P}_i$ and $\bold{\epsilon}_i=(\epsilon_{i1},\cdots,\epsilon_{iT})^{'}\in \mathbb{R}^T$ for each $i$. The notation $\mathbb{S}^{m-1}$ represents the unit sphere in the $m$-dimensional Euclidean space.

lemmaSet $m=T-p.$ Let $\bold{O}_i$ be a $T\times T$ orthogonal matrix such that \begin{align} \bold{P}_i=\bold{O}_i \begin{pmatrix} \bold{I}_m & \bold{0}\\ \bold{0} & \bold{0} \end{pmatrix} \bold{O}_i',\ \ \ 1\leq i \leq N. \end{align} Write $\bold{O}_i=(\bold{U}_{i}, \bold{V}_{i})$ for each $i$, where $\bold{U}_{i}$ is a $T\times m$ submatrix. Let $\{\epsilon_{ij};\, 1\leq i\leq N, 1\leq j \leq T\}$ be independent random variables with $\epsilon_{ij}\sim N(0, \sigma_i^2)$, $\sigma_i>0$, for all $i$ and $j$. Write $\bold{\epsilon}_i=(\epsilon_{i1},\cdots,\epsilon_{iT})^{'}\in \mathbb{R}^T$ for each $i$. Let $\bold{s}_1, \cdots, \bold{s}_N$ be i.i.d. random vectors uniformly distributed on $\mathbb{S}^{m-1}.$ Then $(\frac{\bold{P}_1\bold{\epsilon}_1}{\|\bold{P}_1\bold{\epsilon}_1\|}, \cdots, \frac{\bold{P}_N\bold{\epsilon}_N}{\|\bold{P}_N\bold{\epsilon}_N\|})$ and $(\bold{U}_1\bold{s}_1, \cdots, \bold{U}_{N}\bold{s}_N)$ have the same distribution.

\noindentProof of Lemma (ref). By the scale-invariance of $\frac{\bold{P}_i\bold{\epsilon}_i}{\|\bold{P}_i\bold{\epsilon}_i\|}$, without loss of generality, assume $\sigma_1=\cdots=\sigma_N=1$. Evidently, $(\bold{U}_i, \bold{0})\bold{\epsilon}_i=\bold{U}_i\bold{\eta}_i$ for each $i$, where $\bold{\eta}_i=(\epsilon_{i1},\cdots,\epsilon_{im})'$. By the orthogonality and (ref),

align[align omitted — 114 chars of source]

By the orthogonal invariance of normal distributions and (ref) again, $\bold{P}_i\bold{\epsilon}_i=(\bold{U}_i, \bold{0})\bold{O}_i'\bold{\epsilon}_i$ has the same distribution as that of $(\bold{U}_i, \bold{0})\bold{\epsilon}_i=\bold{U}_i\bold{\eta}_i$. Then $\frac{\bold{P}_i\bold{\epsilon}_i}{\|\bold{P}_i\bold{\epsilon}_i\|}$, as a function of $\bold{P}_i\bold{\epsilon}_i$, has the same distribution as that of

eqnarray*[eqnarray* omitted — 216 chars of source]

for each $i$ by the first identity of (ref). The desired conclusion then follows from the independence among $\{\bold{\epsilon}_1, \cdots, \bold{\epsilon}_N\}$. $\Box$

lemmaLet $\bold{U}_i$'s be as in Lemma (ref). The following holds. (i) Set $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ for any $1\leq i<j \leq N.$ Then both $\bold{M}_{ij}$ and $\bold{I}_m-\bold{M}_{ij}$ are non-negative definite. (ii) Let $\bold{M}$ be a $m\times m$ non-negative definite matrix satisfying that $\bold{I}_m-\bold{M}$ is non-negative definite. Then $\bold{I}_m-\bold{M}^2$ is also non-negative definite.
lemmaLet $\{\bold{M}_i, i=1,2\}$ be non-negative definite matrices. Assume $\bold{M}_1$ is idempotent, that is, $\bold{M}_1^2=\bold{M}_1$. Then $\mbox{tr}\,(\bold{M}_1\bold{M}_2) \leq \mbox{tr}\,(\bold{M}_2)$.
lemmaLet $\bold{M}_1$ and $\bold{M}_2$ be $n\times n$ non-negative definite matrices. Then, $\mbox{tr}(\bold{M}_1\bold{M}_2) \geq 0$ and $[\mbox{tr}(\bold{M}_1\bold{M}_2)]^2 \leq r\cdot \mbox{tr}((\bold{M}_1\bold{M}_2)^2)$, where $r:=\textrm{rank}(\bold{M}_1\bold{M}_2)$ $\leq$ $n$.

Recall notation $(2m-1)!!=1\cdot3\cdots (2m-1)$ for any integer $m\geq 1$. By convention we set $(-1)!!=1$.

lemma[Lemma 2.4 from Jiang09]. Suppose $m\geq 2$ and $Z_1, \cdots, Z_m$ are i.i.d. $N(0, 1)$-distributed random variables. Define $U_i=Z_i^2/(Z_1^2+\cdots +Z_m^2)$ for $1\leq i \leq m$. Let $a_1, \cdots, a_m$ be nonnegative integers. Set $a=a_1+\cdots +a_m$. Then \begin{eqnarray*} E(U_1^{a_1} U_2^{a_2}\cdots U_m^{a_m})=\frac{\prod_{i=1}^m(2a_i-1)!!}{\prod_{i=1}^a(m+2i-2)}. \end{eqnarray*}
lemmaLet $m\geq 2$ and $\{Z_i;\, 1\leq i \leq m\}$ be i.i.d. $N(0,1)$-distributed random variables. Set $\bold{d}=(Z_1, \cdots, Z_m)'/(Z_1^2+ \cdots +Z_m^2)^{1/2}$. Let $\bold{M}$ be a symmetric matrix. Then \begin{eqnarray*} &(i)& E(\bold{d}'\bold{M}\bold{d})=\frac{1}{m}\cdot tr(\bold{M});\\ &(ii)& E[(\bold{d}'\bold{M}\bold{d})^2]=\frac{1}{m(m+2)}\cdot \big\{2\,tr(\bold{M}^2)+[tr(\bold{M})]^2\big\};\\ &(iii)& Var(\bold{d}'\bold{M}\bold{d})=\frac{2}{m(m+2)}\cdot tr(\bold{M}^2)-\frac{2}{m^2(m+2)}\cdot \big[tr(\bold{M})\big]^2. \end{eqnarray*}
lemmaLet $\{Z_i;\, 1\leq i \leq m\}$ be i.i.d. $N(0,1)$-distributed random variables. Set $\bold{d}=(Z_1, \cdots, Z_m)'/(Z_1^2+ \cdots +Z_m^2)^{1/2}$ for $1\leq i \leq m$. Let $\bold{M}$ be a $m\times m$ symmetric matrix. Let $\tau\geq 1$ be given. Then, \begin{eqnarray*} E\big[|\bold{d}'\bold{M}\bold{d}-E(\bold{d}'\bold{M}\bold{d})|^{\tau}\big]\leq \frac{C_\tau}{m^{\tau}}\cdot \Big\{tr(\bold{M}^2)-\frac{1}{m}[tr(\bold{M})]^2\Big\}^{\tau/2} \end{eqnarray*} for all $m\geq 4\tau+1$, where $C_{\tau}>0$ is a constant depending on $\tau$ only.
lemmaLet $\{Z_i;\, 1\leq i \leq m\}$ be i.i.d. $N(0,1)$-distributed random variables. Set $\bold{d}=(Z_1, \cdots, Z_m)'/(Z_1^2+ \cdots +Z_m^2)^{1/2}$ for $1\leq i \leq m$. Let $\bold{a}\in \mathbb{R}^m$ be a vector and $\bold{M}$ be a $m\times m$ symmetric matrix. Let $\tau\geq 1$ be given. Then, $E(|\bold{a}'\bold{d}|^{2\tau})\leq C_\tau\|\bold{a}\|^{2\tau}/m^{\tau}$ and \begin{eqnarray*} E\big(\bold{d}'\bold{M}\bold{d}\big)^{\tau}\leq \frac{C_\tau}{m^{\tau}}\cdot \Big\{|tr(\bold{M})|^{\tau} + \Big[tr(\bold{M}^2)-\frac{1}{m}[tr(\bold{M})]^2\Big]^{\tau/2}\Big\} \end{eqnarray*} for all $m\geq 2\tau+1$, where $C_{\tau}>0$ is a constant depending on $\tau$ only.
lemmaLet $\{\bold{h}, \bold{h}_1, \bold{h}_2\}$ be i.i.d. $\mathbb{R}^m$-valued random vectors, where $\bold{h}$ has the same distribution as $\bold{d}$ in Lemma (ref). Let $\bold{A}$, $\bold{B}$ and $\bold{C}$ be $m\times m$ matrices. Then (i) $E\big[(\bold{h}'\bold{A}\bold{h})(\bold{h}'\bold{B}\bold{h})\big] =\frac{1}{m(m+2)}\big[2\,\mbox{tr}(\bold{A}\bold{B})+\mbox{tr}(\bold{A})\cdot\mbox{tr}(\bold{B}) \big]$ if $\bold{A}$ and $\bold{B}$ are symmetric. (ii) $\mbox{Var}[(\bold{h}_1'\bold{C}\bold{h}_2)^2]\leq \frac{K}{m^{5/2}}\cdot {[\mbox{tr}((\bold{C}\bold{C}')^4)]}^{1/2}$, where $K>0$ is a constant. (iii) $\mbox{Cov}\big[(\bold{h}'\bold{A}\bold{h}_1)^2, (\bold{h}'\bold{B}\bold{h}_2)^2\big]=\frac{2}{m^3(m+2)}\cdot \mbox{tr}(\bold{A}\bold{A}'\bold{B}\bold{B}')-\frac{2}{m^4(m+2)}\, \mbox{tr}(\bold{A}\bold{A}')\cdot\mbox{tr}(\bold{B}\bold{B}')$.

A quick reminder is that, although we assume that $\bold{A}$ and $\bold{B}$ are symmetric in (i) above, we do no need that $\bold{A}$, $\bold{B}$ or $\bold{C}$ are symmetric in (ii) and (iii).

lemmaReview $\bold{P}_i$ in (ref) and $\hat{\rho}_{ij}$ in (ref). Recall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). The following statements hold for all $i\ne j.$ (i) $E\hat{\rho}_{ij}=0$ and $E(\hat{\rho}_{ij}^2)=\frac{1}{m^2}\cdot\mbox{tr}(\bold{P}_i\bold{P}_j)$. (ii) $E[\hat{\rho}_{ij}|\bold{s}_i]=0$ and $E[\hat{\rho}_{ij}^2|\bold{s}_i]=\frac{1}{m}\cdot \bold{s}_i'\bold{M}_{ij}\bold{s}_i$.

In the following we will use notation $\mbox{Var}(\xi_2|\xi_1)$ for conditional variance, which is defined by $E(\xi_2^2|\xi_1)-[E(\xi_2|\xi_1)]^2$ for any random variables $\xi_1$ and $\xi_2.$

lemmaReview $\bold{P}_i$ in (ref) and $\hat{\rho}_{ij}$ in (ref). Recall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). The following statements are true for all $i\ne j.$ (i) $E\big[(\hat{\rho}_{ij})^4\big|\bold{s}_i\big] = \frac{3}{m(m+2)}\cdot\big(\bold{s}_i'\bold{M}_{ij}\bold{s}_i\big)^2.$ (ii) $E[(\hat{\rho}_{ij})^4]=\frac{3}{m^2(m+2)^2}\cdot \big\{2\,\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]+ [\mbox{tr}(\bold{P}_i\bold{P}_j)]^2\big\}$. (iii) $\mbox{Var}(\hat{\rho}_{ij}^2|\bold{s}_i)=\frac{2(m-1)}{m^2(m+2)}\cdot \big(\bold{s}_i'\bold{M}_{ij}\bold{s}_i\big)^2$. (iv) $\mbox{Var}(\hat{\rho}_{ij}^2) = \frac{6}{m^2(m+2)^2}\cdot \mbox{tr}[(\bold{P}_i\bold{P}_j)^2]+\frac{2(m^2-2m-2)}{m^4(m+2)^2}\cdot [\mbox{tr}(\bold{P}_i\bold{P}_j)]^2.$

Intermezzo 1: calculations of variances of sums related to sample correlation coefficients

In (ref) and (ref), we see parameters $p, N, T$ and variables $\bold{x}_i$. In the rest of the paper, we will use or develop many inequalities where a constant $C$ will appear frequently. The constant $C$ does not depend on $p, N, T$ or $\bold{x}_i$'s and it can be different from line to line. The proofs of the lemma in this section will be given in Section (ref).

lemmaReview the notations $p, T, N$ and $\bold{P}_i$ in (ref). Let $p$ be fixed and $m=T-p\geq 1$. For any set $S\subset \{1,2,\cdots,N\}$ with $q=|S|\in \{1, \cdots, N-1\}$, define $\bold{P}_S=\sum_{k\in S}\bold{P}_k.$ Let $j\notin S.$ Then there exists a constant $K>0$ depending on $p$ but not on $N$, $T$ or $\bold{P}_i$ such that the following statements hold uniformly for all $1\leq i< j \leq N$ and $N\geq 4.$ \begin{flalign} &\ \ \ \ \ (i)\ \frac{1}{T}\cdot \big|[tr(\bold{P}_i\bold{P}_j)]^2-T^2\big|\leq K.& \nonumber\\ &\ \ \ \ \ (ii)\ \big|tr((\bold{P}_i\bold{P}_j)^2)-T\big|\leq K.&\nonumber\\ &\ \ \ \ \ (iii) \ \frac{1}{Tq^2}\cdot \big|[tr(\bold{P}_S\bold{P}_j)]^2-T^2q^2\big|\leq K.&\nonumber\\ &\ \ \ \ \ (iv)\ \frac{1}{q^2}\cdot \big|tr((\bold{P}_S\bold{P}_j)^2)-Tq^2\big|\leq K.&\nonumber\\ &\ \ \ \ \ (v)\ Statements (i)-(iv) still hold if symbol\ “T"\, is replaced by \, “m".&\nonumber \end{flalign}
lemmaRecall $\bold{U}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). Let $\bold{e}$ have the uniform distribution on $\mathbb{S}^{m-1}$. Then there is a constant $C>0$ free of $N, T$ and $p$ such that $\sup_{1\leq i<j \leq N}\mbox{Var}\big((\bold{e}'\bold{M}_{ij}\bold{e})^2\big)\leq Cm^{-2}$ as $N\geq C$.
lemmaReview $\bold{P}_i$ in (ref) and $\hat{\rho}_{ij}$ in (ref). Recall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). Define $\bold{P}_{j\blacktriangle}=\sum_{i=1}^{j-1}\bold{P}_i$ for $2\leq j\leq N.$ By Lemma (ref), \begin{eqnarray*} X_j:=\sum_{i=1}^{j-1}[T\hat{\rho}_{ij}^2-E(T\hat{\rho}_{ij}^2|\bold{s}_i)]=\sum_{i=1}^{j-1}T\hat{\rho}_{ij}^2-\frac{T}{m}\sum_{i=1}^{j-1} \bold{s}_i'\bold{M}_{ij}\bold{s}_i \end{eqnarray*} for $2\leq j\leq N.$ Then, \begin{eqnarray*} \frac{1}{T^2}E(X_j^2)&=&\frac{2m-8}{m^3(m+2)^2}\sum_{i=1}^{j-1}tr((\bold{P}_i\bold{P}_j)^2)+ \frac{2m^2+4}{m^4(m+2)^2}\sum_{i=1}^{j-1}[tr(\bold{P}_i\bold{P}_j)]^2+\\ &&\frac{2}{m^3(m+2)} \Big\{tr\big((\bold{P}_{j\blacktriangle}\bold{P}_{j})^2\big)- \frac{1}{m}\cdot\big[tr(\bold{P}_{j\blacktriangle}\bold{P}_{j})\big]^2\Big\}. \end{eqnarray*}

A quick comment is that the last term above is non-negative by Lemma (ref).

lemmaLet $X_j$ be defined as in Lemma (ref) for $2\leq j\leq N.$ Assume $p$ is fixed and $N=o(T^2)$ as $N\to\infty$. Then \begin{eqnarray*} \lim_{N\to \infty}\frac{1}{N^2}\sum_{j=2}^NE(X_j^2) = 1. \end{eqnarray*}
lemmaRecall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). Assume $p$ is fixed and $T=T_N\to \infty$. Then \begin{eqnarray*} {\rm Var}\Big[\sum_{j=2}^N\Big(\sum_{i=1}^{j-1}\bold{s}_i'\bold{M}_{ij}\bold{s}_i\Big)^2\Big] =O\Big(\frac{N^5}{T^2}\Big). \end{eqnarray*} In particular, the variance above is of order $o(N^4T^2)$ if $N=o(T^4)$ as $N\to\infty$.
lemmaRecall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref). If $p$ is fixed and $T=T_N\to \infty$, then \begin{eqnarray*} {\rm Var}\Big\{\sum_{j=2}^N{\rm tr}\Big[\Big(\sum_{i=1}^{j-1}\bold{U}_j'\bold{U}_i\bold{s}_i \bold{s}_i'\bold{U}_i'\bold{U}_j\Big)^2\Big]\Big\}=O\Big(\frac{N^4}{T^2}+\frac{N^5}{T^3}\Big) \end{eqnarray*} as $N\to\infty$. In particular, the variance is of the order $o(N^4)$ if $N=o(T^3)$.

Intermezzo 2: preliminary verifications of the Lindeberg-Feller condition towards proof of Theorem (ref)

lemmaRecall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref) and $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ from Lemma (ref). Let $\mu_N$ be as in (ref). Set $B_j=\frac{T}{m}\sum_{i=1}^{j-1} \bold{s}_i'\bold{M}_{ij}\bold{s}_i$ for $2\leq j \leq N$. If $p$ is fixed and $N=o(T^2)$ as $N\to \infty$, then \begin{eqnarray*} \frac{1}{N}\Big[ \Big(\sum_{j=2}^NB_j\Big)-\mu_N\Big] \to 0 \end{eqnarray*} in probability as $N\to\infty$.

\noindentProof of Lemma (ref). By (ref), $\bold{U}_{i}\bold{U}_{i}'=\bold{P}_i$. Then

align[align omitted — 143 chars of source]

It follows from Lemma (ref)(i) that

eqnarray*[eqnarray* omitted — 150 chars of source]

for $2\leq j \leq N.$ Therefore,

eqnarray*[eqnarray* omitted — 315 chars of source]

by the fact that $\mbox{tr}(\bold{P}_i)=T-p=m$ for each $i.$ On the other hand,

eqnarray*[eqnarray* omitted — 264 chars of source]

where $\bold{M}_{i\blacktriangledown}:=\sum_{j=i+1}^N\bold{M}_{ij}.$ By independence among $\{\bold{s}_i\}$'s and Lemma (ref),

align[align omitted — 551 chars of source]

From Lemma (ref), we know $\bold{I}_m-\bold{M}_{ij}$ is non-negative for each $i \ne j.$ Since the sum of non-negative definite matrices is still non-negative definite, we see that $(N-i)\bold{I}_m-\bold{M}_{i\blacktriangledown}$ is also non-negative definite. By Lemma (ref)(ii), $(N-i)^2\bold{I}_m-\bold{M}_{i\blacktriangledown}^2$ is non-negative definite. In particular,

align[align omitted — 88 chars of source]

Moreover,

align[align omitted — 155 chars of source]

by (ref). Now we estimate $\mbox{tr}(\bold{P}_i\bold{P}_j)$.

Recall (ref). Set $\bold{A}_i=\bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'$ for $1\leq i \leq N.$ Then $\bold{A}_i$ is a $T\times T$ idempotent matrix with rank $p$ and $\mbox{tr}(\bold{A}_i)=p$ for each $i$. Since $\bold{P}_i=\bold{I}_T-\bold{A}_i$, we see

align*[align* omitted — 59 chars of source]

where $\bold{B}_{ij}:=\bold{A}_i\bold{A}_j-\bold{A}_i-\bold{A}_j.$ By Lemma (ref),

align*[align* omitted — 51 chars of source]

for any non-negative definite matrices $\bold{F}_1$ and $\bold{F}_2$. As a result, $\mbox{tr}(\bold{A}_i\bold{A}_j) \geq 0$. Easily, $\mbox{tr}(\bold{A}_i\bold{A}_j)\leq p$ by Lemma (ref). Thus,

align*[align* omitted — 54 chars of source]

Therefore, we have $\mbox{tr}(\bold{P}_i\bold{P}_j) \geq T-2p$. Hence, $\mbox{tr}(\bold{M}_{i\blacktriangledown})\geq (N-i)(T-2p)$ by (ref). This and (ref) tell us that

eqnarray*[eqnarray* omitted — 246 chars of source]

by recalling the notation $m=T-p.$ Plugging this into (ref) we get

eqnarray*[eqnarray* omitted — 128 chars of source]

By the Chebyshev inequality, for any $\tau>0$,

eqnarray*[eqnarray* omitted — 219 chars of source]

which goes to zero provided $N=o(T^2).$ $\square$

Let $\{\bold{s}_1, \cdots, \bold{s}_j\}$ for $1\leq j \leq N$ be defined in Lemma (ref), which are i.i.d. random vectors uniformly distributed on $\mathbb{S}^{m-1}$. Set

align[align omitted — 139 chars of source]

which is the $\sigma$-algebra generated by $\{\bold{s}_1, \cdots, \bold{s}_j\}$ for $1\leq j \leq N$. Here $\Omega$ is the sample space on which random variables $\{\epsilon_{ij}\}$ are defined on.

lemmaLet $X_j$ be defined as in Lemma (ref) and $\mathcal{F}_j$ be as in (ref). Assume $N=o(T^3)$. Define \begin{eqnarray*} Z_N=\frac{1}{N^2} \sum_{j=2}^NE [X_j^2|\mathcal{F}_{j-1}]. \end{eqnarray*} Then ${\rm Var}(Z_N)\rightarrow 0$ as $N\to \infty.$

\noindentProof of Lemma (ref). Set $\bold{H}_{ij}=\bold{U}_i'\bold{U}_j\bold{s}_j \bold{s}_j'\bold{U}_j'\bold{U}_i$ for $1\leq i, j \leq N$, where $\bold{U}_i$'s and $\bold{s}_j$'s are defined as in Lemma (ref). Then $\bold{s}_i'\bold{H}_{ij}\bold{s}_i=\bold{s}_j'\bold{C}_{ij}\bold{s}_j$, where

eqnarray*[eqnarray* omitted — 96 chars of source]

Since $\bold{s}_i'\bold{U}_i'\bold{U}_j\bold{s}_j=(\bold{s}_i'\bold{U}_i'\bold{U}_j\bold{s}_j)' =\bold{s}_j'\bold{U}_j'\bold{U}_i\bold{s}_i\in \mathbb{R}$, we have

align[align omitted — 172 chars of source]

By Lemma (ref)(i) and the independence between $\bold{s}_i$ and $\bold{s}_j$, we have that

eqnarray*[eqnarray* omitted — 207 chars of source]

for $i<j$, where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. Then

align[align omitted — 307 chars of source]

for $2\leq j \leq N$, where

align[align omitted — 173 chars of source]

In view of the independence among $\bold{s}_i$'s, it is easy to check from Lemma (ref) that

eqnarray*[eqnarray* omitted — 175 chars of source]

Since $\mbox{tr}(\bold{C}_{ij})=\mbox{tr}(\bold{U}_j'\bold{U}_i\bold{s}_i \bold{s}_i'\bold{U}_i'\bold{U}_j)=\bold{s}_i'\bold{M}_{ij}\bold{s}_i$, we have that

eqnarray*[eqnarray* omitted — 97 chars of source]

This, (ref) and Lemma (ref) imply

align[align omitted — 273 chars of source]

From (ref),

eqnarray*[eqnarray* omitted — 347 chars of source]

by the definition of $\bold{C}_{ij}$. Thus, we conclude from (ref) that

eqnarray*[eqnarray* omitted — 288 chars of source]

It follows that

eqnarray*[eqnarray* omitted — 351 chars of source]

Review $T=m+p$. Since ${\rm Var}(\xi_1+\xi_2)\leq 2{\rm Var}(\xi_1)+2{\rm Var}(\xi_2)$ for any random variables $\xi_1$ and $\xi_2$, to show ${\rm Var}(Z_N)\rightarrow 0$, it is enough to prove the following two facts.

align[align omitted — 312 chars of source]

Under restriction $N=o(T^3)$, the assertion (ref) is confirmed in Lemma (ref) and (ref) is proved in Lemma (ref). The proof is completed. $\square$

lemmaLet $X_j$ be defined as in Lemma (ref) and $\mathcal{F}_j$ be as in (ref). Assume $N=o(T^4)$. Then \begin{eqnarray*} \frac{1}{N^4} \sum_{j=2}^N E(X_j^4|\mathcal{F}_{j-1})\to 0 \end{eqnarray*} in probability as $N\to \infty.$

\noindentProof of Lemma (ref). It suffices to show

align[align omitted — 70 chars of source]

as $N\to\infty.$ By (ref) and (ref),

align[align omitted — 147 chars of source]

for $2\leq j \leq N$, where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ and

eqnarray*[eqnarray* omitted — 161 chars of source]

Notice

eqnarray*[eqnarray* omitted — 211 chars of source]

where

eqnarray*[eqnarray* omitted — 95 chars of source]

for $1\leq i<j\leq N.$ By Lemma (ref),

align[align omitted — 317 chars of source]

for any $1\leq i<j\leq N.$ We rewrite (ref) to have

eqnarray*[eqnarray* omitted — 242 chars of source]

Therefore,

align[align omitted — 113 chars of source]

Note that $A_j$ is the sum of independent random variables. By (ref) with $\tau=4$,

align[align omitted — 418 chars of source]

where the second inequality follows from Lemma (ref), and where the fact that $\mbox{tr}(\bold{H}_{ij}^2)\geq \frac{1}{m}[\mbox{tr}(\bold{H}_{ij})]^2$ from Lemma (ref) is used in the last step. Easily, $\mbox{tr}(\bold{H}_{ij}^2)=(\bold{s}_j'\bold{M}_{ji}\bold{s}_j)^2.$ Take another expectation to see

align[align omitted — 131 chars of source]

By Lemma (ref) with $\tau=4$ and the fact that $\mbox{tr}(\bold{M}^2)\geq \frac{1}{m}[\mbox{tr}(\bold{M})]^2$ for any $m\times m$ symmetric matrix $\bold{M}$ from Lemma (ref), we obtain

eqnarray*[eqnarray* omitted — 363 chars of source]

It is used before that $\mbox{tr}(\bold{M}_{ji})=\mbox{tr}(\bold{P}_i\bold{P}_j)$ and $\mbox{tr}[\bold{M}_{ji}^2]=\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]$. By Lemma (ref), both quantities are bounded by $m.$ Hence, $E\big[(\bold{s}'\bold{M}_{ji}\bold{s})^4\big]\leq C$ uniformly for all $1\leq i<j\leq N$. We conclude from (ref) that

align[align omitted — 62 chars of source]

uniformly for all $2\leq j \leq N$.

Now we estimate $B_j$. Replace “$\bold{H}_{ij}$" in (ref) with “$\bold{M}_{ij}$" to see

align[align omitted — 393 chars of source]

where the last step holds by (i) and (ii) from Lemma (ref).

Finally, by (ref),

eqnarray*[eqnarray* omitted — 266 chars of source]

where

eqnarray*[eqnarray* omitted — 75 chars of source]

Since $\mbox{tr}(\bold{M}_{ji})=\mbox{tr}(\bold{P}_i\bold{P}_j)$, by defining

eqnarray*[eqnarray* omitted — 72 chars of source]

we have $\mbox{tr}(\bold{M}_{j\blacktriangle})=\mbox{tr}(\bold{P}_{j\blacktriangle}\bold{P}_j)$. Recall (ref), $\bold{U}_{i}\bold{U}_{i}'=\bold{P}_i$. Easily,

align*[align* omitted — 111 chars of source]

for any $1\leq i, j, k\leq N$. It follows that

align*[align* omitted — 275 chars of source]

On the other hand, recall $m=T-p.$ By Lemma (ref)(v), there exists a constant $K$ not depending on $T$ or $N$ such that

eqnarray*[eqnarray* omitted — 220 chars of source]

for every $2\leq j \leq N.$ It follows from the triangle inequality that

align*[align* omitted — 174 chars of source]

for $2\leq j \leq N.$ Consequently, by taking $\tau=4$ in Lemma (ref) we have that

eqnarray*[eqnarray* omitted — 385 chars of source]

uniformly for all $2\leq j \leq N$. Combining this with (ref), (ref) and (ref), we arrive at

eqnarray*[eqnarray* omitted — 143 chars of source]

uniformly for all $2\leq j \leq N$ as $N$ is large (reviewing $m=T-p$ and $T=T_N\to \infty).$ As a result,

eqnarray*[eqnarray* omitted — 94 chars of source]

as $N\to \infty$ as long as $N=o(T^4).$ We obtain (ref). $\Box$

lemmaLet $X_j$ be defined as in Lemma (ref). Assume $N=o(T^2)$ as $N\to \infty.$ Then $\frac{1}{N}\sum_{j=2}^NX_j\to N(0, 1)$ in distribution as $N\to \infty.$

\noindentThe Proof of Lemma (ref). Reviewing Lemma (ref), we know

eqnarray*[eqnarray* omitted — 188 chars of source]

for $2\leq j\leq N.$ Let $\mathcal{F}_j$ be as in (ref). Next we will verify that, for each $N\geq 2$, $\{X_j;\, 2\leq j \leq N\}$ forms a sequence of martingale differences with respect to the $\sigma$-algebras $\{\mathcal{F}_{j};\, 1\le j\le N-1\}.$ Define $J_1=0$ and

align*[align* omitted — 52 chars of source]

for $2\leq j \leq N-1$. By Lemma (ref), $\hat{\rho}_{ij}$ depends on $\bold{s}_i$ and $\bold{s}_j$ only. From independence of $\{\bold{s}_1, \cdots, \bold{s}_N\}$ and Lemma (ref),

eqnarray*[eqnarray* omitted — 153 chars of source]

for $2\leq j \leq N,$ where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. Therefore,

align[align omitted — 77 chars of source]

forms a martingale difference with respect to the $\sigma$-algebras $\{\mathcal{F}_{j}; 2\le j\le N\}.$

Now, in order to prove

eqnarray*[eqnarray* omitted — 54 chars of source]

in distribution as $N\to\infty$, we will employ the Lindeberg-Feller central limit theorem (see, for example, p. 476 from Billingsley95 or p. 344 from Durrett). To achieve so, it is enough to verify that

equation[equation omitted — 99 chars of source]

in probability and

align[align omitted — 94 chars of source]

in probability as $N\to\infty.$ Lemma (ref) has showed (ref). Now, to prove (ref), it suffices to show

align[align omitted — 53 chars of source]

and

align[align omitted — 61 chars of source]

as $N\to\infty.$ Lemma (ref) proves (ref) under the assumption $N=o(T^2)$. The assertion (ref) is confirmed in Lemma (ref) by assuming $N=o(T^3)$. Inspect all restrictions between $N$ and $T$ in the lemmas used earlier, the condition $N=o(T^2)$ meets all requirement. The proof is then completed. $\square$

Finale: proofs of Theorems (ref) and (ref)

With the preparations in Sections in (ref)-(ref), we now are ready to prove the central limit theorem stated in Theorem (ref). The main idea is to write the sum of squares of sample correlation coefficients as sums of martingale differences. Then the Lindeberg-Feller martingale CLT is applied.

\noindentProof of Theorem (ref). Review $J_1=0$ and

align*[align* omitted — 52 chars of source]

for $2\leq j \leq N-1$. Then $S_N=\sum_{j=2}^NJ_j.$ Review $\mathcal{F}_0$ and $\mathcal{F}_j$ in (ref). By Lemma (ref), the conditional expectation,

eqnarray*[eqnarray* omitted — 109 chars of source]

for $2\leq j \leq N,$ where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. As in (ref),

align*[align* omitted — 64 chars of source]

forms a martingale difference with respect to the $\sigma$-algebras $\{\mathcal{F}_{j}; 2\le j\le N\}.$ Therefore $\frac{1}{N}(S_N-\mu_N)$ can be further written by

eqnarray*[eqnarray* omitted — 133 chars of source]

From Lemma (ref),

eqnarray*[eqnarray* omitted — 76 chars of source]

in probability as $N\to\infty$. By Lemma (ref),

eqnarray*[eqnarray* omitted — 54 chars of source]

in distribution as $N\to\infty.$ The proof then follows from the Slutsky lemma. $\square$

\noindentProof of Theorem (ref). Set $m=T-p$. First,

align[align omitted — 157 chars of source]

It is easy to see

eqnarray*[eqnarray* omitted — 71 chars of source]

as $N\to\infty$. It follows that

eqnarray*[eqnarray* omitted — 252 chars of source]

By Lemma (ref)(i) and (ii), there exists a constant $K>0$ depending on $p$ but not on $N$, $T$ or $\bold{P}_i$'s such that

eqnarray*[eqnarray* omitted — 165 chars of source]

uniformly for all $1\leq i< j \leq N$ and $N\geq 4.$ Therefore, by the definition of $v_{Nij}$, we have

eqnarray*[eqnarray* omitted — 185 chars of source]

uniformly for all $1\leq i< j \leq N$ as $N\to\infty.$ Immediately,

align[align omitted — 88 chars of source]

uniformly for all $1\leq i< j \leq N$ as $N\to\infty.$ Now write $\frac{1}{v_{Nij}}=\frac{1}{\sqrt{2}}(1+\omega_{Nij})$. Then

align[align omitted — 84 chars of source]

as $N\geq 4.$ By Lemma (ref), $E\hat{\rho}_{ij}^2=m^{-1}\mu_{Nij}.$ It follows that

eqnarray*[eqnarray* omitted — 363 chars of source]

where $S_N$ is defined as in Theorem (ref). By the Slutsky lemma and Theorem (ref),

eqnarray*[eqnarray* omitted — 67 chars of source]

in distribution as $N\to \infty.$ Recall (ref). To prove $Q_N\to N(0, 1)$ in distribution, by the Slutsky lemma again, it is enough to show

align[align omitted — 141 chars of source]

in probability as $N\to \infty$. Since $\hat{\rho}_{ij}^2$ and $\hat{\rho}_{kl}^2$ are independent if $\{i, j\}\cap \{k, l\}=\emptyset$, then

eqnarray*[eqnarray* omitted — 152 chars of source]

where the last sum runs over all $(k, l)$ with $1\leq k<l\leq N$ and $\{i, j\}\cap \{k, l\}\ne \emptyset$. The total number of such $(k, l)$'s is no more than $2N+ 2N=4N.$ Since $|\mbox{Cov}(U, V)|\leq [\mbox{Var}(U)]^{1/2}\cdot [\mbox{Var}(V)]^{1/2}$ for any random variables $U$ and $V$, we have from (ref) that

align[align omitted — 277 chars of source]

By Lemma (ref), $\mbox{tr}(\bold{P}_i\bold{P}_j)\leq \mbox{tr}(\bold{P}_i)= m$ and $\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]\leq \mbox{tr}(\bold{P}_i)= m$. By Lemma (ref)(iv),

eqnarray*[eqnarray* omitted — 227 chars of source]

Thus, $\mbox{Var}(\hat{\rho}_{ij}^2)\leq \frac{8}{m^2}$ for all $1\leq i<j \leq N.$ Combing this with (ref), we get

eqnarray*[eqnarray* omitted — 67 chars of source]

by the assumption $N=o(T^2).$ This implies (ref). The proof is completed.$\square$

The proofs of Theorems (ref), (ref) and (ref)

Theorems (ref)-(ref) will be proved via approximating $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ for any $p\geq 0$ by $L_N$ for the case $p=0$. The latter one has the asymptotic result known in CJF13. The main job is reduced to show the difference between the two versions of $L_N$ is small enough.

Prelude: auxilary results towards proofs of Theorems (ref), (ref) and (ref)

The results stated in this section will be proved in Section (ref).

lemmaLet $m\geq 1$ and $\{\xi_i;\, 1\leq i \leq m\}$ be i.i.d. random variables with $E\xi_1=0$,\, $E\xi_1^2=1$ and $E(|\xi_1|^\tau)<\infty$ for some $\tau\geq 2$. Let $\{a_i;\, 1\leq i \leq m\}$ be constants such that $a_1^2+\cdots + a_m^2=1.$ Then, there exists a constant $K>0$ satisfying \begin{eqnarray*} P(|a_1\xi_1+\cdots+ a_m\xi_m|\geq x) \leq \frac{K}{x^\tau} \end{eqnarray*} for all $x\geq 3.$

It is easy to see that the bound in the lemma is tight by simply taking $a_1=1$ and $a_2=\cdots=a_m=0.$

lemmaLet $m\geq 1$ and $\{\xi_i;\, 1\leq i \leq m\}$ be i.i.d. random variables with $E\xi_1=0$,\, $E\xi_1^2=1$ and $Ee^{\omega|\xi_1|}<\infty$ for some $\omega>0$. Let $\{a_i;\, 1\leq i \leq m\}$ be constants satisfying $a_1^2+\cdots + a_m^2=1.$ Then, there exists $K>0$ such that \begin{eqnarray*} P(|a_1\xi_1+\cdots+ a_m\xi_m|\geq x) \leq K\cdot e^{-x/K} \end{eqnarray*} for all $x\geq 0.$

The above inequality is tight, which can be seen by taking $a_1=1$ and $a_i=0$ for $2\leq i \leq m.$

Recall the definition of subgaussian random variables defined before the statement of Theorem (ref).

lemmaLet $m\geq 1$ and $\{\xi_i;\, 1\leq i \leq m\}$ be i.i.d. subgaussian random variables. Let $\{a_i;\, 1\leq i \leq m\}$ be constants such that $a_1^2+\cdots + a_m^2=1.$ Then, there exists a positive constant $K$ not depending on $m$ or $\{a_i;\, 1\leq i \leq m\}$ such that \begin{eqnarray*} P(|a_1\xi_1+\cdots+ a_m\xi_m|\geq x) \leq 2\cdot e^{-K x^2} \end{eqnarray*} for all $x>0.$

The upper bound in the lemma is optimized, which can be seen evidently by choosing $a_1=1$ and $a_2=\cdots=a_m=0$.

Intermezzo: approximation of sample correlation coefficients by simple versions

Recall the setting in (ref), (ref) and (ref). Let $p$ be fixed. Let $\bold{e}_i=\frac{\bold{\epsilon}_i}{\|\bold{\epsilon}_i\|}$ and $\tilde{\rho}_{ij}=\bold{e}_i'\bold{e}_j$ for $1\leq i, j \leq N$. In this section, we always assume that $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables. The “continuous" requirement guarantees that $\hat{\rho}_{ij}$ in (ref) is well-defined. See the comment below (ref).

lemmaAssume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables. Let $\hat{\rho}_{ij}$ be defined as in (ref). Set $\bold{A}_i=\bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'$ for $1\leq i \leq N$. Then, \begin{eqnarray*} \max_{1\leq i <j \leq N}\big|\hat{\rho}_{ij} - \tilde{\rho}_{ij}\big| \leq 14\cdot \big(\max_{1\leq i \leq j \leq N}\bold{e}_j'\bold{A}_i\bold{e}_j\big). \end{eqnarray*}

\noindentProof of Lemma (ref). Notice $\bold{P}_i\bold{\epsilon}_i=\bold{\epsilon}_i-\bold{A}_i\bold{\epsilon}_i.$ It follows that

eqnarray*[eqnarray* omitted — 365 chars of source]

Take $i=j$ to see that

eqnarray*[eqnarray* omitted — 122 chars of source]

since $\bold{A}_i^2=\bold{A}_i$. In particular, $\|\bold{\epsilon}_i\|^2- \bold{\epsilon}_i'\bold{A}_i\bold{\epsilon}_i\geq 0$, hence

align[align omitted — 76 chars of source]

for each $i$. Combining the last two identities, we have

eqnarray*[eqnarray* omitted — 387 chars of source]

Dividing the numerator and denominator by $\|\bold{\epsilon}_i\|\cdot \|\bold{\epsilon}_j\|$, we have

align[align omitted — 299 chars of source]

Write

eqnarray*[eqnarray* omitted — 108 chars of source]

It is easy to see $|1-(1-x)^{-1/2}|\leq 2|x|$ if $|x|\leq \frac{1}{2}.$ Then $(1-x)^{-1/2}\leq 1+2|x|$ as $|x|\leq \frac{1}{2}.$ For brevity of notation, set $h_{ij}=\big(1-\bold{e}_i'\bold{A}_i\bold{e}_i\big)^{-1/2} \big(1-\bold{e}_j'\bold{A}_j\bold{e}_j\big)^{-1/2}.$ Then

eqnarray*[eqnarray* omitted — 335 chars of source]

provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$, and at the same time $h_{ij}\leq 2$ by definition. From (ref),

align[align omitted — 219 chars of source]

By the Cauchy-Schwartz inequality and the fact $\bold{A}_i^2=\bold{A}_i$,

eqnarray*[eqnarray* omitted — 119 chars of source]

Similarly, $|\bold{e}_i'\bold{A}_i\bold{e}_j|\leq \|\bold{A}_i\bold{e}_i\|\cdot\|\bold{A}_i\bold{e}_j\|$ and $|\bold{e}_i'\bold{A}_j\bold{e}_j| \leq \|\bold{A}_j\bold{e}_i\|\cdot\|\bold{A}_j\bold{e}_j\|$ since $\bold{A}_i^2=\bold{A}_i$. Consequently

eqnarray*[eqnarray* omitted — 446 chars of source]

by the fact $|\tilde{\rho}_{ij}| \leq 1$. Use the trivial fact that $2xy\leq x^2+y^2$ to see

eqnarray*[eqnarray* omitted — 467 chars of source]

provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$, where in the last inequality we use the fact that each term is bounded by $\max_{1\leq i \leq j \leq N}\bold{e}_j'\bold{A}_i\bold{e}_j$. We then have from (ref) that

align[align omitted — 179 chars of source]

provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$. If $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i> \frac{1}{2}$, (ref) holds automatically due to the facts that $|\hat{\rho}_{ij}|\leq 1$ and $|\tilde{\rho}_{ij}|\leq 1$ for all $i,j$. The proof is completed. $\Box$

propAssume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables with $E\epsilon_{11}=0$ and $E(|\epsilon_{11}|^\tau)<\infty$ for some $\tau\geq 4$. Suppose $T/(N^{8/\tau}\log N)\to \infty$ as $N\to\infty$. Then \begin{equation*} \sqrt{T\log N}\cdot\max_{1\leq i <j \leq N}\big|\hat{\rho}_{ij} - \tilde{\rho}_{ij}\big|\to 0 \end{equation*} in probability as $N\to\infty.$

\noindentProof of Proposition (ref). To prove the result, by the homogeneity of $\hat{\rho}_{ij}$ from (ref), without loss of generality, we assume $E(\bold{\epsilon}_{11}^2)=1$. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\alpha_N\to 0$ as $N\to\infty$ by assumption. From Lemma (ref), for any $h\in (0, 14)$,

equation[equation omitted — 228 chars of source]

Next we estimate the last probability.

For any $v>0$,

align[align omitted — 317 chars of source]

Now

align[align omitted — 325 chars of source]

Note that $\|\bold{\epsilon}_1\|^2=\sum_{j=1}^T\bold{\epsilon}_{1j}^2$. Since $E(\bold{\epsilon}_{11}^2)=1$ and $E(|\bold{\epsilon}_{11}|^\tau)<\infty$, we see

align[align omitted — 323 chars of source]

by (ref), where the Markov inequality is applied in the second inequality. Write $\bold{Q}=\Gamma'\bold{D}\Gamma$ where $\Gamma=(\gamma_{ij})_{T\times T}$ is an orthogonal matrix and

eqnarray*[eqnarray* omitted — 101 chars of source]

Then $\bold{\epsilon}_1'\bold{Q}\bold{\epsilon}_1=\sum_{k=1}^p\big(\sum_{j=1}^T\gamma_{kj}\epsilon_{1j}\big)^2$. It follows that

align[align omitted — 363 chars of source]

where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. Thus, from Lemma (ref) we have that there exists some $K>0$, such that

eqnarray*[eqnarray* omitted — 142 chars of source]

Join this with (ref), (ref) and (ref) to get

eqnarray*[eqnarray* omitted — 180 chars of source]

By taking $v=\frac{h}{28}$, we have from the above and (ref) that

eqnarray*[eqnarray* omitted — 300 chars of source]

which goes to zero by the assumption that $T/(N^{8/\tau}\log N)\to \infty$ as $N\to\infty$. $\Box$

propAssume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables with $E\epsilon_{11}=0$ and $Ee^{\omega|\epsilon_{11}|}<\infty$ for some $\omega>0$. Suppose $\log N=o(T^{1/5})$ as $N\to\infty$. Then \begin{eqnarray*} \sqrt{T\log N}\cdot\max_{1\leq i <j \leq N}\big|\hat{\rho}_{ij} - \tilde{\rho}_{ij}\big|\to 0 \end{eqnarray*} in probability as $N\to\infty.$

\noindentProof of Proposition (ref). To prove the result, by the homogeneity of $\hat{\rho}_{ij}$ from (ref), without loss of generality, assume $E(\bold{\epsilon}_{11}^2)=1$. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\alpha_N\to 0$ as $N\to\infty$ by assumption. From (ref), we have that, for any $h\in (0, 14)$,

align[align omitted — 224 chars of source]

as $N$ is sufficiently large. So to finish the proof it suffices to show the second probability goes to zero.

For any $v>0$, by (ref) and (ref),

align[align omitted — 303 chars of source]

Since $E(\bold{\epsilon}_{11}^2)=1$, by large deviations, there exists a constant $\eta_0>0$ such that

align[align omitted — 161 chars of source]

for large enough $T$; see, for example, DZ98. By (ref),

eqnarray*[eqnarray* omitted — 213 chars of source]

where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. From Lemma (ref), there exists $K>0$ such that

eqnarray*[eqnarray* omitted — 332 chars of source]

as $T$ is large enough, where $v':=(v/p)^{1/2}$ and $K>0$ is a constant not depending on $p$, $N$, $T$ or $\gamma_{kj}$'s. It is easy to see the above goes to zero if $T/(\log N)^5\to \infty.$ It follows that

eqnarray*[eqnarray* omitted — 105 chars of source]

provided $T/(\log N)^5\to \infty.$ The proof is completed. $\Box$

propAssume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous and subgaussian random variables. If $\log N=o(T^{1/3})$, then \begin{eqnarray*} \sqrt{T\log N}\cdot\max_{1\leq i <j \leq N}\big|\hat{\rho}_{ij} - \tilde{\rho}_{ij}\big|\to 0 \end{eqnarray*} in probability as $N\to\infty.$

\noindentProof of Proposition (ref). First, the subgaussian assumption implies that $Ee^{\omega|\epsilon_{11}|^2}<\infty$ for some $\omega>0$. Hence $Ee^{t|\epsilon_{11}|}<\infty$ for all $t>0$. We will use the same notation as in the proof of Proposition (ref). Reviewing (ref) and (ref), to get our desired result, it suffices to show that $N^2\cdot \max_{1\leq i \leq N}P\big(\bold{e}_1'\bold{A}_i\bold{e}_1> 2\alpha_N v\big)\to 0$ for any $v>0$.

By (ref), for each $i$,

eqnarray*[eqnarray* omitted — 199 chars of source]

where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. From Lemma (ref), there exists $K>0$ such that

eqnarray*[eqnarray* omitted — 129 chars of source]

Therefore, by (ref) and (ref), there exists a constant $\beta_0>0$ such that

eqnarray*[eqnarray* omitted — 173 chars of source]

It is easy to see the above goes to zero if $\log N=o(T^{1/3}).$ $\Box$

Assume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables. Set $\bar{\bold{\epsilon}}_i=(1/T)\sum_{j=1}^T\epsilon_{ij}$ for all $i$. Define $\bold{1}=(1, \dots, 1)'\in \mathbb{R}^T.$ The Pearson correlation coefficient $\rho_{ij}$ is then defined by

align[align omitted — 235 chars of source]

for $1\leq i, j\leq N.$ Similar to the clarification below (ref), the “i.i.d. continuous" assumption justifies that $\rho_{ij}$ is well-defined.

propLet $\tilde{\rho}_{ij}$ be as in Lemma (ref). Assume $E\epsilon_{11}=0$ and $E(|\epsilon_{11}|^\tau)<\infty$ for some $\tau\geq 2$. If $\frac{N^{4/\alpha}\log N}{T}\to 0$, then \begin{eqnarray*} \sqrt{T\log N}\cdot \max_{1\leq i <j \leq N}\big|\tilde{\rho}_{ij}-\rho_{ij}\big|\to 0 \end{eqnarray*} in probability as $N\to\infty$.

\noindentProof of Proposition (ref). The proof consists of two steps. In the first step we will show

align[align omitted — 185 chars of source]

By using this we will prove the desired conclusion in the second step.

{\it Step 1}. Set $\delta_i=\sqrt{T}\cdot\frac{\bar{\bold{\epsilon}}_i}{\|\bold{\epsilon}_i\|}$ for each $i.$ Recall $\tilde{\rho}_{ij}=\frac{\bold{\epsilon}_i'\bold{\epsilon}_j} {\|\bold{\epsilon}_i\|\cdot\|\bold{\epsilon}_j\|}$. Write

eqnarray*[eqnarray* omitted — 319 chars of source]

It follows that

align[align omitted — 209 chars of source]

By the inequality $(1-x)^{-1/2}\leq 1+2|x|$ for $|x|\leq \frac{1}{2}$ appeared in the proof of Lemma (ref), we have

eqnarray*[eqnarray* omitted — 150 chars of source]

provided $\max_{1\leq i \leq N} |\delta_i|<\frac{1}{2}.$ Under this restriction, $(1-\delta_i^2)^{-1/2}\leq \frac{2}{\sqrt{3}} \leq 2$ for each $i$. Therefore,

eqnarray*[eqnarray* omitted — 110 chars of source]

Since $|\tilde{\rho}_{ij}|\leq 1$, the above two estimates joining with (ref) implies that

eqnarray*[eqnarray* omitted — 209 chars of source]

as $\max_{1\leq i \leq N} |\delta_i| \leq \frac{1}{2}.$ Moreover, the above naturally holds if $\max_{1\leq i \leq N} |\delta_i| > \frac{1}{2}.$ This leads to (ref).

{\it Step 2}. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\lim_{N\to\infty}\alpha_N= 0$. From {\it Step 1}, for any $t>0$,

eqnarray*[eqnarray* omitted — 278 chars of source]

Therefore, to show $\sqrt{T\log N}\cdot\max_{1\leq i< j \leq N}|\rho_{ij}-\tilde{\rho}_{ij}| \to 0$, it is enough to prove that

eqnarray*[eqnarray* omitted — 80 chars of source]

for any $s>0.$ In fact,

eqnarray*[eqnarray* omitted — 340 chars of source]

where $\{\xi_j;\, 1\leq j \leq T\}$ are i.i.d. random variables with the same distribution of $\bold{\epsilon}_{11}$. The reason we switch the notations from $\{\bold{\epsilon}_{ij}\}$'s to $\{\xi_j;\, 1\leq j \leq T\}$ is for the brevity of symbols. By (ref),

eqnarray*[eqnarray* omitted — 102 chars of source]

By the Markov inequality and (ref) as used in (ref),

eqnarray*[eqnarray* omitted — 199 chars of source]

since $\alpha_N=1/\sqrt{T\log N}$. Combing the above assertions, we arrive at

eqnarray*[eqnarray* omitted — 165 chars of source]

which converges to zero provided $\frac{N(\log N)^{\alpha/4}}{T^{\alpha/4}}\to 0$, or equivalently, $\frac{N^{4/\alpha}\log N}{T}\to 0$ $\Box$

Finale: proofs of Theorems (ref), (ref) and (ref)

With preparations earlier, we are now ready to prove the main theorems on the maximum statistics of sample correlation coefficients.

\noindentProof of Theorem (ref). Under the condition $E|\epsilon_{11}|^{6}<\infty$, Jiang2004 and Zhou2007 show that

align[align omitted — 52 chars of source]

converges weakly to a distribution with distribution function $F(y)$, where $L_N'=\max_{1\leq i< j \leq N}|\rho_{ij}|$ and $\rho_{ij}$ is as in (ref). Set $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). Observe that

align[align omitted — 204 chars of source]

Since $E|\epsilon_{11}|^{\tau}<\infty$ with $\tau>8$, by using the assumption $T/N\to c\in (0, \infty)$, we see that $\lim_{N\to\infty}T/(N^{8/\tau}\log N)=\infty$. Hence, by the Proposition (ref)

eqnarray*[eqnarray* omitted — 110 chars of source]

in probability as $N\to\infty$. On the other hand, since $T/N\to c\in (0, \infty)$ and $\tau > 8$ we have $\frac{N^{4/\alpha}\log N}{T}\to 0$. Then, by Proposition (ref),

eqnarray*[eqnarray* omitted — 103 chars of source]

in probability as $N\to\infty$. From (ref) we see that

align[align omitted — 62 chars of source]

in probability. Set $\Delta=L_N-L_N'$. Then

align[align omitted — 83 chars of source]

The Slutsky lemma and (ref) say that $(T/\log N)^{1/2}L_N'\to 2$ in probability. Consequently,

eqnarray*[eqnarray* omitted — 191 chars of source]

in probability by (ref). These together with (ref) conclude that

align[align omitted — 50 chars of source]

By the Slutsky lemma again, this fact and (ref) imply the desired result. $\Box$

\noindentProof of Theorem (ref). By assumption, $Ee^{\omega|\epsilon_{11}|}<\infty$ and $\log N=o(T^{1/5})$. Using Theorem 3 and Remark 2.1 from Cai_J2011 with $``\mu=0"$ and $``\alpha=1"$, we get

align[align omitted — 55 chars of source]

converges weakly to a distribution with distribution function $F(y)$, where $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). By Proposition (ref),

eqnarray*[eqnarray* omitted — 110 chars of source]

in probability as $N\to\infty.$ Recall $L_N=\max_{1\leq i< j \leq N}|\hat{\rho}_{ij}|$. By the triangle inequality, the above says that

eqnarray*[eqnarray* omitted — 54 chars of source]

in probability as $N\to\infty.$ Repeating the argument from (ref) to (ref), we obtain

eqnarray*[eqnarray* omitted — 41 chars of source]

as $N\to\infty.$ The conclusion follows from (ref). $\Box$

\noindentProof of Theorem (ref). By assumption, $\log N=o(T^{1/3})$. Taking $``\mu=0"$ and $``\alpha=2"$ in Theorem 3 and Remark 2.1 from Cai_J2011, we have

align[align omitted — 55 chars of source]

converges weakly to distribution function $F(y)$ for $y \in \mathbb{R}$, where $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). Recall $L_N=\max_{1\leq i< j \leq N}|\hat{\rho}_{ij}|$. By Proposition (ref), under the restriction $\log N=o(T^{1/3})$,

eqnarray*[eqnarray* omitted — 110 chars of source]

in probability as $N\to\infty.$ By the triangle inequality, the above says that

eqnarray*[eqnarray* omitted — 54 chars of source]

in probability as $N\to\infty.$ From the argument between (ref) and (ref), we have

align[align omitted — 52 chars of source]

as $N\to\infty.$ This and (ref) yield the conclusion. $\Box$

The proof of Theorem (ref)

We create a new method to prove Theorem (ref) which gives the asymptotic independence between the sum $S_N$ and the maximum $L_N$. The idea is employing the inclusion-exclusion formula twice. We expect this method to work for other problems regarding asymptotic independence between sums of and maxima of weakly dependent random variables.

Prelude: auxiliary results towards proof of Theorem (ref)

The results stated in this section are about the estimates of probabilities of events related to Gaussian random variables. They are useful in their own right. Their proofs will be presented in Section (ref).

lemmaFor each $N\geq 1$, let $T=T_N\geq 2$ be an integer. Suppose $\bold{s}_1$ and $\bold{s}_2$ are i.i.d. random vectors uniformly distributed on $\mathbb{S}^{T-1}$. Given $y \in \mathbb{R}$, set $l_N=T^{-1/2}\cdot (4\log N -\log\log N+y)^{1/2}$ which makes sense for large $N$. Assume $\log N=o(\sqrt{T})$ as $N\to\infty.$ Then \begin{eqnarray*} \lim_{N\to\infty}N^2\cdot P(\bold{s}_1'\bold{s}_2\geq l_N)= \frac{1}{2\sqrt{2\pi}}e^{-y/2}. \end{eqnarray*}
lemmaSuppose $\bold{s}_1$ and $\bold{s}_2$ are two i.i.d. random vectors uniformly distributed on $\mathbb{S}^{T-1}$ with $T\geq 2.$ Let $\{\xi_1, \cdots, \xi_k\}$ be random variables (not necessarily independent), each of which has the same distribution as that of $\bold{s}_1'\bold{s}_2$. Then \begin{eqnarray*} P\big(\max_{1\leq i \leq k}|\xi_i|\geq t\big) \leq k\cdot e^{-Tt^2/4}+(2k)\cdot e^{-cT} \end{eqnarray*} for all $t>\frac{2}{\sqrt{T}}$, where $c>0$ is a constant free of $k$, $t$ and $T$.
lemmaLet $\{Z, Z_1, \cdots, Z_k\}$ be i.i.d. standard normals. Let $\delta \in (0, 1)$ be given. Set $v_i=\sqrt{\delta}Z+\sqrt{1-\delta}Z_i$ for $1\leq i \leq k$. Then \begin{eqnarray*} P\big(\min_{1\leq i \leq k}v_i>x\big)\leq \frac{1}{y}\exp\Big(-\frac{y^2}{2\delta}\Big)+ \frac{1}{(x-y)^k}\cdot \exp\Big[-\frac{k(x-y)^2}{2(1-\delta)}\Big] \end{eqnarray*} for all $x>y>0$.
lemma(Slepian's lemma from Slepian) Suppose $(U_1, \cdots, U_k)'$ and $(V_1, \cdots, V_k)'$ are two $\mathbb{R}^k$-valued centered Gaussian random vectors such that $EU_i^2=EV_i^2$ and $E(U_iU_j)\leq E(V_iV_j)$ for all $1\leq i, j \leq k.$ Then, for any real numbers $t_1, \cdots, t_k$, \begin{eqnarray*} P(U_i\leq t_i\ for all\ 1\leq i \leq k) \leq P(V_i\leq t_i\,for all\ 1\leq i \leq k). \end{eqnarray*}
lemmaSuppose $\bold{a}_1, \cdots, \bold{a}_k$ are constant unit vectors on $\mathbb{S}^{T-1}$ for some $T\geq 2$. Let $\bold{s}$ be a vector with the uniform distribution on $\mathbb{S}^{T-1}.$ Assume $\max_{1\leq i< j \leq k}|\bold{a}_i'\bold{a}_j|\leq \delta$ for some $\delta\in [0, 1)$. Then \begin{eqnarray*} P\big(\min_{1\leq i \leq k}|\bold{a}_i'\bold{s}|>z\big) &\leq& \frac{2^k}{y}\cdot\exp\Big(-\frac{y^2}{2\delta}\Big)+2\exp(-cT)\\ & &+ \frac{2^k}{(z\sqrt{rT}-y)^k}\cdot \exp\Big[-\frac{k\big(z\sqrt{rT}-y\big)^2}{2(1-\delta)}\Big] \end{eqnarray*} for all $z>0$, $y \in (0, z\sqrt{rT})$, $r\in (0, 1)$ and $c$ is a constant depending on $r$ only.
lemmaLet $\hat{\rho}_{ij}$ be as in (ref). Suppose assumption (ref) holds with $\{\epsilon_{ij};\, 1\leq i\leq N, 1\leq j \leq T\}$ being Gaussian random variables. Recall $\bold{U}_i$ and $\bold{s}_i$ from Lemma (ref). Let $\tau\geq 2$ be given. Then \begin{eqnarray*} E\big(\big|\hat{\rho}_{ij}^2-E(\hat{\rho}_{ij}^2|\bold{s}_i)\big|^{\tau}\big) \leq \frac{K}{m^{\tau}}, \end{eqnarray*} \begin{eqnarray*} E\Big[\,\Big|\sum_{j=i+1}^N (\hat{\rho}_{ij}^2-E\hat{\rho}_{ij}^2)\Big|^{\tau}\Big] \leq K\cdot \Big[\frac{(N-i)^{\tau/2}}{m^{\tau}}+ \frac{(N-i)^{\tau}}{m^{2\tau}}\Big] \end{eqnarray*} and \begin{eqnarray*} E\Big[\,\Big|\sum_{i=1}^{j-1} (\hat{\rho}_{ij}^2-E\hat{\rho}_{ij}^2)\Big|^{\tau}\Big] \leq K\cdot \Big[\frac{(j-1)^{\tau/2}}{m^{\tau}} + \frac{(j-1)^{\tau}}{m^{2\tau}}\Big] \end{eqnarray*} for all $1\leq i < j\leq N$ and $N\geq 3$, where $K$ is a constant depending on $\tau$ only.

Intermezzo: key steps in the proof of Theorem (ref)

After collecting some useful facts in Section (ref), we are now ready to prove Theorem (ref). To make the discussion easier to follow, we give the outline first.

First, Let $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Review the framework between (ref) and (ref). In particular,

align[align omitted — 183 chars of source]

for $1\leq i, j \leq N.$ Assume (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. Then

align[align omitted — 112 chars of source]

are i.i.d. uniformly distributed over the $T$-dimensional unit sphere $\mathbb{S}^{T-1}.$ For fixed $y\in \mathbb{R}$, set

align[align omitted — 77 chars of source]

Here is the structure of the proof of Theorem (ref).

1. Let $\bold{e}_i$ be as in (ref). Define $\tilde{L}_N=\max_{1\leq i<j \leq N}|\bold{e}_i'\bold{e}_j|$. To show that $TL_N^2-4\log N +\log\log N$ and $(S_N-\mu_N)/N$ are asymptotically independent, it is enough to prove that $T\tilde{L}_N^2-4\log N +\log\log N$ and $(S_N-\mu_N)/N$ are asymptotically independent (Lemma (ref)). The benefit of this step is that $\{\bold{e}_i'\bold{e}_j;\ 1\leq i<j\leq N\}$ are identically distributed. This is not true for $\{\hat{\rho}_{ij};\, 1\leq i<j \leq N\}$ appeared in definition of $L_N$.

2. Review (ref). To show the asymptotic independence, it suffices to prove

align[align omitted — 123 chars of source]

for any real numbers $x$ and $y$, where $\Phi(x)$ is the cdf of $N(0, 1)$ and is also the limiting distribution function of $\frac{1}{N}(S_N-\mu_N)$; $F(y)$ is the Gumbel distribution and is also the limiting distribution function of $\tilde{L}_N$. Recall the definition of $\tilde{L}_N$, we are able to write the event in (ref) as the union of $\binom{N}{2}$ many events which are exchangeable. Then, by using the inclusion-exclusion formula, the probability in (ref) is sandwiched between two bounds [(ref) and (ref)]. The advantage is that we reduce the probability on the global maximum “$\tilde{L}_N$" to sums of probabilities on “local maxima".

3. In dealing with the “local maxima", each probability in the sum is of the form $P(\frac{1}{N}(S_N-\mu_N)\leq x, |\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$, where $n$ is a fixed number free of $N$ and $T$, and where the indices $\{(i_l, j_l);\, 1\leq l \leq n\}$ are different. Review $S_N$ is the sum of $(\bold{e}_i'\bold{e}_j)^2$ over all $1\leq i<j \leq N$. Remove the terms related to $\{\bold{e}_{i_l}'\bold{e}_{j_l};\, 1\leq l \leq n\}$ from $S_N$, in other words, eliminate the terms $(\bold{e}_i'\bold{e}_j)^2$ for all $(i, j)$ with $\{i, j\}\cap \{i_l, j_l\} \ne \emptyset$ for some $1\leq l\leq n$. Then the resulting sum is independent of $\{\bold{e}_{i_l}'\bold{e}_{j_l};\, 1\leq l \leq n\}$, and hence $P(\frac{1}{N}(S_N-\mu_N)\leq x, |\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ is asymptotically the product of $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$. Of course we have to handle the “loss" after removing the terms. It turns out that the removed terms are very concentrated at their mean values by the second and third conclusions from Lemma (ref). So the probability $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and the modified version $P(\frac{1}{N}(\tilde{S}_N-\tilde{\mu}_N)\leq x)$ are asymptotically equal. The total errors in the above approximations is negligible (Lemma (ref)).

4. In step 3, we have showed that

center[center omitted — 130 chars of source]

is asymptotically the product of $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ in (ref) and (ref), where $A_N=\{\frac{1}{N}(S_N-\mu_N)\leq x\}$ and $B_{I}=\{|\bold{e}_i'\bold{e}_j|>l_N\}$ for $I=(i, j)$. We will use one more time the inclusion-exclusion formula to regroup the sum of probabilities $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ and change it to $P(\max_{1\leq i<j \leq N}|\bold{e}_{i}'\bold{e}_{j}|>l_N)$. Note that the original upper bound becomes the lower bound of $P(\max_{1\leq i<j \leq N}|\bold{e}_{i}'\bold{e}_{j}|>l_N)$ and similarly the original lower bound becomes the new upper-bound. There are some “middle" terms in between the bounds, we have to show they are negligible. This is guaranteed by Lemma (ref).

Now let us execute the steps streamlined above.

lemmaLet $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Let $\{\bold{e}_i;\ 1\leq i \leq N\}$ be defined in (ref). Set $\tilde{L}_N=\max_{1\leq i<j \leq N}|\bold{e}_i'\bold{e}_j|$. Assume $N=o(T^2)$ and (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. If $T\tilde{L}_N^2-4\log N +\log\log N$ and $\frac{1}{N}(S_N-\mu_N)$ are asymptotically independent, then $TL_N^2-4\log N +\log\log N$ and $\frac{1}{N}(S_N-\mu_N)$ are also asymptotically independent.

\noindentProof of Lemma (ref). Let $m=T-p.$ Under assumption (ref) with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$, we know $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}.$ Define $\tilde{\rho}_{ij}=\bold{e}_i'\bold{e}_j$ for $1\leq i<j \leq N.$ To organize the proof clearly, we list the relevant quantities as follows.

eqnarray*[eqnarray* omitted — 299 chars of source]

where $\bold{P}_i$ is defined in (ref). By Theorems (ref) and (ref), the following hold.

align[align omitted — 202 chars of source]

By Theorem 6 from CJF13, the assertion (ref) is also true if “$L_N$" is replaced by “$\tilde{L}_N$". To show asymptotic independence, it is enough to show

align[align omitted — 138 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$, where $\Phi(x)=(2\pi)^{-1/2}\int_{-\infty}^xe^{-t^2/2}\,dt.$ Let $l_N$ be as in (ref). Due to (ref) and (ref) we know (ref) is equivalent to that

align[align omitted — 121 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. By assumption, we know that

align[align omitted — 128 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. We show next that (ref) implies (ref).

By Proposition (ref),

eqnarray*[eqnarray* omitted — 110 chars of source]

in probability as $N\to\infty$ provided $\log N=o(T^{1/3})$ as $N\to\infty$. By the triangle inequality,

eqnarray*[eqnarray* omitted — 68 chars of source]

in probability.

Given $\epsilon \in (0,1)$. Set

eqnarray*[eqnarray* omitted — 93 chars of source]

for $N\geq 3.$ Then

align[align omitted — 58 chars of source]

Now,

align[align omitted — 172 chars of source]

On $\Omega_N$, if $L_N>l_N$ then

align[align omitted — 112 chars of source]

Define

eqnarray*[eqnarray* omitted — 83 chars of source]

which makes sense for large $N.$ Use the formula $\sqrt{x}-\sqrt{y}=(x-y)/(\sqrt{x}+\sqrt{y})$ for any $x\geq 0$ and $y\geq 0$ to see

align[align omitted — 303 chars of source]

as $N \to\infty.$ Thus,

eqnarray*[eqnarray* omitted — 67 chars of source]

as $N$ is sufficiently large. This and (ref) conclude that

eqnarray*[eqnarray* omitted — 44 chars of source]

as $N$ is sufficiently large. Review (ref). We have

eqnarray*[eqnarray* omitted — 174 chars of source]

Immediately from (ref) and (ref) we get

eqnarray*[eqnarray* omitted — 122 chars of source]

for any $\epsilon \in (0,1).$ Inspect that the left-hand side of the above does not depend on $\epsilon$. Letting $\epsilon \downarrow 0$, we obtain

align[align omitted — 124 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}.$ In the following we will show the lower limit.

Evidently,

align[align omitted — 156 chars of source]

Set

eqnarray*[eqnarray* omitted — 84 chars of source]

Similar to (ref), it is checked that

eqnarray*[eqnarray* omitted — 86 chars of source]

as $N\to\infty$. Therefore,

eqnarray*[eqnarray* omitted — 65 chars of source]

as $N$ is sufficiently large. It is straightforward to verify that

eqnarray*[eqnarray* omitted — 184 chars of source]

as $N$ is sufficiently large, where the last inclusion follows from the definition of $\Omega_N$. By (ref),

eqnarray*[eqnarray* omitted — 161 chars of source]

Thus, from (ref) and (ref) we get

eqnarray*[eqnarray* omitted — 122 chars of source]

for any $\epsilon \in (0,1).$ Sending $\epsilon \downarrow 0$ we see

eqnarray*[eqnarray* omitted — 112 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}.$ This together with (ref) concludes (ref). $\Box$

We need some notations now. Let $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Let $\{\bold{e}_i;\ 1\leq i \leq N\}$ be as in (ref). Define

align[align omitted — 190 chars of source]

for any $I=(i,j)\in \Lambda_N$. To make a clear presentation, we impose a trivial ordering for elements in $\Lambda_N$. For any $I_1=(i_1, j_1)\in \Lambda_N$ and $I_2=(i_2, j_2) \in \Lambda_N$, we say $I_1<I_2$ if $i_1<i_2$ or $i_1=i_2$ but $j_1<j_2$.

lemmaRecall the notations from (ref) and (ref). Assume $\log N=o(\sqrt{T})$ as $N\to\infty$. Assume $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1},$ which is particularly true if (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. Set \begin{eqnarray*} H(N, n)=\sum_{I_1< I_2< \cdots < I_{n}\in \Lambda_N}P(B_{I_1}B_{I_2}\cdots B_{I_{n}}). \end{eqnarray*} Then $\lim_{n\to\infty}\limsup_{N\to\infty}H(N, n)=0.$

\noindentProof of Lemma (ref). For $I_l$ appeared in $H(N, n)$, write $I_l=(i_l, j_l)$ for $l=1, \cdots, n.$ Now we classify the indices $I_1< I_2< \cdots < I_{n}\in \Lambda_N$ in the definition of $H(N, n)$ into three cases. Let $\Gamma_{N,1}$ be the set of indices $(I_1, \cdots, I_n)$ such that no two of the $2n$ indices $\{i_l, j_l\,; l=1, \cdots, n\}$ are identical. Let $\Gamma_{N,2}$ be the set of indices $(I_1, \cdots, I_n)$ such that either $i_1=\cdots =i_n$ or $j_1=\cdots = j_n.$ Let $\Gamma_{N,3}$ be the set of indices $I_1< I_2< \cdots < I_{n}\in \Lambda_N$ excluding $\Gamma_{N,1}\cup \Gamma_{N,2}$. In the following we will estimate

eqnarray*[eqnarray* omitted — 102 chars of source]

for $j=1, 2, 3$ one by one. We will see $F_1$ contributes essentially the sum in the expression of $H(N, n)$ by an easy argument; the term $F_2$ is negligible and its computation is trivial; the term $F_3$ is also negligible but its estimate is most involved.

{\it Step 1: the estimate of $F_1$}. Recall $B_{I}=\{|\bold{e}_i'\bold{e}_j|>l_N\}$ if $I=(i,j)\in \Lambda_N$, where $l_N$ is defined in (ref). By the definition of $\Gamma_{N,1}$, we know that $B_{I_1}, B_{I_2} \cdots, B_{I_n}$ are independent. By Lemma (ref) and the symmetry of $\bold{e}_1'\bold{e}_2$,

align[align omitted — 143 chars of source]

for all $N\geq 3$. Then, by the elementary fact $\binom{k}{n}=\frac{1}{n!}k(k-1)\cdots (k-n+1)\leq \frac{k^n}{n!}$ for all $k> n\geq 1.$

align[align omitted — 106 chars of source]

{\it Step 2: the estimate of $F_2$}. Evidently, the size of $\Gamma_{N,2}$ is no more than $\binom{N}{1}\cdot\binom{N}{n}\cdot 2\leq 2N^{n+1}$. We first claim that $\{\bold{e}_1'\bold{e}_2, \bold{e}_1'\bold{e}_3, \cdots, \bold{e}_1'\bold{e}_n\}$ are independent. In fact, let $\bold{e}$ be uniformly distributed on $\mathbb{S}^{T-1}$. Then, $\bold{a}'\bold{e}$ has the same distribution as that of $(1, 0, \cdots, 0)'\bold{e}$ for any $a\in \mathbb{S}^{T-1}$ (see, e.g., Theorem 1.5.7(i) and the argument for (5) on p.147 from mui). Since $\bold{e}_1, \cdots, \bold{e}_n$ are i.i.d. random vectors, we know that, conditioning on $\bold{e}_1$, the random variables $\{\bold{e}_1'\bold{e}_2, \bold{e}_1'\bold{e}_3, \cdots, \bold{e}_1'\bold{e}_n\}$ are i.i.d. with a common distribution of $(1, 0, \cdots, 0)'\bold{e}$. In particular, their conditional distributions do not depend on $\bold{e}_1$. This proves the claim. Consequently,

align[align omitted — 227 chars of source]

by (ref).

{\it Step 3: the estimate of $F_3$}. Fix a tuple $(I_1, I_2, \cdots , I_{n})\in \Gamma_{N,3}$. By the ordering imposed on $\Lambda_N$, we see that $i_1\leq i_2\leq \cdots \leq i_n$. There are two different cases: (1) $i_1<i_2$; (2) there exists $2\leq k\leq n-1$ such that $i_1=\cdots =i_k<i_{k+1}$.

Under case (1), let $\mathcal{F}_1$ be the set of random vectors $\{\bold{e}_{j_1}, \bold{e}_{i_l}, \bold{e}_{j_l};\, 2\leq l\leq n\}$ (the first index is “$j_1$" which is different from the third one “$j_l$"). Then, by independence and the property “take out what is known" for the conditional probability,

eqnarray*[eqnarray* omitted — 235 chars of source]

As a fact used earlier, the conditional distribution of $\bold{e}_{i_1}'\bold{e}_{j_1}$ given $\bold{e}_{j_1}$ and the unconditional distribution of $\bold{e}_{i_1}'\bold{e}_{j_1}$ are identical. Therefore, by (ref),

align[align omitted — 115 chars of source]

Let us study case (2). Without loss of generality, for notational clarity, we assume $i_1=\cdots =i_k=1$ and $i_{k+1}=2$. Denote by $\mathcal{F}_2$ the set of random vectors $\{\bold{e}_{i_l}, \bold{e}_{j_l};\, 1\leq l\leq n\}$ excluding $\bold{e}_1$. Then use conditional probability and independence to see

align[align omitted — 253 chars of source]

where $P_1$ stands for the condition probability given $\mathcal{F}_2$. By independence, the last probability in (ref) is computed by treating $\bold{e}_1$ as a random variable while fixing the values of $\bold{e}_{j_1}, \cdots, \bold{e}_{j_{k}}$. To study the $P_1$, we need to understand the relationship among $\{\bold{e}_{j_1}, \cdots, \bold{e}_{j_{k}}\}.$ To do so, set

eqnarray*[eqnarray* omitted — 177 chars of source]

By Lemma (ref) and the fact $k\leq n$,

align[align omitted — 150 chars of source]

as $N$ is sufficiently large provided $\log N=o(T)$. Notice that

align[align omitted — 215 chars of source]

We claim that, for any $\epsilon\in (0,1)$, there exists an integer $N_{\epsilon}\geq 1$ such that

align[align omitted — 148 chars of source]

as $N\geq N_{\epsilon}$. On $\Omega_N$, we know $\max_{j_1\leq l_1<l_2\leq j_{k}}|\bold{e}_{j_{l_1}}'\bold{e}_{j_{l_2}}|< \delta_N.$ Take $r=1-\frac{\epsilon}{4k}$, $y=(\log N)^{1/4}$, $z=l_N$, $\delta=\delta_N$ in Lemma (ref). Observe that $2^k/y\to 0$. By (ref), $l_N\sim 2\sqrt{(\log N)/T}$ as $N\to\infty$, and hence $y=o(z\sqrt{rT})$. Also, $\delta_N\to 0$ since $\log N=o(T)$.

eqnarray*[eqnarray* omitted — 85 chars of source]

Thus, by the lemma, use the facts that $2rk>2k-\frac{3}{4}\epsilon$ and that $z\sqrt{rT}-y\to \infty$ to get

eqnarray*[eqnarray* omitted — 411 chars of source]

as $N\geq N_{\epsilon}$ thanks to the assumption $\log N=o(\sqrt{T})$, where $N_{\epsilon}\geq 1$ is an integer depending on $\epsilon$ only. This leads to (ref).

Now, combining (ref) and (ref), we arrive at

eqnarray*[eqnarray* omitted — 145 chars of source]

as $N\geq N_{\epsilon}$. This together with (ref) and (ref) implies

eqnarray*[eqnarray* omitted — 155 chars of source]

as $N$ is sufficiently large. In summary, by using the above conclusion and (ref), for any $\epsilon\in (0, 1)$ and any $(I_1, \cdots, I_n)\in \Gamma_{n,3}$,

eqnarray*[eqnarray* omitted — 161 chars of source]

as $N\geq N_{\epsilon}$, where $k_1$ is the number of elements on the $i_1$-th row of $\{I_1, \cdots, I_n\}$. In words, when we consider $P(B_{I_1}B_{I_2}\cdots B_{I_{n}})$ based on the positions of $I_j$'s appeared in the upper triangular matrix $\Lambda_N=\{(i, j);\, 1\leq i<j \leq N\}$, after reducing the first row we see the connection between the old and new probabilities. Similarly, let $k_j$ be the number of elements from $\{I_1, \cdots, I_n\}$ on the $j$-th row for $j\geq 1.$ Then

eqnarray*[eqnarray* omitted — 342 chars of source]

where the two intersections above run over all elements from $\{I_1, \cdots, I_n\}$ excluding the first two rows. Continue the process recursively to see

eqnarray*[eqnarray* omitted — 136 chars of source]

where $b$ is the total number of rows of $\{I_1, \cdots, I_n\}$ in the upper triangular matrix $\Lambda_N=\{(i, j);\, 1\leq i<j \leq N\}$. Obviously, $k_1+\cdots +k_b=n$ and $b\leq n.$ Therefore, for each $\epsilon \in (0, 1)$,

eqnarray*[eqnarray* omitted — 118 chars of source]

for $N\geq N_{\epsilon}$. This gives that

align[align omitted — 231 chars of source]

Recall $I_l=(i_l, j_l)$ for each $1\leq l \leq n.$ In view of the definition of $\Gamma_{N,3}$, there are at least two of the $2n$ indices from $\{(i_l, j_l);\, 1\leq l \leq n\}$ are identical for any $(I_1, \cdots, I_n)\in \Gamma_{N,3}$. Let $\kappa=|\{i_l, j_l; 1\leq l \leq n\}|$ for such $(I_1, I_2, \cdots , I_{n})$. Easily, $n+1\leq \kappa\leq 2n-1$. To see how many such $(I_1, \cdots, I_n)$ with $|\{i_l, j_l; 1\leq l \leq N\}|=\kappa$, first pick $\kappa$ many indices from $\{1, 2, \cdots, N\}$, which has the total number of ways $\binom{N}{\kappa}\leq N^{\kappa}$, then use the $\kappa$ many indices to make a $(I_1, \cdots, I_n)\in \Gamma_{N,3}$. The total number of ways to do so is no more than $\kappa^{2n}$. Therefore,

eqnarray*[eqnarray* omitted — 116 chars of source]

As a consequence, for each $\epsilon \in (0, 1)$, from (ref) we have

eqnarray*[eqnarray* omitted — 108 chars of source]

as $N\geq N_{\epsilon}$. Take $\epsilon=\frac{1}{2n}$ to see $\lim_{N\to\infty}F_3= 0.$ Joining this with (ref) and (ref), we eventually arrive at

align[align omitted — 74 chars of source]

for each $n\geq 3.$ The desired conclusion then follows by sending $n\to \infty.$ $\Box$

lemmaRecall the notations from (ref) and (ref). Assume (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. If $N=o(T^2)$ as $N\to \infty$, then \begin{eqnarray*} \sum_{I_1< I_2< \cdots < I_{n}\in \Lambda_N}\big[P(A_NB_{I_1}B_{I_2}\cdots B_{I_{n}}) - P(A_N)\cdot P(B_{I_1}B_{I_2}\cdots B_{I_{n}})\big]\to 0 \end{eqnarray*} as $N\to\infty$ for each $n\geq 1.$

\noindentProof of Lemma (ref). From assumption that (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$, we know from (ref) that $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}$. For $I_1< I_2< \cdots < I_{n}\in \Lambda_N$, write $I_l=(i_l, j_l)$ for $l=1,2,\cdots, n$. Set

eqnarray*[eqnarray* omitted — 142 chars of source]

for $n\geq 1.$ It is easy to check that $|\Lambda_{n,N}| =\sum_{l=1}^n(N-i_l+j_l-2)$. Since $i_l<j_l$ for each $l$, we see that

eqnarray*[eqnarray* omitted — 77 chars of source]

Recall $m=T-p$ and

eqnarray*[eqnarray* omitted — 168 chars of source]

where $\bold{P}_i$ is defined as in (ref). Define

eqnarray*[eqnarray* omitted — 86 chars of source]

for $N\geq 3$ and

eqnarray*[eqnarray* omitted — 74 chars of source]

for $N\geq n\geq 1.$ Observe that $B_{I_1}B_{I_2}\cdots B_{I_{n}}$ is an event generated by random vectors $\{\bold{e}_i, \bold{e}_j ;\, (i, j) \in \Lambda_{n,N}\}$. A crucial observation is that $S_N-S_{N,n}$ is independent of $B_{I_1}B_{I_2}\cdots B_{I_{n}}$. It is easy to see that

eqnarray*[eqnarray* omitted — 216 chars of source]

For any integer $\tau\geq 2$, from a convex inequality we have

eqnarray*[eqnarray* omitted — 317 chars of source]

by Lemma (ref), where the constant $C$ is free of $N$ and $T$, and where the last step follows from the assumption $N=o(T^2)$. Similarly,

eqnarray*[eqnarray* omitted — 86 chars of source]

Lastly, by Lemma (ref) again,

eqnarray*[eqnarray* omitted — 213 chars of source]

Therefore,

eqnarray*[eqnarray* omitted — 88 chars of source]

Fix $\epsilon\in (0,1).$ By the Markov inequality,

align[align omitted — 196 chars of source]

for all $N\geq n^2$, where $C'$ is a constant depending on $\epsilon$ but free of $N$, $T$ or indices $\{I_1, \cdots, I_n\}.$

Fix $I_1< I_2< \cdots < I_{n}\in \Lambda_N$. By (ref) and the definition of $A_N(x)$,

eqnarray*[eqnarray* omitted — 505 chars of source]

by the independence between $S_N-S_{N,n}$ and $B_{I_1}B_{I_2}\cdots B_{I_{n}}$. Now

eqnarray*[eqnarray* omitted — 426 chars of source]

Combing the two inequalities to get

align[align omitted — 202 chars of source]

Similarly,

eqnarray*[eqnarray* omitted — 431 chars of source]

In other words, by independence,

eqnarray*[eqnarray* omitted — 227 chars of source]

Furthermore,

eqnarray*[eqnarray* omitted — 323 chars of source]

The above two strings of inequalities imply

eqnarray*[eqnarray* omitted — 204 chars of source]

which joining with (ref) yields

eqnarray*[eqnarray* omitted — 225 chars of source]

where

eqnarray*[eqnarray* omitted — 105 chars of source]

In particular,

align[align omitted — 108 chars of source]

as $N\to\infty$ by Theorem (ref). As a consequence,

eqnarray*[eqnarray* omitted — 466 chars of source]

where

eqnarray*[eqnarray* omitted — 102 chars of source]

as defined in Lemma (ref). From (ref), we know $\limsup_{N\to\infty}H(N, n)\leq C/n!$, where $C$ is a universal constant. Picking $\tau=6n$, and using the trivial fact $\binom{r}{s}\leq r^s$ for any integers $1\leq s \leq r$, we have that

eqnarray*[eqnarray* omitted — 149 chars of source]

Hence, from (ref)

eqnarray*[eqnarray* omitted — 206 chars of source]

for any $\epsilon>0$. The desired result follows by sending $\epsilon \downarrow 0.$ $\square$

Finale: proof of Theorem (ref)

We now are ready to assemble everything together.

\noindentProof of Theorem (ref). Recall $\{\bold{e}_i;\ 1\leq i \leq N\}$ in (ref). By assumption (ref), we see that $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}$. As in Lemma (ref), define

eqnarray*[eqnarray* omitted — 76 chars of source]

Let $m=T-p.$ Recall

eqnarray*[eqnarray* omitted — 205 chars of source]

By Theorem 3 and Remark 2.1 from Cai_J2011 and Theorem (ref), the following hold.

align[align omitted — 204 chars of source]

To show asymptotic independence, by Lemma (ref), it is enough to show

eqnarray*[eqnarray* omitted — 136 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$, where $\Phi(x)=(2\pi)^{-1/2}\int_{-\infty}^xe^{-t^2/2}\,dt.$ Review (ref) to see

align[align omitted — 74 chars of source]

which makes sense for large $N$. Because of (ref) and (ref), the above is equivalent to that

align[align omitted — 128 chars of source]

for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. Review notations $ \Lambda_N$, $A_N$ and $B_{I}$ for any $I=(i,j)\in \Lambda_N$ in (ref). Write

align[align omitted — 127 chars of source]

Here the notation $A_NB_I$ stands for $A_N\cap B_I$. From the inclusion-exclusion principle,

align[align omitted — 293 chars of source]

and

align[align omitted — 291 chars of source]

for any integer $k\geq 1$. Reviewing the definition

eqnarray*[eqnarray* omitted — 102 chars of source]

for $n\geq 1$ in Lemma (ref), we have from the lemma that

align[align omitted — 72 chars of source]

Set

eqnarray*[eqnarray* omitted — 173 chars of source]

for $n\geq 1.$ By Lemma (ref),

align[align omitted — 61 chars of source]

for each $n\geq 1$. The assertion (ref) implies that

align[align omitted — 463 chars of source]

where the inclusion-exclusion formula is used again in the last inequality, that is,

eqnarray*[eqnarray* omitted — 275 chars of source]

for all $k\geq 1$. By the definition of $l_N$ and (ref),

eqnarray*[eqnarray* omitted — 130 chars of source]

as $N\to\infty$. By (ref), $P(A_N)\to \Phi(x)$ as $N\to\infty.$ From (ref), by fixing $k$ first and sending $N\to \infty$ we get from (ref) that

eqnarray*[eqnarray* omitted — 161 chars of source]

Now, let $k\to \infty$ and use (ref) to see

align[align omitted — 132 chars of source]

By applying the same argument to (ref), we see that the counterpart of (ref) becomes

eqnarray*[eqnarray* omitted — 463 chars of source]

where in the last step we use the inclusion-exclusion principle such that

eqnarray*[eqnarray* omitted — 279 chars of source]

for all $k\geq 1$. Review (ref) and repeat the earlier procedure to see

eqnarray*[eqnarray* omitted — 120 chars of source]

by sending $N\to \infty$ and then sending $k\to\infty.$ This and (ref) yield (ref). The proof is completed. $\Box$

Acknowledgment

Professors Feng and Liu thank NSFC grants 11501092 and 11571068 for partially support. Professor Jiang thanks NSF Grants DMS-1406279 and DMS-1916014 for partially support.