Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
175,514 characters · 30 sections · 59 citation commands
Max-Sum tests for cross-sectional dependence of high-dimensional panel data
In this paper we will study the cross-sectional dependence for the following linear regression model for panel data
for $i=1,\cdots,N$ and $t=1,\cdots,T$, where $i$ represents households, individuals, firms, etc., and $t$ represents time. In the literature of panel data, the index $i$ stands for {\it sections}. For each section $i$, the corresponding model is a standard multiple linear regression model, where $y_{it}\in \mathbb{R}$ is the dependent variable and $x_{it}\in \mathbb{R}^p$ is the regressor with slope parameter $\beta_i\in \mathbb{R}^p$. The first coordinate of $x_{it}$ is one if there is an intercept in the linear regression model (ref). The value of $\beta_i$ may vary across $i$. In (ref), we assume $\{\epsilon_{it};\, 1\leq t\leq T\}$ are independent and identically distributed (i.i.d.) for each section $i$. However, across sections the random errors may be dependent, that is, $\{\epsilon_{it};\, 1\leq i\leq N\}$ may be dependent for some $t.$ Such dependence is referred to as cross-sectional dependence. The objective of this paper is to test if there exists cross-sectional dependence by using a few of new methods. Before stating our results, we will introduce some background next.
In statistics and econometrics, panel data or longitudinal data are multi-dimensional data involving measurements over time, which contain observations of various phenomena over multiple time periods for the same unit, for instance, a household or a firm. In the study of panel data models, the cross-sectional dependence is an important concept, described as the interaction between cross-sectional units, which could arise from the behavioral interaction between units.
Stephan Stephan1934 argues that “in dealing with social data, we know that by virtue of their very social character, persons, groups and their characteristics are interrelated and not independent. " However, to make theoretical study easier, experts assume cross-sectional independence in various model setups HsiaoP2012, Pesaran2004. If data across individuals are dependent, inferences under the assumption of cross-sectional independence would be inaccurate and misleading; see HsiaoP2012, pesaran2015testing and the literature therein. To this end, testing the existence of cross-sectional dependence is an important task, which has attracted more attention in recent years, see, for instance, Chudik2013, moscone2009a, pesaran2015testing, Pesaran2015,ST12.
Perhaps the most widely known test for cross-sectional independence is the Lagrange Multiplier (LM) statistic proposed by Breusch and Pagan breusch1980the in 1980 (Google records 5353 citations currently). Their test statistic is the sum of squares of sample correlation coefficients between the residuals from the ordinary least square (OLS). Precisely, for each $i$, let $\hat{\beta}_i$ be the standard estimator of $\beta_i$ in the linear regression for observations $\{(y_{it},x_{it});\, t=1,\cdots, T\}$ and the quantity $\hat{\bold{\epsilon}}_{it}=y_{it}-x_{it}'\hat{\beta}_i$ denotes the residual. For each $i,j=1,\cdots,N$, define the sample correlation $\hat{\rho}_{ij}$ by
Breusch and Pagan breusch1980the propose the Lagrange multiplier test statistic defined by
To get the rejection region, we need to figure out the limiting distribution of $S_N$ as $N$ goes to infinity. Under the null hypothesis that there is no cross-sectional dependence, that is, $\{\epsilon_{it};\, 1\leq i\leq N, 1\leq t \leq T\}$ from (ref) are independent, the asymptotic distribution of $S_N$ is understood when the cross-sectional dimension $N$ is fixed and the time dimension $T$ goes to infinity. In fact, assuming that $\epsilon_{it}$'s are normally distributed, Breusch and Pagan breusch1980the show that, for fixed $N$,
in distribution as $T\to\infty$, where $d=N(N-1)/2.$ If $N$ is relatively large, the above chi-square approximation is not accurate Pesaran2004. A natural amendment is approximating $\chi^2(d)$ by the standard normal distribution: $(\chi^2(d)-d)/\sqrt{2d}$ goes to the standard normal distribution as $N$ goes to infinity. However, as both $N$ and $T$ are very large, taking limit by sending $T \to\infty$ followed by sending $N\to\infty$ is not legitimate mathematically, and the approximation may not be accurate statistically (our Remark (ref) shows such an example). For this consideration Pesaran Pesaran2004 and Pesaran {\it et al.} pesaran2008a provide two versions of normalization of $S_N$ and conjecture that both versions satisfy the central limit theorem (CLT); some of the insights why the CLTs hold can be seen, for example, from ST12 and Pesaran2015. In this paper we prove the two conjectures in Theorems (ref) and (ref). This enables us to carry out the test for cross-sectional dependence through $S_N$ in (ref). In the future, when $S_N$ is used to be a test statistic, we call it the sum test.
On the other hand, when data are sparse, experts in recent years realize that a better test than sum statistics is the maximum of sample correlation coefficients. This is confirmed in, for example, Cai2014; see also caiLiu2011Adaptive, cai2013two-sample and caiZhang2016. With this philosophy in mind, to test the cross-sectional dependence when the residuals $\hat{\bold{\epsilon}}_{it}$ are sparse, we propose statistic
where $\hat{\rho}_{ij}$ is defined as in (ref). Later, when $L_N$ is used to be a test statistic, we refer it to as the max test. Its limiting distribution is obtained as both $N$ and $T$ go to infinity under various moment conditions (Theorems (ref), (ref) and (ref)). The corresponding rejection region based on the test statistic $L_N$ is given after Theorem (ref).
In practice it is hard to tell or differentiate if a set of data is sparse. We then combine the sum test $S_N$ and the max test $L_N$ to propose another test $C_N$, which is the minimum of the $p$-values corresponding to the tests based on $S_N$ and $L_N$. We prove in Theorem (ref) that, under normalization, $S_N$ and $L_N$ are asymptotically independent as both $N$ and $T$ go to infinity. Hence the limiting distribution of $C_N$ is identified. In further discussions, when $C_N$ is used to be a test statistic, we name it the max-sum test. From simulation we see this test, taking care of both sparsity and non-sparsity cases, is better than the sum test and the max test. The tool of deriving asymptotic independence between the sum and the maximum of random variables is new to our knowledge. It seems a universal method to handle asymptotic independence between random variables of this nature.
To sum up, to test cross-sectional dependence for panel data models, in this paper we study three types of tests, {\it i.e.}, the sum test, the max test and the max-sum test. To carry the test, we have solved two open problems about the CLTs for the sum of squares of residuals; the limiting distributions of the maxima of residuals are systematically studied; a new method of studying asymptotic independence is created to develop part of the above theory successfully.
Review model (ref) that $y_{it}=x_{it}'\beta_i+\epsilon_{it}$ for $i=1,\cdots,N$ and $t=1,\cdots,T$, where $i$ indexes the cross-sectional units and $t$ indexes the observations. In this model, $y_{it}\in \mathbb{R}$ is the dependent variable, and $x_{it}\in \mathbb{R}^p$ is the non-random, exogenous regressor with slope parameter $\beta_i\in \mathbb{R}^p$ that are allowed to vary across $i$. We assume $\{\epsilon_{it};\, 1\leq t\leq T\}$ are i.i.d. real-valued random variables for each section $i$. However, across sections the random errors may be dependent, that is, $\{\epsilon_{it};\, 1\leq i\leq N\}$ may be dependent for some $t.$ Such dependence is called cross-sectional dependence. Set
for $i=1,2,\cdots, N.$ Then $\bold{x}_i$ is a $T\times p$ matrix; both $\bold{y}_i$ and $\bold{\epsilon}_{i}$ are $T$-dimensional vectors. Throughout the paper we assume that the $T$ entries of $\epsilon_i$ are i.i.d. with mean zero for each $i$. Recalling (ref), the cross-sectional independence is the same as saying that
In general, although sometimes we assume $\bold{\epsilon}_{1}$ has the normal distribution, we do not need the exact distribution of $\bold{\epsilon}_{1}$ but rather its moments.
First, we list some notations used in the rest of the paper. Reviewing (ref), for each $i=1,\cdots,N$, let
where $\bold{I}_T$ is the $T\times T$ identity matrix and $\bold{P}_i$ is a $T\times T$ projection matrix with $\bold{P}_i^2=\bold{P}_i$ and the rank of $\bold{P}_i$ is $T-p$. For each $i,j=1,\cdots,N$, let $\hat{\rho}_{ij}$ denote the sample correlation coefficient computed by the Ordinary Least Squares (OLS) residuals $(\hat{\bold{\epsilon}}_{i1}, \cdots, \hat{\bold{\epsilon}}_{iT})^T$ and $(\hat{\bold{\epsilon}}_{j1}, \cdots, \hat{\bold{\epsilon}}_{jT})^T$ where $\hat{\bold{\epsilon}}_{it}=y_{it}-x_{it}'\hat{\beta}_i$ for each $i$ and $t$. Under model (ref), it is easy to see that
for each $i$. Thus, by (ref),
In this paper, to test the null hypothesis (ref), we will study three types of tests as follows:
respectively, where
Here, $F(y)= \exp(-e^{-y/2}/\sqrt{8\pi})$ is the extreme-value distribution function of type I, also called the Gumble distribution in literature, and $\Phi(y)$ is the distribution function of $N(0,1)$.
For the sum test in (ref), we will establish that, under $H_0$ in (ref), $(S_N-\mu_N)/N$ converges weakly to the standard normal distribution when both $N$ and $T$ go to infinity with a certain restriction (Theorem (ref)), hence a level-$\alpha$ test will be performed through rejecting $H_0$ when $(S_N-\mu_N)/N$ is larger than the $1-\alpha$ quantile $z_{\alpha}= \Phi^{-1}(1-\alpha)$ of the standard normal distribution.
For the max test in (ref), under $H_0$, we will establish that $TL_N^2-4\log N+\log\log N$ has an asymptotic extreme-value distribution as both $N$ and $T$ go to infinity (Theorems (ref), (ref) and (ref)). We do not impose normality assumptions but rather moment conditions. Recall $F(y)$ is defined below (ref). A level-$\alpha$ test will then be performed by rejecting $H_0$ when $TL_N^2-4\log N+\log\log N$ is larger than the $1-\alpha$ quantile $q_{\alpha}=-\log(8\pi)-2\log\log(1-\alpha)^{-1}$ of $F(y)$.
Furthermore, for the max-sum test in (ref), its asymptotic distribution under $H_0$ is constructed based on the asymptotic independence between $(S_N-\mu_N)/N$ and $TL_N^2-4\log N+\log\log N$ as both $N$ and $T$ go to infinity (Theorem (ref) and Corollary (ref)). So a level-$\alpha$ test will be performed through rejecting $H_0$ when $C_N<1-\sqrt{1-\alpha}$.
In this paper, for the panel data model (ref) we study the cross-sectional dependence. The asymptotic distributions of three test statistics based on residuals are established. As application, three hypothesis tests are accomplished. A real data analysis by using our results is provided. We will now further elaborate below.
In the theoretical part, we have solved two open problems on the sum of squares of residuals conjectured by economists (Pesaran2004, pesaran2008a; see also Pesaran2015, ST12). We have developed an extreme-value theory for the maximum of residuals. Further, a new method is developed to show the sum and the maximum are asymptotically independent. There are not many results in literature to show asymptotic independence between sums of and maxima of random variables. Close references are BBQ98, Xu2016. Our method, being different from earlier literature, provides a general and novel tool for showing asymptotic independence between sums of and maxima of random variables.
In application, we propose three tests on the cross-sectional dependence for high-dimensional panel data: the sum test, the max test and the max-sum test. The max test is the first high-dimensional max test for cross-sectional dependence in panel data models, which is good for sparse residuals while existing test statistics of sum types tend to fail. The sum test is useful for non-sparse residuals, which is clearly demonstrated by simulation in, for example, Pesaran2004, Pesaran2015, pesaran2008a, ST12. We are able to derive the limiting distribution of the sums in this paper.
Furthermore, the max-sum test is constructed based on the asymptotic independence between the max and the sum statistics aforementioned. It is the first max-sum test for studying cross-sectional dependence for high-dimensional panel data. The advantage is that the test works well for both sparse and non-sparse residuals. Comparing the pros and cons of the max test and the sum test, the max-sum test definitely overcomes both disadvantages. Our simulations reveal this fact clearly; see Figure (ref) and its interpretation at the last part of Section (ref). The max-sum test is particularly useful considering it is hard to quantify or determine in practice whether a data set is sparse or not.
We now present the main theoretical results based on the three types of tests in the order of the sum test, the max test and the max-sum test. Their proofs are presented in Section (ref).
Recall that the sum test described in (ref) is a classical one for testing cross-sectional dependence in panel data models. However, the asymptotic theory has not been established yet. Pesaran from Pesaran2004, pesaran2008a conjectures that $S_N$ satisfies the central limit theorem. Some insights on this aspect are given, for example, in ST12 and Pesaran2015. In the following we will present our solution to the problem as well as another one in which the details are given below. The following assumption will be needed throughout the paper. Recall a random variable $V$ is said to be continuous if $P(V=v)=0$ for every $v\in \mathbb{R}.$
If the $T$ entries of $\bold{\epsilon}_{i}$ are i.i.d. continuous random variables, by using a conditional argument, we then trivially have $P(\bold{a}'\bold{\epsilon}_{i}=0)=0$ for any $\bold{a}\in \mathbb{R}^T\backslash\{\bold{0}\}$. This implies that $P(\bold{M}\bold{\epsilon}_{i}= \bold{0})=0$ for any $l\times T$ matrix $\bold{M}\ne 0$ and $l\geq 1$. So $\hat{\rho}_{ij}$ in (ref) is well-defined if $T>p$ because $\mbox{rank}(\bold{P}_i)=T-p$; see the explanation below (ref).
Now we present our solutions to Pesaran's conjectures from Pesaran2004, pesaran2008a as follows. For mathematical rigor, we assume the parameter $T$ depends on $N$. The notation $N_k(\mu, \bold{\Sigma})$ stands for the $k$-dimensional multivariate normal distribution with mean vector $\mu$ and covariance matrix $\bold{\Sigma}$. Although the linear regression model in (ref) requires $p\geq 1$, that is, there are at least one regressors, the following theorem applies to the case that
see (ref).
From the assumption, it is allowed that the number of cross-sectional units $N$ is much larger than the number of observations $T$, for example, $N$ is of order $T^{3/2}$. We apply the framework of the Lindeberg-Feller martingale CLT to study $S_N$. Although the method is simple and is easy to follow, the technical steps are very involved due to the complex nature of the sample correlation coefficients $\hat{\rho}_{ij}$ in (ref). Many computations focus on the conditional means, variances and higher moments.
Considering a possible better convergence rate than the CLT given in Theorem (ref), pesaran2008a revises the statistic $S_N$ and proposes a new one as follows.
where
Pesaran {\it et al}. pesaran2008a conjecture that $Q_N$ also satisfies the CLT. We confirm it in the next theorem.
Our simulation in Figure (ref) shows that the effects of the two approximations in Theorems (ref) and (ref) are too close to be distinguishable. Theorem (ref) allows us to perform a level-$\alpha$ test by rejecting the null hypothesis from (ref) when $(S_N-\mu_N)/N$ is larger than the $1-\alpha$ quantile $z_{\alpha}= \Phi^{-1}(1-\alpha)$ of $N(0,1)$. Theorem (ref) establishes the central limit theorem of $Q_N$ under the same null hypothesis. This provides a theoretical guarantee for $Q_N$ that has been used in econometrics; see, for example, Chudik2013, moscone2009a, pesaran2015testing.
Now we make some comments.
The scenario in the following remark is not practical. We consider it purely for mathematical purposes. They serve for further discussions.
Recall model (ref) and notations in (ref). As in Section (ref), we assume that $\bold{x}_i\bold{x}_i'$ is invertible for each $1\leq i\leq N$ and the quantity $T$ depends on $N$. From (ref) and assumption (ref), we know $\{\hat{\rho}_{ij}:1\leq i,j \leq N\}$ are invariant of $\sigma^2_1,\cdots,\sigma^2_N$, so we are able to assume, without loss of generality, $\sigma_1=\cdots=\sigma_N\equiv 1$. Review $\epsilon_{11}$ in assumption (ref) and $L_N$ in (ref). As explained in (ref), we will also consider the case $p=0$. The main results in this section are presented as follows.
We say $\xi$ is a {\it subgaussian} random variable if there exists $\sigma>0$ such that $Ee^{t\xi}\leq e^{\sigma^2t^2/2}$ for all $t\in \mathbb{R}.$ By the Markov inequality, it is easy to see $P(|\xi|\geq x)\leq 2e^{-x^2/(2\sigma^2)}$ for all $x>0$. As a consequence, $Ee^{\theta \xi^2} < \infty$ for all $\theta < \frac{1}{2\sigma^2}$. Obviously, bounded random variables and Gaussian random variables are all subgaussian random variables.
The strategy of the proofs of Theorems (ref)-(ref) is to approximate $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ for any $p\geq 0$ by $L_N$ for the case $p=0$, in which the limiting behavior is understood in CJF13. We have to show the difference between the two versions of $L_N$ is negligible.
Theorems (ref)-(ref) indicate that we get the same asymptotic distribution of the max statistic $TL_N^2-4\log N +\log\log N$ under different moment assumptions. This allows us to have a flexibility to work on different pairs of $(N, T)$. Under null hypothesis (ref) and the assumptions imposed in the above three theorems, we conclude that a level-$\alpha$ test by rejecting the null hypothesis when $TL_N^2-4\log N+\log\log N$ is larger than the $1-\alpha$ quantile $q_{\alpha}= -\log(8\pi)-2\log\log(1-\alpha)^{-1}$ of $F(y)$.
Review the accounts before the statement of Theorem (ref). We have the following conclusion on asymptotic independence.
By employing a new trick we prove the asymptotic independence between the maximum and the sum of random variables in Theorem (ref). This method is expected to be used in many of such type of problems. In fact, there are few literature to prove asymptotic independence between sums of and maxima of random variables. Some close references are BBQ98, Xu2016. The method here is new. It gives a novel tool to establish asymptotic independence between sums of and maxima of random variables.
To understand the idea quickly, we start with the set-up (ref). The first observation is that the maximum of many random variables, seemingly a global property, can be understood from their local property, that is, the maxima of subsets of random variables with fixed sizes. This step is done through the inclusion-exclusion formula. The second observation is that any such subset of random variables and the sum are independent with very high probability. Consequently the probability of the intersection of the events related to local maxima and the sum can be written as the product of two individual probabilities. Then we use the inclusion-exclusion formula one more time to get the product of two individual probabilities up to negligible errors. Details are shown at the beginning of Section (ref).
An immediate application is given below. By Theorems (ref) and (ref), we know that
Let $\Phi(x)$ be the distribution function of $N(0,1).$ Trivially, both $F(y)$ and $\Phi(x)$ are continuous functions. Set $P_{S_N}=1-\Phi\{(S_N-\mu_N)/N\}$ and $P_{L_N}=1-F(TL_N^2-4\log N+\log\log N)$. By Theorem (ref), (ref) and (ref), we see that $P_{L_N}$ and $P_{S_N}$ are asymptotically independent and each limit is $U[0, 1]$, the uniform distribution over $[0, 1].$ So the following holds easily.
According to Corollary (ref), the proposed max-sum test in (ref) allows us to perform a level-$\alpha$ test by rejecting the null hypothesis (ref) if $C_N<1-\sqrt{1-\alpha}$.
We now conduct simulations to compare the finite sample performance of the tests studied in this paper and another test in literature. The tests we have worked on in this paper are based on $S_N$, $L_N$, $C_N$, $Q_N$ in (ref), (ref), (ref), (ref), respectively. The other one is based on $CD$ from Pesaran2004 defined by
Notice $Q_N$ here is the same as the notation LM$_{\textrm{adj}}$ from pesaran2008a. In the following we will explain our simulation designs and state our simulation findings.
We consider the data generating process used in pesaran2008a, which is specified as
for $i=1, \cdots, N$ and $t=1, \cdots, T$. Comparing the notations in model (ref), we have $x_{it}=(1, x_{2it}, \cdots, x_{pit})^T\in \mathbb{R}^{p}$ and $\beta_i=(\alpha_i, \beta_{2i}, \cdots, \beta_{pi})\in \mathbb{R}^{p}.$
Now we independently generate $\alpha_i\sim N(0,1)$ and $\beta_{li}\sim N(1,0.04)$. The covariates are generated by
for $i=1,\cdots, N,~t=-50, -49, \cdots, T$ and $l=2,\cdots, p$ with $x_{li,-51}=0$, where $v_{lit}\sim N(0,\zeta_{li}^2/(1-0.6^2))$ and $\zeta_{li}^2\sim \chi^2_6/6$. In this case, $\zeta_{li}^2$'s are independently sampled first, then $v_{lit}$'s are independently generated by conditioning on the values of $\zeta_{li}^2$.
Now we generate $\epsilon_{it}$'s under null hypothesis (ref). Let $\epsilon_{it}=\sigma_i w_{it}$, where $w_{it}$'s are generated from three different distributions: (i) $N(0,1)$; (ii) $t_6/\sqrt{6/4}$; (iii) $(\chi^2_5-5)/\sqrt{10}$. Here $t_d$ is the $t$-distribution of degree $d$ and $\chi^2_d$ is the chi-square distribution of degree $d.$ The normalization in (ii) and (iii) is such that each new random variable has mean zero and variance one. Let $\sigma_i^2\sim \frac{p}{2}\chi^2_2$, as in the dynamic setup of pesaran2008a.
We turn to produce data under the alternative hypothesis. Let $\bold{\eta}_t:=(\eta_{1t},\cdots,\eta_{Nt})'$ be generated from the above three different distributions under the null hypothesis. Set $\bold{\epsilon}_{.t}=(\epsilon_{1t},\cdots,\epsilon_{Nt})'={\bm\Sigma}^{1/2}\bold{\eta}_t$. Please differentiate the notation $\bold{\epsilon}_{.t}$ here and $\bold{\epsilon}_{i}$ in (ref). We consider the following two cases of the covariance matrix ${\bm\Sigma}={\bf D}^{1/2}{\bf R}{\bf D}^{1/2}$ with ${\bf D}=\mathrm {diag}\{\sigma^2_1,\cdots,\sigma^2_N\}$.
To ensure that the covariance matrix ${\bm\Sigma}={\bf D}^{1/2}{\bf R}{\bf D}^{1/2}$ is positive definite, we replace the correlation matrix ${\bf R}$ with ${\bf R}+\lambda {\bf I}_N$, where $\lambda:=|\lambda_{\min}({\bf R})|+0.05$ and $\lambda_{\min}({\bf R})$ is the minimum eigenvalue of ${\bf R}$. Then, we consider two choices of the sample size $T=50,100$, and three choices of the dimension $N=50,100,200$.
We now present simulation results on the tests of $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. All the conclusions are based on 1,000 replications. The empirical sizes and powers of these tests in non-sparse and sparse cases are summarized in Tables (ref) to (ref). The power curves are plotted in Figure (ref). We next analyze them in detail.
Table (ref) indicates that all methods have empirical sizes not much larger than 5%. Here, the max test $L_N$ and the max-sum test $C_N$ tend to have smaller empirical sizes than the remaining ones, especially as $T$ is relatively small. This is not very surprising because it is common for maximum methods designed for raw data models; see, for example, Liu2008.
Tables (ref) and (ref) show the information of empirical powers in both non-sparse and sparse cases. Review the sum test $S_N$ and the sum-based test $Q_N$ are originally proposed in breusch1980the and pesaran2008a, respectively. The two are well studied in this paper. Tables (ref) and (ref) show that $S_N$ and $Q_N$ perform best in non-sparse cases in terms of empirical powers, but very poorly in sparse cases. On the contrary, the proposed max test $L_N$ performs the best in sparse cases, but very poorly in dense cases. Interestingly, it can be seen from Figure (ref) that the empirical power performance of our proposed max-sum test $C_N$ is always very close to the optimal one among all of the tests, regardless of the local alternative being sparse or not. This shows a very appealing property for the test $C_N$ which compromises the tests for residuals in both sparse and non-sparse cases. In fact, it is hard to tell in reality if residuals are sparse or not.
Figure (ref) shows the changes of the powers of all the tests as the degree of sparsity changes. Now we explain the procedure to generate the empirical power curves in Figure (ref). In fact, the horizonal direction in the plot is $n$, the degree of sparsity to be defined; the vertical direction represents powers. Specifically, the simulation is designed as follows. Review the general simulation design in Section (ref). We take $T=50$, $N=200$, $p=2$, $n=2,\cdots,16$; $w_{it}$ are generated from normal distributions; a subset $A\subset\{1,\cdots,N\}$ is randomly selected with cardinality $n$; ${\bf R}=(\rho_{ij})_{1\le i,j\le N}$, where $\rho_{ij}=1$ if $i=j$; for $i\neq j$, set $\rho_{ij}=0$ if $i\not \in A$ or $j\not \in A$, and $\rho_{ij}$ has the uniform distribution over $(\sqrt{8(\log n)^{-1}\log N/T}, \sqrt{10(\log n)^{-1}\log N/T})$ if $i\in A$ and $j \in A$. So a larger $n$ means a lower level of sparsity.
Figure (ref) indicates that the empirical power of the max-sum test $C_N$ is always very close to the maximum power of all tests for all $n$. By contrast, the empirical power curves of the remaining methods are monotone, {\it i.e.}, the empirical powers of both $S_N$ and $Q_N$ generally increase with the decrease of sparsity. On the contrary, the empirical power of the max test increases with the increase of sparsity. However, every test excluding the max-sum test $C_N$, favors either the sparse case or the non-sparse case, not both cases simultaneously.
In this section, we apply the five tests to the securities in the Standard $\&$ Poor (S$\&$P) 500 index of large cap U.S. equity market. As seen earlier, they are $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. This demonstrates the practical usefulness of the proposed tests. The S$\&$P 500 index is primarily intended as a leading indicator of U.S. equities. The composition of this index is monitored by Standard and Poor to ensure the widest possible overall market representation while reducing the index turnover to a minimum. In this section, we consider 374 securities that have been included in the S$\&$P 500 index during the whole period from January 2005 to November 2018.
In particular, the panel data on the safe rate of return, and the market factors are obtained from Ken French's data library web page. The one-month US treasury bill rate is chosen as the risk-free rate ($r_{ft}$), the value-weighted return on all NYSE, AMEX, and NASDAQ stocks from CRSP is used as a proxy for the market return ($r_{mt}$), the average return on the three small portfolios minus the average return on the three big portfolios ($SMB_t$), and the average return on two value portfolios minus the average return on two growth portfolios ($HML_t$). SMB and HML are based on the stocks listed on the NYSE, AMEX and NASDAQ. All data are measured in percent per month. During January 2005 to November 2018, a total of 163 consecutive observations are obtained.
The Fama-French three-factor model Fama1993Common is given as follows: \[ y_{it}=r_{it}-r_{ft}=\beta_{0i}+\beta_{1i} (r_{mt}-r_{ft})+\beta_{2i}SMB_t+\beta_{3i} HML_t+\epsilon_{it} \] for each $1\leq i\leq N$ and $1\leq t\leq T$ with $N=374$. We are interested in the following null hypothesis:
That is, we are testing that the 374 variables are independent.
Now we evaluate the performance of the five tests in Section (ref), that is, $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. We randomly sample $T=15,25,35$ observations from the 163 monthly returns. At each value of $T$, the experiment is repeated 1000 times. It is trivial to see
Similarly,
This says that, although there is a dependency when sample $15$ numbers from a total of $163$ numbers for $1000$ times, comparing to $10^{15}$, the number of repeats $1000$ is still reasonable. The same also applies to the cases $T=25$ and $T=35$.
The results are summarized in Table (ref). It suggests that all tests except the max test always reject the null hypothesis of cross-sectional independence. So this indicates the definite cross-sectional dependence among stock returns under the three-factor model by Fama-French. In particular, the max test rejects the null hypothesis when $T$ grows to 35, but never reject it when $T$ reduces to 15. To understand this phenomenon, we point out a well known fact that there may exist a large number of underlying dependencies between stocks in the same industry or relevant industries. This leads us to believe that this is indeed a non-sparse case in which the sum and max-sum tests are more valid.
In this paper we study three tests: the sum test, the max test and the max-sum test, where the latter two are new ones. Two conjectures on the sum test have been settled. A new method to show asymptotic independence between the maximum and the sum of squares of a given set of random variables is established. Now we make some comments.
1. Under the Gaussian assumption, we obtain the CLTs for $S_N$ in Theorems (ref) and (ref). However, the Gaussian assumption is not needed in the study on the maxima of sample correlations in Theorems (ref), (ref) and (ref). One question is whether the Gaussian assumption can be removed from Theorems (ref) and (ref). Our proofs rely on the framework in Lemma (ref) where the normal assumption is essential. Another question is about the restriction between $N$ and $T$ in Theorems (ref) and (ref). Can the assumption “$T/\sqrt{N}\to \infty$" be relaxed? What is the behavior of $S_N$ for other regimes of relationship between $N$ and $T$?
2. The linear regression in (ref) is one of many panel data models; see, for example, the book length treatment in Baltagi2013, Hsiao2014, Pesaran2015, Wooldridge2010, among others. Some of other models can be studied similarly for the properties we have pursued in this paper. We leave them as a future work to our authors.
3. A new way is established to show the asymptotic independence between the sum of and the maximum of a set of random variables. The detail of the method is elaborated at the beginning of Section (ref). We expect this method will also work for other set of random variables of similar feature.
4. For the sum $S_N=\sum_{1\le i<j\le N}T\hat{\rho}_{ij}^2$, Theorem (ref) states that the central limit theorem of $S_N$ involves with projection matrices $\bold{P}_i$ defined via data; see (ref). However, interestingly enough, as shown in Theorems (ref), (ref) and (ref), the behavior of the maximum $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ does not depend on $\bold{P}_i$. Only parameters $N$ and $T$ participate in the limiting process.
5. As seen in Section (ref), a simulation study is carried out for tests based on $S_N$, $L_N$, $C_N$, $Q_N$, $CD$ in (ref), (ref), (ref), (ref), (ref), respectively. It shows that the max-sum test is always very close to the maximum power of all tests for both sparse and non-sparse residuals. By contrast, the empirical power curves of other methods favor only for one of the two types of residuals. In fact, in practice, it is hard to differentiate if a set of numbers is sparse or not. This implies the max-sum test is also desirable for other statistical models as long as three things are known: the central limit theorem holds for the sum of a set of random variables; the maximum of the set of random variables is the Gumbel distribution asymptotically; the sum and the maximum are asymptotically independent.
There are three subsections in this part. In each subsection we first accumulate some first hand or second hand of understanding before the proofs of main theorems are presented. Considering many proofs are involved, we postpone some of them in Appendix. They are interesting in their own right.
In this paper we use the following notation. For a sequence of random variables $\{U_N;\, N \geq 1\}$ and a sequence of constants $\{a_N;\, N\geq 1\}$, the notation $U_N=o_p(a_N)$ means that $U_N/a_N \to 0$ in probability as $N\to\infty;$ we write $U_N=O_p(a_N)$ if $\{U_N/a_N;\, N\geq 1\}$ is stochastically bounded, that is, $\lim_{A\to \infty}\limsup_{N\to \infty}P(|U_N/a_N|\geq A)=0.$ In particular, if $U_N=O_p(a_N)$ then $U_N=o_p(a_Nb_N)$ for any sequence of numbers $\{b_N;\, N\geq 1\}$ with $\lim_{N\to\infty}b_N=\infty.$ We write $a_N\sim b_N$ if $\lim_{N\to\infty}\frac{a_N}{b_N}=1$ for any two sequence of numbers $\{a_N;\, N\geq 1\}$ and $\{b_N;\, N\geq 1\}$.
The proof of Theorem (ref) is lengthy. The main tool is the Lindeberg-Feller central limit theorem for martingales. Automatically many computations of conditional means and variances as well as higher moments are needed for sample correlation coefficients $\hat{\rho}_{ij}$. They are non-trivial. To make the proof organized, we decide to put key steps in a few of sections. This may best facilitate the understanding of readers.
The proofs of the results in this section will be presented in Section (ref).
The following is the Marcinkiewicz-Zygmund inequality; see, e.g., p. 386 and p. 387 from CT88.
In this subsection we develop some identities and inequalities regarding moments of random vectors with the uniform distribution on high-dimensional unit spheres. The tools and methods are of independent interest. The proof of Lemma (ref) is given in this section to show the main idea and starting point. The remaining proofs of other lemmas will be presented in Section (ref).
Review the setting above (ref) and notation $\bold{P}_i$ and $\bold{\epsilon}_i=(\epsilon_{i1},\cdots,\epsilon_{iT})^{'}\in \mathbb{R}^T$ for each $i$. The notation $\mathbb{S}^{m-1}$ represents the unit sphere in the $m$-dimensional Euclidean space.
\noindentProof of Lemma (ref). By the scale-invariance of $\frac{\bold{P}_i\bold{\epsilon}_i}{\|\bold{P}_i\bold{\epsilon}_i\|}$, without loss of generality, assume $\sigma_1=\cdots=\sigma_N=1$. Evidently, $(\bold{U}_i, \bold{0})\bold{\epsilon}_i=\bold{U}_i\bold{\eta}_i$ for each $i$, where $\bold{\eta}_i=(\epsilon_{i1},\cdots,\epsilon_{im})'$. By the orthogonality and (ref),
By the orthogonal invariance of normal distributions and (ref) again, $\bold{P}_i\bold{\epsilon}_i=(\bold{U}_i, \bold{0})\bold{O}_i'\bold{\epsilon}_i$ has the same distribution as that of $(\bold{U}_i, \bold{0})\bold{\epsilon}_i=\bold{U}_i\bold{\eta}_i$. Then $\frac{\bold{P}_i\bold{\epsilon}_i}{\|\bold{P}_i\bold{\epsilon}_i\|}$, as a function of $\bold{P}_i\bold{\epsilon}_i$, has the same distribution as that of
for each $i$ by the first identity of (ref). The desired conclusion then follows from the independence among $\{\bold{\epsilon}_1, \cdots, \bold{\epsilon}_N\}$. $\Box$
Recall notation $(2m-1)!!=1\cdot3\cdots (2m-1)$ for any integer $m\geq 1$. By convention we set $(-1)!!=1$.
A quick reminder is that, although we assume that $\bold{A}$ and $\bold{B}$ are symmetric in (i) above, we do no need that $\bold{A}$, $\bold{B}$ or $\bold{C}$ are symmetric in (ii) and (iii).
In the following we will use notation $\mbox{Var}(\xi_2|\xi_1)$ for conditional variance, which is defined by $E(\xi_2^2|\xi_1)-[E(\xi_2|\xi_1)]^2$ for any random variables $\xi_1$ and $\xi_2.$
In (ref) and (ref), we see parameters $p, N, T$ and variables $\bold{x}_i$. In the rest of the paper, we will use or develop many inequalities where a constant $C$ will appear frequently. The constant $C$ does not depend on $p, N, T$ or $\bold{x}_i$'s and it can be different from line to line. The proofs of the lemma in this section will be given in Section (ref).
A quick comment is that the last term above is non-negative by Lemma (ref).
\noindentProof of Lemma (ref). By (ref), $\bold{U}_{i}\bold{U}_{i}'=\bold{P}_i$. Then
It follows from Lemma (ref)(i) that
for $2\leq j \leq N.$ Therefore,
by the fact that $\mbox{tr}(\bold{P}_i)=T-p=m$ for each $i.$ On the other hand,
where $\bold{M}_{i\blacktriangledown}:=\sum_{j=i+1}^N\bold{M}_{ij}.$ By independence among $\{\bold{s}_i\}$'s and Lemma (ref),
From Lemma (ref), we know $\bold{I}_m-\bold{M}_{ij}$ is non-negative for each $i \ne j.$ Since the sum of non-negative definite matrices is still non-negative definite, we see that $(N-i)\bold{I}_m-\bold{M}_{i\blacktriangledown}$ is also non-negative definite. By Lemma (ref)(ii), $(N-i)^2\bold{I}_m-\bold{M}_{i\blacktriangledown}^2$ is non-negative definite. In particular,
Moreover,
by (ref). Now we estimate $\mbox{tr}(\bold{P}_i\bold{P}_j)$.
Recall (ref). Set $\bold{A}_i=\bold{x}_i(\bold{x}_i'\bold{x}_i)^{-1}\bold{x}_i'$ for $1\leq i \leq N.$ Then $\bold{A}_i$ is a $T\times T$ idempotent matrix with rank $p$ and $\mbox{tr}(\bold{A}_i)=p$ for each $i$. Since $\bold{P}_i=\bold{I}_T-\bold{A}_i$, we see
where $\bold{B}_{ij}:=\bold{A}_i\bold{A}_j-\bold{A}_i-\bold{A}_j.$ By Lemma (ref),
for any non-negative definite matrices $\bold{F}_1$ and $\bold{F}_2$. As a result, $\mbox{tr}(\bold{A}_i\bold{A}_j) \geq 0$. Easily, $\mbox{tr}(\bold{A}_i\bold{A}_j)\leq p$ by Lemma (ref). Thus,
Therefore, we have $\mbox{tr}(\bold{P}_i\bold{P}_j) \geq T-2p$. Hence, $\mbox{tr}(\bold{M}_{i\blacktriangledown})\geq (N-i)(T-2p)$ by (ref). This and (ref) tell us that
by recalling the notation $m=T-p.$ Plugging this into (ref) we get
By the Chebyshev inequality, for any $\tau>0$,
which goes to zero provided $N=o(T^2).$ $\square$
Let $\{\bold{s}_1, \cdots, \bold{s}_j\}$ for $1\leq j \leq N$ be defined in Lemma (ref), which are i.i.d. random vectors uniformly distributed on $\mathbb{S}^{m-1}$. Set
which is the $\sigma$-algebra generated by $\{\bold{s}_1, \cdots, \bold{s}_j\}$ for $1\leq j \leq N$. Here $\Omega$ is the sample space on which random variables $\{\epsilon_{ij}\}$ are defined on.
\noindentProof of Lemma (ref). Set $\bold{H}_{ij}=\bold{U}_i'\bold{U}_j\bold{s}_j \bold{s}_j'\bold{U}_j'\bold{U}_i$ for $1\leq i, j \leq N$, where $\bold{U}_i$'s and $\bold{s}_j$'s are defined as in Lemma (ref). Then $\bold{s}_i'\bold{H}_{ij}\bold{s}_i=\bold{s}_j'\bold{C}_{ij}\bold{s}_j$, where
Since $\bold{s}_i'\bold{U}_i'\bold{U}_j\bold{s}_j=(\bold{s}_i'\bold{U}_i'\bold{U}_j\bold{s}_j)' =\bold{s}_j'\bold{U}_j'\bold{U}_i\bold{s}_i\in \mathbb{R}$, we have
By Lemma (ref)(i) and the independence between $\bold{s}_i$ and $\bold{s}_j$, we have that
for $i<j$, where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. Then
for $2\leq j \leq N$, where
In view of the independence among $\bold{s}_i$'s, it is easy to check from Lemma (ref) that
Since $\mbox{tr}(\bold{C}_{ij})=\mbox{tr}(\bold{U}_j'\bold{U}_i\bold{s}_i \bold{s}_i'\bold{U}_i'\bold{U}_j)=\bold{s}_i'\bold{M}_{ij}\bold{s}_i$, we have that
This, (ref) and Lemma (ref) imply
From (ref),
by the definition of $\bold{C}_{ij}$. Thus, we conclude from (ref) that
It follows that
Review $T=m+p$. Since ${\rm Var}(\xi_1+\xi_2)\leq 2{\rm Var}(\xi_1)+2{\rm Var}(\xi_2)$ for any random variables $\xi_1$ and $\xi_2$, to show ${\rm Var}(Z_N)\rightarrow 0$, it is enough to prove the following two facts.
Under restriction $N=o(T^3)$, the assertion (ref) is confirmed in Lemma (ref) and (ref) is proved in Lemma (ref). The proof is completed. $\square$
\noindentProof of Lemma (ref). It suffices to show
as $N\to\infty.$ By (ref) and (ref),
for $2\leq j \leq N$, where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$ and
Notice
where
for $1\leq i<j\leq N.$ By Lemma (ref),
for any $1\leq i<j\leq N.$ We rewrite (ref) to have
Therefore,
Note that $A_j$ is the sum of independent random variables. By (ref) with $\tau=4$,
where the second inequality follows from Lemma (ref), and where the fact that $\mbox{tr}(\bold{H}_{ij}^2)\geq \frac{1}{m}[\mbox{tr}(\bold{H}_{ij})]^2$ from Lemma (ref) is used in the last step. Easily, $\mbox{tr}(\bold{H}_{ij}^2)=(\bold{s}_j'\bold{M}_{ji}\bold{s}_j)^2.$ Take another expectation to see
By Lemma (ref) with $\tau=4$ and the fact that $\mbox{tr}(\bold{M}^2)\geq \frac{1}{m}[\mbox{tr}(\bold{M})]^2$ for any $m\times m$ symmetric matrix $\bold{M}$ from Lemma (ref), we obtain
It is used before that $\mbox{tr}(\bold{M}_{ji})=\mbox{tr}(\bold{P}_i\bold{P}_j)$ and $\mbox{tr}[\bold{M}_{ji}^2]=\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]$. By Lemma (ref), both quantities are bounded by $m.$ Hence, $E\big[(\bold{s}'\bold{M}_{ji}\bold{s})^4\big]\leq C$ uniformly for all $1\leq i<j\leq N$. We conclude from (ref) that
uniformly for all $2\leq j \leq N$.
Now we estimate $B_j$. Replace “$\bold{H}_{ij}$" in (ref) with “$\bold{M}_{ij}$" to see
where the last step holds by (i) and (ii) from Lemma (ref).
Finally, by (ref),
where
Since $\mbox{tr}(\bold{M}_{ji})=\mbox{tr}(\bold{P}_i\bold{P}_j)$, by defining
we have $\mbox{tr}(\bold{M}_{j\blacktriangle})=\mbox{tr}(\bold{P}_{j\blacktriangle}\bold{P}_j)$. Recall (ref), $\bold{U}_{i}\bold{U}_{i}'=\bold{P}_i$. Easily,
for any $1\leq i, j, k\leq N$. It follows that
On the other hand, recall $m=T-p.$ By Lemma (ref)(v), there exists a constant $K$ not depending on $T$ or $N$ such that
for every $2\leq j \leq N.$ It follows from the triangle inequality that
for $2\leq j \leq N.$ Consequently, by taking $\tau=4$ in Lemma (ref) we have that
uniformly for all $2\leq j \leq N$. Combining this with (ref), (ref) and (ref), we arrive at
uniformly for all $2\leq j \leq N$ as $N$ is large (reviewing $m=T-p$ and $T=T_N\to \infty).$ As a result,
as $N\to \infty$ as long as $N=o(T^4).$ We obtain (ref). $\Box$
\noindentThe Proof of Lemma (ref). Reviewing Lemma (ref), we know
for $2\leq j\leq N.$ Let $\mathcal{F}_j$ be as in (ref). Next we will verify that, for each $N\geq 2$, $\{X_j;\, 2\leq j \leq N\}$ forms a sequence of martingale differences with respect to the $\sigma$-algebras $\{\mathcal{F}_{j};\, 1\le j\le N-1\}.$ Define $J_1=0$ and
for $2\leq j \leq N-1$. By Lemma (ref), $\hat{\rho}_{ij}$ depends on $\bold{s}_i$ and $\bold{s}_j$ only. From independence of $\{\bold{s}_1, \cdots, \bold{s}_N\}$ and Lemma (ref),
for $2\leq j \leq N,$ where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. Therefore,
forms a martingale difference with respect to the $\sigma$-algebras $\{\mathcal{F}_{j}; 2\le j\le N\}.$
Now, in order to prove
in distribution as $N\to\infty$, we will employ the Lindeberg-Feller central limit theorem (see, for example, p. 476 from Billingsley95 or p. 344 from Durrett). To achieve so, it is enough to verify that
in probability and
in probability as $N\to\infty.$ Lemma (ref) has showed (ref). Now, to prove (ref), it suffices to show
and
as $N\to\infty.$ Lemma (ref) proves (ref) under the assumption $N=o(T^2)$. The assertion (ref) is confirmed in Lemma (ref) by assuming $N=o(T^3)$. Inspect all restrictions between $N$ and $T$ in the lemmas used earlier, the condition $N=o(T^2)$ meets all requirement. The proof is then completed. $\square$
With the preparations in Sections in (ref)-(ref), we now are ready to prove the central limit theorem stated in Theorem (ref). The main idea is to write the sum of squares of sample correlation coefficients as sums of martingale differences. Then the Lindeberg-Feller martingale CLT is applied.
\noindentProof of Theorem (ref). Review $J_1=0$ and
for $2\leq j \leq N-1$. Then $S_N=\sum_{j=2}^NJ_j.$ Review $\mathcal{F}_0$ and $\mathcal{F}_j$ in (ref). By Lemma (ref), the conditional expectation,
for $2\leq j \leq N,$ where $\bold{M}_{ij}=\bold{U}_i'\bold{U}_j\bold{U}_j'\bold{U}_i$. As in (ref),
forms a martingale difference with respect to the $\sigma$-algebras $\{\mathcal{F}_{j}; 2\le j\le N\}.$ Therefore $\frac{1}{N}(S_N-\mu_N)$ can be further written by
From Lemma (ref),
in probability as $N\to\infty$. By Lemma (ref),
in distribution as $N\to\infty.$ The proof then follows from the Slutsky lemma. $\square$
\noindentProof of Theorem (ref). Set $m=T-p$. First,
It is easy to see
as $N\to\infty$. It follows that
By Lemma (ref)(i) and (ii), there exists a constant $K>0$ depending on $p$ but not on $N$, $T$ or $\bold{P}_i$'s such that
uniformly for all $1\leq i< j \leq N$ and $N\geq 4.$ Therefore, by the definition of $v_{Nij}$, we have
uniformly for all $1\leq i< j \leq N$ as $N\to\infty.$ Immediately,
uniformly for all $1\leq i< j \leq N$ as $N\to\infty.$ Now write $\frac{1}{v_{Nij}}=\frac{1}{\sqrt{2}}(1+\omega_{Nij})$. Then
as $N\geq 4.$ By Lemma (ref), $E\hat{\rho}_{ij}^2=m^{-1}\mu_{Nij}.$ It follows that
where $S_N$ is defined as in Theorem (ref). By the Slutsky lemma and Theorem (ref),
in distribution as $N\to \infty.$ Recall (ref). To prove $Q_N\to N(0, 1)$ in distribution, by the Slutsky lemma again, it is enough to show
in probability as $N\to \infty$. Since $\hat{\rho}_{ij}^2$ and $\hat{\rho}_{kl}^2$ are independent if $\{i, j\}\cap \{k, l\}=\emptyset$, then
where the last sum runs over all $(k, l)$ with $1\leq k<l\leq N$ and $\{i, j\}\cap \{k, l\}\ne \emptyset$. The total number of such $(k, l)$'s is no more than $2N+ 2N=4N.$ Since $|\mbox{Cov}(U, V)|\leq [\mbox{Var}(U)]^{1/2}\cdot [\mbox{Var}(V)]^{1/2}$ for any random variables $U$ and $V$, we have from (ref) that
By Lemma (ref), $\mbox{tr}(\bold{P}_i\bold{P}_j)\leq \mbox{tr}(\bold{P}_i)= m$ and $\mbox{tr}[(\bold{P}_i\bold{P}_j)^2]\leq \mbox{tr}(\bold{P}_i)= m$. By Lemma (ref)(iv),
Thus, $\mbox{Var}(\hat{\rho}_{ij}^2)\leq \frac{8}{m^2}$ for all $1\leq i<j \leq N.$ Combing this with (ref), we get
by the assumption $N=o(T^2).$ This implies (ref). The proof is completed.$\square$
Theorems (ref)-(ref) will be proved via approximating $L_N=\max_{1\le i<j\le N} |\hat{\rho}_{ij}|$ for any $p\geq 0$ by $L_N$ for the case $p=0$. The latter one has the asymptotic result known in CJF13. The main job is reduced to show the difference between the two versions of $L_N$ is small enough.
The results stated in this section will be proved in Section (ref).
It is easy to see that the bound in the lemma is tight by simply taking $a_1=1$ and $a_2=\cdots=a_m=0.$
The above inequality is tight, which can be seen by taking $a_1=1$ and $a_i=0$ for $2\leq i \leq m.$
Recall the definition of subgaussian random variables defined before the statement of Theorem (ref).
The upper bound in the lemma is optimized, which can be seen evidently by choosing $a_1=1$ and $a_2=\cdots=a_m=0$.
Recall the setting in (ref), (ref) and (ref). Let $p$ be fixed. Let $\bold{e}_i=\frac{\bold{\epsilon}_i}{\|\bold{\epsilon}_i\|}$ and $\tilde{\rho}_{ij}=\bold{e}_i'\bold{e}_j$ for $1\leq i, j \leq N$. In this section, we always assume that $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables. The “continuous" requirement guarantees that $\hat{\rho}_{ij}$ in (ref) is well-defined. See the comment below (ref).
\noindentProof of Lemma (ref). Notice $\bold{P}_i\bold{\epsilon}_i=\bold{\epsilon}_i-\bold{A}_i\bold{\epsilon}_i.$ It follows that
Take $i=j$ to see that
since $\bold{A}_i^2=\bold{A}_i$. In particular, $\|\bold{\epsilon}_i\|^2- \bold{\epsilon}_i'\bold{A}_i\bold{\epsilon}_i\geq 0$, hence
for each $i$. Combining the last two identities, we have
Dividing the numerator and denominator by $\|\bold{\epsilon}_i\|\cdot \|\bold{\epsilon}_j\|$, we have
Write
It is easy to see $|1-(1-x)^{-1/2}|\leq 2|x|$ if $|x|\leq \frac{1}{2}.$ Then $(1-x)^{-1/2}\leq 1+2|x|$ as $|x|\leq \frac{1}{2}.$ For brevity of notation, set $h_{ij}=\big(1-\bold{e}_i'\bold{A}_i\bold{e}_i\big)^{-1/2} \big(1-\bold{e}_j'\bold{A}_j\bold{e}_j\big)^{-1/2}.$ Then
provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$, and at the same time $h_{ij}\leq 2$ by definition. From (ref),
By the Cauchy-Schwartz inequality and the fact $\bold{A}_i^2=\bold{A}_i$,
Similarly, $|\bold{e}_i'\bold{A}_i\bold{e}_j|\leq \|\bold{A}_i\bold{e}_i\|\cdot\|\bold{A}_i\bold{e}_j\|$ and $|\bold{e}_i'\bold{A}_j\bold{e}_j| \leq \|\bold{A}_j\bold{e}_i\|\cdot\|\bold{A}_j\bold{e}_j\|$ since $\bold{A}_i^2=\bold{A}_i$. Consequently
by the fact $|\tilde{\rho}_{ij}| \leq 1$. Use the trivial fact that $2xy\leq x^2+y^2$ to see
provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$, where in the last inequality we use the fact that each term is bounded by $\max_{1\leq i \leq j \leq N}\bold{e}_j'\bold{A}_i\bold{e}_j$. We then have from (ref) that
provided $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i\leq \frac{1}{2}$. If $\max_{1\leq i \leq N}\bold{e}_i'\bold{A}_i\bold{e}_i> \frac{1}{2}$, (ref) holds automatically due to the facts that $|\hat{\rho}_{ij}|\leq 1$ and $|\tilde{\rho}_{ij}|\leq 1$ for all $i,j$. The proof is completed. $\Box$
\noindentProof of Proposition (ref). To prove the result, by the homogeneity of $\hat{\rho}_{ij}$ from (ref), without loss of generality, we assume $E(\bold{\epsilon}_{11}^2)=1$. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\alpha_N\to 0$ as $N\to\infty$ by assumption. From Lemma (ref), for any $h\in (0, 14)$,
Next we estimate the last probability.
For any $v>0$,
Now
Note that $\|\bold{\epsilon}_1\|^2=\sum_{j=1}^T\bold{\epsilon}_{1j}^2$. Since $E(\bold{\epsilon}_{11}^2)=1$ and $E(|\bold{\epsilon}_{11}|^\tau)<\infty$, we see
by (ref), where the Markov inequality is applied in the second inequality. Write $\bold{Q}=\Gamma'\bold{D}\Gamma$ where $\Gamma=(\gamma_{ij})_{T\times T}$ is an orthogonal matrix and
Then $\bold{\epsilon}_1'\bold{Q}\bold{\epsilon}_1=\sum_{k=1}^p\big(\sum_{j=1}^T\gamma_{kj}\epsilon_{1j}\big)^2$. It follows that
where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. Thus, from Lemma (ref) we have that there exists some $K>0$, such that
Join this with (ref), (ref) and (ref) to get
By taking $v=\frac{h}{28}$, we have from the above and (ref) that
which goes to zero by the assumption that $T/(N^{8/\tau}\log N)\to \infty$ as $N\to\infty$. $\Box$
\noindentProof of Proposition (ref). To prove the result, by the homogeneity of $\hat{\rho}_{ij}$ from (ref), without loss of generality, assume $E(\bold{\epsilon}_{11}^2)=1$. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\alpha_N\to 0$ as $N\to\infty$ by assumption. From (ref), we have that, for any $h\in (0, 14)$,
as $N$ is sufficiently large. So to finish the proof it suffices to show the second probability goes to zero.
For any $v>0$, by (ref) and (ref),
Since $E(\bold{\epsilon}_{11}^2)=1$, by large deviations, there exists a constant $\eta_0>0$ such that
for large enough $T$; see, for example, DZ98. By (ref),
where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. From Lemma (ref), there exists $K>0$ such that
as $T$ is large enough, where $v':=(v/p)^{1/2}$ and $K>0$ is a constant not depending on $p$, $N$, $T$ or $\gamma_{kj}$'s. It is easy to see the above goes to zero if $T/(\log N)^5\to \infty.$ It follows that
provided $T/(\log N)^5\to \infty.$ The proof is completed. $\Box$
\noindentProof of Proposition (ref). First, the subgaussian assumption implies that $Ee^{\omega|\epsilon_{11}|^2}<\infty$ for some $\omega>0$. Hence $Ee^{t|\epsilon_{11}|}<\infty$ for all $t>0$. We will use the same notation as in the proof of Proposition (ref). Reviewing (ref) and (ref), to get our desired result, it suffices to show that $N^2\cdot \max_{1\leq i \leq N}P\big(\bold{e}_1'\bold{A}_i\bold{e}_1> 2\alpha_N v\big)\to 0$ for any $v>0$.
By (ref), for each $i$,
where $v':=(v/p)^{1/2}.$ Note that $\sum_{j=1}^T\gamma_{kj}^2=1$ for each $1\leq k \leq p$ by orthogonality. From Lemma (ref), there exists $K>0$ such that
Therefore, by (ref) and (ref), there exists a constant $\beta_0>0$ such that
It is easy to see the above goes to zero if $\log N=o(T^{1/3}).$ $\Box$
Assume $\{\epsilon_{ij};\, i\geq 1, j \geq 1 \}$ are i.i.d. continuous random variables. Set $\bar{\bold{\epsilon}}_i=(1/T)\sum_{j=1}^T\epsilon_{ij}$ for all $i$. Define $\bold{1}=(1, \dots, 1)'\in \mathbb{R}^T.$ The Pearson correlation coefficient $\rho_{ij}$ is then defined by
for $1\leq i, j\leq N.$ Similar to the clarification below (ref), the “i.i.d. continuous" assumption justifies that $\rho_{ij}$ is well-defined.
\noindentProof of Proposition (ref). The proof consists of two steps. In the first step we will show
By using this we will prove the desired conclusion in the second step.
{\it Step 1}. Set $\delta_i=\sqrt{T}\cdot\frac{\bar{\bold{\epsilon}}_i}{\|\bold{\epsilon}_i\|}$ for each $i.$ Recall $\tilde{\rho}_{ij}=\frac{\bold{\epsilon}_i'\bold{\epsilon}_j} {\|\bold{\epsilon}_i\|\cdot\|\bold{\epsilon}_j\|}$. Write
It follows that
By the inequality $(1-x)^{-1/2}\leq 1+2|x|$ for $|x|\leq \frac{1}{2}$ appeared in the proof of Lemma (ref), we have
provided $\max_{1\leq i \leq N} |\delta_i|<\frac{1}{2}.$ Under this restriction, $(1-\delta_i^2)^{-1/2}\leq \frac{2}{\sqrt{3}} \leq 2$ for each $i$. Therefore,
Since $|\tilde{\rho}_{ij}|\leq 1$, the above two estimates joining with (ref) implies that
as $\max_{1\leq i \leq N} |\delta_i| \leq \frac{1}{2}.$ Moreover, the above naturally holds if $\max_{1\leq i \leq N} |\delta_i| > \frac{1}{2}.$ This leads to (ref).
{\it Step 2}. Set $\alpha_N=1/\sqrt{T\log N}$. Then $\lim_{N\to\infty}\alpha_N= 0$. From {\it Step 1}, for any $t>0$,
Therefore, to show $\sqrt{T\log N}\cdot\max_{1\leq i< j \leq N}|\rho_{ij}-\tilde{\rho}_{ij}| \to 0$, it is enough to prove that
for any $s>0.$ In fact,
where $\{\xi_j;\, 1\leq j \leq T\}$ are i.i.d. random variables with the same distribution of $\bold{\epsilon}_{11}$. The reason we switch the notations from $\{\bold{\epsilon}_{ij}\}$'s to $\{\xi_j;\, 1\leq j \leq T\}$ is for the brevity of symbols. By (ref),
By the Markov inequality and (ref) as used in (ref),
since $\alpha_N=1/\sqrt{T\log N}$. Combing the above assertions, we arrive at
which converges to zero provided $\frac{N(\log N)^{\alpha/4}}{T^{\alpha/4}}\to 0$, or equivalently, $\frac{N^{4/\alpha}\log N}{T}\to 0$ $\Box$
With preparations earlier, we are now ready to prove the main theorems on the maximum statistics of sample correlation coefficients.
\noindentProof of Theorem (ref). Under the condition $E|\epsilon_{11}|^{6}<\infty$, Jiang2004 and Zhou2007 show that
converges weakly to a distribution with distribution function $F(y)$, where $L_N'=\max_{1\leq i< j \leq N}|\rho_{ij}|$ and $\rho_{ij}$ is as in (ref). Set $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). Observe that
Since $E|\epsilon_{11}|^{\tau}<\infty$ with $\tau>8$, by using the assumption $T/N\to c\in (0, \infty)$, we see that $\lim_{N\to\infty}T/(N^{8/\tau}\log N)=\infty$. Hence, by the Proposition (ref)
in probability as $N\to\infty$. On the other hand, since $T/N\to c\in (0, \infty)$ and $\tau > 8$ we have $\frac{N^{4/\alpha}\log N}{T}\to 0$. Then, by Proposition (ref),
in probability as $N\to\infty$. From (ref) we see that
in probability. Set $\Delta=L_N-L_N'$. Then
The Slutsky lemma and (ref) say that $(T/\log N)^{1/2}L_N'\to 2$ in probability. Consequently,
in probability by (ref). These together with (ref) conclude that
By the Slutsky lemma again, this fact and (ref) imply the desired result. $\Box$
\noindentProof of Theorem (ref). By assumption, $Ee^{\omega|\epsilon_{11}|}<\infty$ and $\log N=o(T^{1/5})$. Using Theorem 3 and Remark 2.1 from Cai_J2011 with $``\mu=0"$ and $``\alpha=1"$, we get
converges weakly to a distribution with distribution function $F(y)$, where $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). By Proposition (ref),
in probability as $N\to\infty.$ Recall $L_N=\max_{1\leq i< j \leq N}|\hat{\rho}_{ij}|$. By the triangle inequality, the above says that
in probability as $N\to\infty.$ Repeating the argument from (ref) to (ref), we obtain
as $N\to\infty.$ The conclusion follows from (ref). $\Box$
\noindentProof of Theorem (ref). By assumption, $\log N=o(T^{1/3})$. Taking $``\mu=0"$ and $``\alpha=2"$ in Theorem 3 and Remark 2.1 from Cai_J2011, we have
converges weakly to distribution function $F(y)$ for $y \in \mathbb{R}$, where $L_N''=\max_{1\leq i< j \leq N}|\tilde{\rho}_{ij}|$ and $\tilde{\rho}_{ij}$ is as in Lemma (ref). Recall $L_N=\max_{1\leq i< j \leq N}|\hat{\rho}_{ij}|$. By Proposition (ref), under the restriction $\log N=o(T^{1/3})$,
in probability as $N\to\infty.$ By the triangle inequality, the above says that
in probability as $N\to\infty.$ From the argument between (ref) and (ref), we have
as $N\to\infty.$ This and (ref) yield the conclusion. $\Box$
We create a new method to prove Theorem (ref) which gives the asymptotic independence between the sum $S_N$ and the maximum $L_N$. The idea is employing the inclusion-exclusion formula twice. We expect this method to work for other problems regarding asymptotic independence between sums of and maxima of weakly dependent random variables.
The results stated in this section are about the estimates of probabilities of events related to Gaussian random variables. They are useful in their own right. Their proofs will be presented in Section (ref).
After collecting some useful facts in Section (ref), we are now ready to prove Theorem (ref). To make the discussion easier to follow, we give the outline first.
First, Let $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Review the framework between (ref) and (ref). In particular,
for $1\leq i, j \leq N.$ Assume (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$. Then
are i.i.d. uniformly distributed over the $T$-dimensional unit sphere $\mathbb{S}^{T-1}.$ For fixed $y\in \mathbb{R}$, set
Here is the structure of the proof of Theorem (ref).
1. Let $\bold{e}_i$ be as in (ref). Define $\tilde{L}_N=\max_{1\leq i<j \leq N}|\bold{e}_i'\bold{e}_j|$. To show that $TL_N^2-4\log N +\log\log N$ and $(S_N-\mu_N)/N$ are asymptotically independent, it is enough to prove that $T\tilde{L}_N^2-4\log N +\log\log N$ and $(S_N-\mu_N)/N$ are asymptotically independent (Lemma (ref)). The benefit of this step is that $\{\bold{e}_i'\bold{e}_j;\ 1\leq i<j\leq N\}$ are identically distributed. This is not true for $\{\hat{\rho}_{ij};\, 1\leq i<j \leq N\}$ appeared in definition of $L_N$.
2. Review (ref). To show the asymptotic independence, it suffices to prove
for any real numbers $x$ and $y$, where $\Phi(x)$ is the cdf of $N(0, 1)$ and is also the limiting distribution function of $\frac{1}{N}(S_N-\mu_N)$; $F(y)$ is the Gumbel distribution and is also the limiting distribution function of $\tilde{L}_N$. Recall the definition of $\tilde{L}_N$, we are able to write the event in (ref) as the union of $\binom{N}{2}$ many events which are exchangeable. Then, by using the inclusion-exclusion formula, the probability in (ref) is sandwiched between two bounds [(ref) and (ref)]. The advantage is that we reduce the probability on the global maximum “$\tilde{L}_N$" to sums of probabilities on “local maxima".
3. In dealing with the “local maxima", each probability in the sum is of the form $P(\frac{1}{N}(S_N-\mu_N)\leq x, |\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$, where $n$ is a fixed number free of $N$ and $T$, and where the indices $\{(i_l, j_l);\, 1\leq l \leq n\}$ are different. Review $S_N$ is the sum of $(\bold{e}_i'\bold{e}_j)^2$ over all $1\leq i<j \leq N$. Remove the terms related to $\{\bold{e}_{i_l}'\bold{e}_{j_l};\, 1\leq l \leq n\}$ from $S_N$, in other words, eliminate the terms $(\bold{e}_i'\bold{e}_j)^2$ for all $(i, j)$ with $\{i, j\}\cap \{i_l, j_l\} \ne \emptyset$ for some $1\leq l\leq n$. Then the resulting sum is independent of $\{\bold{e}_{i_l}'\bold{e}_{j_l};\, 1\leq l \leq n\}$, and hence $P(\frac{1}{N}(S_N-\mu_N)\leq x, |\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ is asymptotically the product of $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$. Of course we have to handle the “loss" after removing the terms. It turns out that the removed terms are very concentrated at their mean values by the second and third conclusions from Lemma (ref). So the probability $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and the modified version $P(\frac{1}{N}(\tilde{S}_N-\tilde{\mu}_N)\leq x)$ are asymptotically equal. The total errors in the above approximations is negligible (Lemma (ref)).
4. In step 3, we have showed that
is asymptotically the product of $P(\frac{1}{N}(S_N-\mu_N)\leq x)$ and $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ in (ref) and (ref), where $A_N=\{\frac{1}{N}(S_N-\mu_N)\leq x\}$ and $B_{I}=\{|\bold{e}_i'\bold{e}_j|>l_N\}$ for $I=(i, j)$. We will use one more time the inclusion-exclusion formula to regroup the sum of probabilities $P(|\bold{e}_{i_1}'\bold{e}_{j_1}|>l_N, \cdots, |\bold{e}_{i_n}'\bold{e}_{j_n}|>l_N )$ and change it to $P(\max_{1\leq i<j \leq N}|\bold{e}_{i}'\bold{e}_{j}|>l_N)$. Note that the original upper bound becomes the lower bound of $P(\max_{1\leq i<j \leq N}|\bold{e}_{i}'\bold{e}_{j}|>l_N)$ and similarly the original lower bound becomes the new upper-bound. There are some “middle" terms in between the bounds, we have to show they are negligible. This is guaranteed by Lemma (ref).
Now let us execute the steps streamlined above.
\noindentProof of Lemma (ref). Let $m=T-p.$ Under assumption (ref) with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$, we know $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}.$ Define $\tilde{\rho}_{ij}=\bold{e}_i'\bold{e}_j$ for $1\leq i<j \leq N.$ To organize the proof clearly, we list the relevant quantities as follows.
where $\bold{P}_i$ is defined in (ref). By Theorems (ref) and (ref), the following hold.
By Theorem 6 from CJF13, the assertion (ref) is also true if “$L_N$" is replaced by “$\tilde{L}_N$". To show asymptotic independence, it is enough to show
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$, where $\Phi(x)=(2\pi)^{-1/2}\int_{-\infty}^xe^{-t^2/2}\,dt.$ Let $l_N$ be as in (ref). Due to (ref) and (ref) we know (ref) is equivalent to that
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. By assumption, we know that
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. We show next that (ref) implies (ref).
By Proposition (ref),
in probability as $N\to\infty$ provided $\log N=o(T^{1/3})$ as $N\to\infty$. By the triangle inequality,
in probability.
Given $\epsilon \in (0,1)$. Set
for $N\geq 3.$ Then
Now,
On $\Omega_N$, if $L_N>l_N$ then
Define
which makes sense for large $N.$ Use the formula $\sqrt{x}-\sqrt{y}=(x-y)/(\sqrt{x}+\sqrt{y})$ for any $x\geq 0$ and $y\geq 0$ to see
as $N \to\infty.$ Thus,
as $N$ is sufficiently large. This and (ref) conclude that
as $N$ is sufficiently large. Review (ref). We have
Immediately from (ref) and (ref) we get
for any $\epsilon \in (0,1).$ Inspect that the left-hand side of the above does not depend on $\epsilon$. Letting $\epsilon \downarrow 0$, we obtain
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}.$ In the following we will show the lower limit.
Evidently,
Set
Similar to (ref), it is checked that
as $N\to\infty$. Therefore,
as $N$ is sufficiently large. It is straightforward to verify that
as $N$ is sufficiently large, where the last inclusion follows from the definition of $\Omega_N$. By (ref),
Thus, from (ref) and (ref) we get
for any $\epsilon \in (0,1).$ Sending $\epsilon \downarrow 0$ we see
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}.$ This together with (ref) concludes (ref). $\Box$
We need some notations now. Let $S_N$, $L_N$ and $\mu_N$ be as in (ref), (ref) and (ref), respectively. Let $\{\bold{e}_i;\ 1\leq i \leq N\}$ be as in (ref). Define
for any $I=(i,j)\in \Lambda_N$. To make a clear presentation, we impose a trivial ordering for elements in $\Lambda_N$. For any $I_1=(i_1, j_1)\in \Lambda_N$ and $I_2=(i_2, j_2) \in \Lambda_N$, we say $I_1<I_2$ if $i_1<i_2$ or $i_1=i_2$ but $j_1<j_2$.
\noindentProof of Lemma (ref). For $I_l$ appeared in $H(N, n)$, write $I_l=(i_l, j_l)$ for $l=1, \cdots, n.$ Now we classify the indices $I_1< I_2< \cdots < I_{n}\in \Lambda_N$ in the definition of $H(N, n)$ into three cases. Let $\Gamma_{N,1}$ be the set of indices $(I_1, \cdots, I_n)$ such that no two of the $2n$ indices $\{i_l, j_l\,; l=1, \cdots, n\}$ are identical. Let $\Gamma_{N,2}$ be the set of indices $(I_1, \cdots, I_n)$ such that either $i_1=\cdots =i_n$ or $j_1=\cdots = j_n.$ Let $\Gamma_{N,3}$ be the set of indices $I_1< I_2< \cdots < I_{n}\in \Lambda_N$ excluding $\Gamma_{N,1}\cup \Gamma_{N,2}$. In the following we will estimate
for $j=1, 2, 3$ one by one. We will see $F_1$ contributes essentially the sum in the expression of $H(N, n)$ by an easy argument; the term $F_2$ is negligible and its computation is trivial; the term $F_3$ is also negligible but its estimate is most involved.
{\it Step 1: the estimate of $F_1$}. Recall $B_{I}=\{|\bold{e}_i'\bold{e}_j|>l_N\}$ if $I=(i,j)\in \Lambda_N$, where $l_N$ is defined in (ref). By the definition of $\Gamma_{N,1}$, we know that $B_{I_1}, B_{I_2} \cdots, B_{I_n}$ are independent. By Lemma (ref) and the symmetry of $\bold{e}_1'\bold{e}_2$,
for all $N\geq 3$. Then, by the elementary fact $\binom{k}{n}=\frac{1}{n!}k(k-1)\cdots (k-n+1)\leq \frac{k^n}{n!}$ for all $k> n\geq 1.$
{\it Step 2: the estimate of $F_2$}. Evidently, the size of $\Gamma_{N,2}$ is no more than $\binom{N}{1}\cdot\binom{N}{n}\cdot 2\leq 2N^{n+1}$. We first claim that $\{\bold{e}_1'\bold{e}_2, \bold{e}_1'\bold{e}_3, \cdots, \bold{e}_1'\bold{e}_n\}$ are independent. In fact, let $\bold{e}$ be uniformly distributed on $\mathbb{S}^{T-1}$. Then, $\bold{a}'\bold{e}$ has the same distribution as that of $(1, 0, \cdots, 0)'\bold{e}$ for any $a\in \mathbb{S}^{T-1}$ (see, e.g., Theorem 1.5.7(i) and the argument for (5) on p.147 from mui). Since $\bold{e}_1, \cdots, \bold{e}_n$ are i.i.d. random vectors, we know that, conditioning on $\bold{e}_1$, the random variables $\{\bold{e}_1'\bold{e}_2, \bold{e}_1'\bold{e}_3, \cdots, \bold{e}_1'\bold{e}_n\}$ are i.i.d. with a common distribution of $(1, 0, \cdots, 0)'\bold{e}$. In particular, their conditional distributions do not depend on $\bold{e}_1$. This proves the claim. Consequently,
by (ref).
{\it Step 3: the estimate of $F_3$}. Fix a tuple $(I_1, I_2, \cdots , I_{n})\in \Gamma_{N,3}$. By the ordering imposed on $\Lambda_N$, we see that $i_1\leq i_2\leq \cdots \leq i_n$. There are two different cases: (1) $i_1<i_2$; (2) there exists $2\leq k\leq n-1$ such that $i_1=\cdots =i_k<i_{k+1}$.
Under case (1), let $\mathcal{F}_1$ be the set of random vectors $\{\bold{e}_{j_1}, \bold{e}_{i_l}, \bold{e}_{j_l};\, 2\leq l\leq n\}$ (the first index is “$j_1$" which is different from the third one “$j_l$"). Then, by independence and the property “take out what is known" for the conditional probability,
As a fact used earlier, the conditional distribution of $\bold{e}_{i_1}'\bold{e}_{j_1}$ given $\bold{e}_{j_1}$ and the unconditional distribution of $\bold{e}_{i_1}'\bold{e}_{j_1}$ are identical. Therefore, by (ref),
Let us study case (2). Without loss of generality, for notational clarity, we assume $i_1=\cdots =i_k=1$ and $i_{k+1}=2$. Denote by $\mathcal{F}_2$ the set of random vectors $\{\bold{e}_{i_l}, \bold{e}_{j_l};\, 1\leq l\leq n\}$ excluding $\bold{e}_1$. Then use conditional probability and independence to see
where $P_1$ stands for the condition probability given $\mathcal{F}_2$. By independence, the last probability in (ref) is computed by treating $\bold{e}_1$ as a random variable while fixing the values of $\bold{e}_{j_1}, \cdots, \bold{e}_{j_{k}}$. To study the $P_1$, we need to understand the relationship among $\{\bold{e}_{j_1}, \cdots, \bold{e}_{j_{k}}\}.$ To do so, set
By Lemma (ref) and the fact $k\leq n$,
as $N$ is sufficiently large provided $\log N=o(T)$. Notice that
We claim that, for any $\epsilon\in (0,1)$, there exists an integer $N_{\epsilon}\geq 1$ such that
as $N\geq N_{\epsilon}$. On $\Omega_N$, we know $\max_{j_1\leq l_1<l_2\leq j_{k}}|\bold{e}_{j_{l_1}}'\bold{e}_{j_{l_2}}|< \delta_N.$ Take $r=1-\frac{\epsilon}{4k}$, $y=(\log N)^{1/4}$, $z=l_N$, $\delta=\delta_N$ in Lemma (ref). Observe that $2^k/y\to 0$. By (ref), $l_N\sim 2\sqrt{(\log N)/T}$ as $N\to\infty$, and hence $y=o(z\sqrt{rT})$. Also, $\delta_N\to 0$ since $\log N=o(T)$.
Thus, by the lemma, use the facts that $2rk>2k-\frac{3}{4}\epsilon$ and that $z\sqrt{rT}-y\to \infty$ to get
as $N\geq N_{\epsilon}$ thanks to the assumption $\log N=o(\sqrt{T})$, where $N_{\epsilon}\geq 1$ is an integer depending on $\epsilon$ only. This leads to (ref).
Now, combining (ref) and (ref), we arrive at
as $N\geq N_{\epsilon}$. This together with (ref) and (ref) implies
as $N$ is sufficiently large. In summary, by using the above conclusion and (ref), for any $\epsilon\in (0, 1)$ and any $(I_1, \cdots, I_n)\in \Gamma_{n,3}$,
as $N\geq N_{\epsilon}$, where $k_1$ is the number of elements on the $i_1$-th row of $\{I_1, \cdots, I_n\}$. In words, when we consider $P(B_{I_1}B_{I_2}\cdots B_{I_{n}})$ based on the positions of $I_j$'s appeared in the upper triangular matrix $\Lambda_N=\{(i, j);\, 1\leq i<j \leq N\}$, after reducing the first row we see the connection between the old and new probabilities. Similarly, let $k_j$ be the number of elements from $\{I_1, \cdots, I_n\}$ on the $j$-th row for $j\geq 1.$ Then
where the two intersections above run over all elements from $\{I_1, \cdots, I_n\}$ excluding the first two rows. Continue the process recursively to see
where $b$ is the total number of rows of $\{I_1, \cdots, I_n\}$ in the upper triangular matrix $\Lambda_N=\{(i, j);\, 1\leq i<j \leq N\}$. Obviously, $k_1+\cdots +k_b=n$ and $b\leq n.$ Therefore, for each $\epsilon \in (0, 1)$,
for $N\geq N_{\epsilon}$. This gives that
Recall $I_l=(i_l, j_l)$ for each $1\leq l \leq n.$ In view of the definition of $\Gamma_{N,3}$, there are at least two of the $2n$ indices from $\{(i_l, j_l);\, 1\leq l \leq n\}$ are identical for any $(I_1, \cdots, I_n)\in \Gamma_{N,3}$. Let $\kappa=|\{i_l, j_l; 1\leq l \leq n\}|$ for such $(I_1, I_2, \cdots , I_{n})$. Easily, $n+1\leq \kappa\leq 2n-1$. To see how many such $(I_1, \cdots, I_n)$ with $|\{i_l, j_l; 1\leq l \leq N\}|=\kappa$, first pick $\kappa$ many indices from $\{1, 2, \cdots, N\}$, which has the total number of ways $\binom{N}{\kappa}\leq N^{\kappa}$, then use the $\kappa$ many indices to make a $(I_1, \cdots, I_n)\in \Gamma_{N,3}$. The total number of ways to do so is no more than $\kappa^{2n}$. Therefore,
As a consequence, for each $\epsilon \in (0, 1)$, from (ref) we have
as $N\geq N_{\epsilon}$. Take $\epsilon=\frac{1}{2n}$ to see $\lim_{N\to\infty}F_3= 0.$ Joining this with (ref) and (ref), we eventually arrive at
for each $n\geq 3.$ The desired conclusion then follows by sending $n\to \infty.$ $\Box$
\noindentProof of Lemma (ref). From assumption that (ref) holds with $\bold{\epsilon}_{i} \sim N_T(\bold{0}, \sigma_i^2\bold{I})$ for each $i$, we know from (ref) that $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}$. For $I_1< I_2< \cdots < I_{n}\in \Lambda_N$, write $I_l=(i_l, j_l)$ for $l=1,2,\cdots, n$. Set
for $n\geq 1.$ It is easy to check that $|\Lambda_{n,N}| =\sum_{l=1}^n(N-i_l+j_l-2)$. Since $i_l<j_l$ for each $l$, we see that
Recall $m=T-p$ and
where $\bold{P}_i$ is defined as in (ref). Define
for $N\geq 3$ and
for $N\geq n\geq 1.$ Observe that $B_{I_1}B_{I_2}\cdots B_{I_{n}}$ is an event generated by random vectors $\{\bold{e}_i, \bold{e}_j ;\, (i, j) \in \Lambda_{n,N}\}$. A crucial observation is that $S_N-S_{N,n}$ is independent of $B_{I_1}B_{I_2}\cdots B_{I_{n}}$. It is easy to see that
For any integer $\tau\geq 2$, from a convex inequality we have
by Lemma (ref), where the constant $C$ is free of $N$ and $T$, and where the last step follows from the assumption $N=o(T^2)$. Similarly,
Lastly, by Lemma (ref) again,
Therefore,
Fix $\epsilon\in (0,1).$ By the Markov inequality,
for all $N\geq n^2$, where $C'$ is a constant depending on $\epsilon$ but free of $N$, $T$ or indices $\{I_1, \cdots, I_n\}.$
Fix $I_1< I_2< \cdots < I_{n}\in \Lambda_N$. By (ref) and the definition of $A_N(x)$,
by the independence between $S_N-S_{N,n}$ and $B_{I_1}B_{I_2}\cdots B_{I_{n}}$. Now
Combing the two inequalities to get
Similarly,
In other words, by independence,
Furthermore,
The above two strings of inequalities imply
which joining with (ref) yields
where
In particular,
as $N\to\infty$ by Theorem (ref). As a consequence,
where
as defined in Lemma (ref). From (ref), we know $\limsup_{N\to\infty}H(N, n)\leq C/n!$, where $C$ is a universal constant. Picking $\tau=6n$, and using the trivial fact $\binom{r}{s}\leq r^s$ for any integers $1\leq s \leq r$, we have that
Hence, from (ref)
for any $\epsilon>0$. The desired result follows by sending $\epsilon \downarrow 0.$ $\square$
We now are ready to assemble everything together.
\noindentProof of Theorem (ref). Recall $\{\bold{e}_i;\ 1\leq i \leq N\}$ in (ref). By assumption (ref), we see that $\{\bold{e}_i;\ 1\leq i \leq N\}$ are i.i.d. uniformly distributed over $\mathbb{S}^{T-1}$. As in Lemma (ref), define
Let $m=T-p.$ Recall
By Theorem 3 and Remark 2.1 from Cai_J2011 and Theorem (ref), the following hold.
To show asymptotic independence, by Lemma (ref), it is enough to show
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$, where $\Phi(x)=(2\pi)^{-1/2}\int_{-\infty}^xe^{-t^2/2}\,dt.$ Review (ref) to see
which makes sense for large $N$. Because of (ref) and (ref), the above is equivalent to that
for any $x\in \mathbb{R}$ and $y \in \mathbb{R}$. Review notations $ \Lambda_N$, $A_N$ and $B_{I}$ for any $I=(i,j)\in \Lambda_N$ in (ref). Write
Here the notation $A_NB_I$ stands for $A_N\cap B_I$. From the inclusion-exclusion principle,
and
for any integer $k\geq 1$. Reviewing the definition
for $n\geq 1$ in Lemma (ref), we have from the lemma that
Set
for $n\geq 1.$ By Lemma (ref),
for each $n\geq 1$. The assertion (ref) implies that
where the inclusion-exclusion formula is used again in the last inequality, that is,
for all $k\geq 1$. By the definition of $l_N$ and (ref),
as $N\to\infty$. By (ref), $P(A_N)\to \Phi(x)$ as $N\to\infty.$ From (ref), by fixing $k$ first and sending $N\to \infty$ we get from (ref) that
Now, let $k\to \infty$ and use (ref) to see
By applying the same argument to (ref), we see that the counterpart of (ref) becomes
where in the last step we use the inclusion-exclusion principle such that
for all $k\geq 1$. Review (ref) and repeat the earlier procedure to see
by sending $N\to \infty$ and then sending $k\to\infty.$ This and (ref) yield (ref). The proof is completed. $\Box$
Professors Feng and Liu thank NSFC grants 11501092 and 11571068 for partially support. Professor Jiang thanks NSF Grants DMS-1406279 and DMS-1916014 for partially support.