Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,994 characters · 16 sections · 42 citation commands
Estimation of Cross--Sectional Dependence in Large Panels
{
Large panel data analysis attracts ever-growing interest in the modern literature of economics and finance. Cross--sectional dependence is popular in large panel data and the relevant literature focuses on testing existence of cross-sectional dependence. A survey on description and testing of cross-sectional dependence is given in SW2012. P2004 utilizes sample correlations to test cross-sectional dependence while BFK2012 extend the classical Lagrangian multiplier (LM) test to the large dimensional case. CGL2012 and HPP2012 consider cross-sectional dependence tests for nonlinear econometric models. As more and more cross-sections are grouped together in panel data, it is quite natural and common for cross-sectional dependence to appear. Cross-sectional independence is an extreme hypothesis. Rejecting such a hypothesis does not provide much information about the relationship between different cross-sections under study. In view of this, measuring the degree of cross-sectional dependence is more important than just testing its presence. As we know, in comparison with cross-sectional dependence tests, few studies contribute to accessing the extent of cross-sectional dependence. Ng2006 uses spacings of cross-sectional correlations to exploit the ratio of correlated subsets over all sections. BKP2016 use a factor model to describe cross-sectional dependence and develop estimators that are based on a method of moments.
In this paper, we will contribute to this area: description and measure of the extent of cross-sectional dependence for large dimensional panel data with $N$ cross-section units and $T$ time series. The first natural question is how to describe cross-sectional dependence in panel data efficiently? To address this issue, the panel data literature mainly discusses two different ways of modelling cross-sectional dependence: the spatial correlation and the factor structure approach (see, for example, SW2012). This paper utilizes the factor model to describe cross-sectional dependence as well as capturing time serial dependence, which can benefit further statistical inference such like forecasting. Actually, the factor model is not only a powerful tool to characterize cross-sectional dependence for economic and financial data, but also efficient in dealing with statistical inference for high dimensional data from a dimension-reduction point of view. Some related studies include FFL2008, FLM2013 and PY2008.
While it is a rare phenomenon to have cross-sectional independence for all $N$ sections, it is also unrealistic to assume that all $N$ sections are dependent. Hence the degree or extent of cross-sectional dependence is more significant in statistical inference for panel data. BN2006 illustrate that more data usually result in worse forecasting due to heterogeneity in the presence of cross-sectional dependence. Grouping strong-correlated cross-sections together is very significant in further study. In factor model, relation among cross-sections is described by common factors and the strength of this relation is reflected via factor loading for each cross-section. Larger factor loading for one cross-section means stronger relation of this cross-section with common factors. In this paper, we suppose that some factor loadings are bounded away from zero while others are around zero. In detail, it is assumed that only $[N^{\alpha_0}] (0\leq\alpha_0\leq 1)$ of all $N$ factor loadings are individually important. Instead of measuring the extent by $\alpha_0N$, we adopt the parametrization $[N^{\alpha_0}]$. The proportion of $[N^{\alpha_0}]$ over the total $N$ is quite small which tends to $0$ as $0<\alpha_0<1$, while $\alpha_0N$ is comparable to $N$ because of the same order. In this sense, our model covers some “sparse" cases that only a small part of the sections are cross-sectionally dependent.
With such parametrization of extent for cross-sectional dependence, one goal is to propose an estimation method for the parameter $\alpha_0$. This paper proposes a unified estimation method which incorporates two classical types of cross-sectional dependence: static and dynamic principal components. In fact, factor model is equivalent to principal component analysis (PCA) in some sense (see FLM2013). Static PCA provides the common factor with most variation while dynamic PCA finds the common factor with largest “aggregated" time-serial covariances. The existing literature, including Baing2002, FFL2008, FLM2013, focuses on common factors from static PCA. In high dimensional time series, researchers prefer using dynamic PCA to derive common factors that can keep time-serial dependence. This is very important in high dimensional time series forecasting, e.g. LY2012.
In this paper, for our panel data $x_{it}, i=1, 2, \ldots, N; t=1, 2, \ldots, T$, an estimator for $\alpha_0$ is proposed based on the criterion of covariance between $\bar x_t$ and $\bar x_{t+\tau}$ for $0\leq\alpha_0\leq 1$, where $\bar x_t=\frac{1}{N}\sum^{N}_{i=1}x_{it}$ and $\tau\geq 0$. When $\tau=0$, it reduces to the approach proposed in BKP2016. However, the criterion with $\tau=0$ can derive consistent estimation for $\alpha_0$ under the restriction $\alpha_0>0.5$. This is due to the interruption of variance of error components. We overcome this disadvantage by benefiting from possibility of disappearance of time-serial covariance in error components in dynamic PCA. The criterion $cov(\bar{x}_t, \bar{x}_{t+\tau})$ with $\tau>0$, under the scenario of common factors from dynamic PCA, is proposed to obtain consistent estimation for all ranges of $\alpha_0$ in $[0, 1]$. Furthermore, joint estimation approach for $\alpha_0$ and another population parameter (necessary in estimation) is established. If this population parameter is observed, marginal estimation is also provided. From the aspect of theoretical contribution, asymptotic distributions of the joint and marginal estimators are both developed.
The main contribution of this paper is summarized as follows.
The rest of the paper is organized as follows. The model and the main assumptions are introduced in Section 2. Section 3 proposes both joint and marginal estimators that are based on the second moment criterion. Asymptotic properties for these estimators are established in Section 4. Section 5 reports the simulation results. Section 6 provides empirical applications to both cross-country macro-variables and stock returns in S$\&$P 500 market. Conclusions are included in Section 7. Main proofs are provided in Appendix A while some lemmas are listed in Appendix B. The proofs of lemmas are given in a supplementary material.
Let $x_{it}$ be a double array of random variables indexed by $i=1,\ldots,N$ and $t=1,\ldots,T$, over space and time, respectively. The aim of this paper is to measure the extent of the cross-sectional dependence of the data $\{x_{it}: i=1,\ldots,N\}$. In panel data analysis, there are two common models to describe cross-sectional dependence: spatial models and factor models. In BKP2016, a static approximate factor model is used. We consider a factor model as follows:
where ${\bf f}_t$ is the $m\times 1$ vector of unobserved factors (with m being fixed),
in which $\boldsymbol{\beta}_{i\ell}=(\beta_{i\ell 1}, \beta_{i\ell 2}, \ldots, \beta_{i\ell m})^{'}$, $\ell=0,1,\ldots,s$ are the associated vectors of unobserved factor loadings and $L$ is the lag operator, here $s$ is assumed to be fixed, and $\mu_i, i=1,2,\ldots,N$ are constants that represent the mean values for all sections, and $\{u_{it}: i=1,\ldots,N; t=1,\ldots,T\}$ are idiosyncratic components.
Clearly, we can write ((ref)) in the static form:
where
This model has been studied in SW2002 and F2009.
The dimension of ${\bf f}_t$ is called the number of dynamic factors and is denoted by $m$. Then the dimension of ${\bf F}_t$ is equal to $r=m(s+1)$. In factor analysis, $\boldsymbol{\beta}_i^{'}{\bf F}_t$ is called the common components of $x_{it}$.
We first introduce the following assumptions.
Now we provide some justification for these two assumptions.
The aim of this paper is to estimate the exponent $\alpha_0=\max_{\ell,k}(\alpha_{\ell k})$, which describes the extent of cross-sectional dependence. As in BKP2016 , we consider the cross-sectional average $\bar x_t=1/N\sum^{N}_{i=1}x_{it}$ and then derive an estimator for $\alpha_0$ from the information of $\{\bar x_t: t=1,2,\ldots,T\}$. BKP2016 use the variance of the cross-sectional average $\bar x_t$ to estimate $\alpha_0$ and carry out statistical inference for an estimator of $\alpha_0$. Specifically they show that
where $\widetilde{\kappa}_0$ is a constant associated with the common components and $c_N$ is a bias constant incurred by the idiosyncratic errors. From ((ref)), we can see that, in order to estimate $\alpha_0$, BKP2016 assume that $2\alpha_0-2>-1$, i.e. $\alpha_0>1/2$. Otherwise, the second term will have a higher order than the first term. So the approach by BKP2016 fails in the case of $0<\alpha_0<1/2$.
This paper is to propose a new estimator that is applicable to the full range of $\alpha_0$, i.e., $0\leq\alpha_0\leq1$. Based on the assumption that the common factors possess serial dependence that is stronger than that of the idiosyncratic components, we construct a so--called covariance criterion $Cov(\bar x_t, \bar x_{t+\tau})$, whose leading term does not include the idiosyncratic components for $0\leq\alpha_0\leq1$. In other words, the advantage of this covariance criterion over the variance criterion $Var(\bar x_t)$ lies on the fact that there is no interruption brought by the idiosyncratic components $\{u_{it}: i=1,2,\ldots,N; t=1,2,\ldots,T\}$ in $Cov(\bar x_t, \bar x_{t+\tau})$.
We define
in which $\mu_v = E[v_{i \ell k}]$, and $s$ and $m$ are the same as in ((ref)). Here $\kappa_{\tau}$ comes from the leading term of $Cov(\bar x_t, \bar x_{t+\tau})$.
Next, we illustrate how the covariance $Cov(\bar x_t, \bar x_{t+\tau})$ implies the extent parameter $\alpha_0$ in detail. Let $[N^{a}]$ ($a\geq 0$) denote the largest integer part not greater than $N^a$. For simplicity, let $[N^{b}]$ ($b\leq0$) denote $\frac{1}{[N^{-b}]}$. Moreover, to simplify the notation, throughout the paper we also use the following notation:
But we would like to remind the reader that $[N^{ka}]$ is actually not equal to $[N^{a}]^k$. Next, we will propose an estimator for $\alpha_0$ under two different scenarios: the joint estimator $(\widetilde{\alpha}_{\tau}, \widetilde{\kappa}_{\tau})$ under the case of some other parameters being unknown while the marginal estimator $\widehat\alpha_{\tau}$ for the case of some other parameters being known.
At first we consider the marginal estimator $\widehat\alpha_{\tau}$ to deal with the case where $\kappa_{\tau}$ is known. The parameter $\kappa_{\tau}$ describes the temporal dependence in the common factors. If we know this information in advance, the estimation of the extent of cross-sectional dependence becomes easy. We propose the following marginal estimation method.
Without loss of generality, we assume that $\alpha_{\ell k}=\alpha_0, \forall \ell=0,1,2,\ldots,s; k=1,2,\ldots,m$. Let Assumption 2 hold. Let $\bar x_{nt}$ be the cross--sectional average of $x_{it}$ over $i=1,2,\ldots,n$ with $n\leq N$. Similarly, $\bar\beta_{n\ell k}:=\frac{1}{n}\sum^{n}_{i=1}\beta_{i\ell k}$. Then
and
where $K_{n\ell k}=\sum^{n}_{i=[N^{\alpha_0}]+1}\beta_{i\ell k}$.
It follows that
A simple calculation indicates that
which implies
where $\kappa_{\tau}$ is defined in ((ref)).
Hence, for $0\leq\alpha_0\leq 1$, $\alpha_0$ can be estimated from ((ref)) using a consistent estimator for ${\rm Cov}(\bar x_t, \bar x_{t+\tau})$ given by
where $\bar x^{(1)}=\frac{1}{T-\tau}\sum^{T-\tau}_{t=1}\bar x_t$ and $\bar x^{(2)}=\frac{1}{T-\tau}\sum^{T-\tau}_{t=1}\bar x_{t+\tau}$ and $\bar{x}_t = \frac{1}{N} \sum_{i=1}^N x_{it}$. Thus, a consistent estimator for $\alpha_0$ is given by
Now we consider the case where $\kappa_{\tau}$ is unknown. Recalling ((ref)), we minimize the following quadratic form in terms of $\alpha$ and $\kappa$:
where $\widehat\sigma_n(\tau)$ is a consistent estimator for $Cov(\bar x_{nt}, \bar x_{n,t+\tau})$ of the form:
with $\bar x^{(1)}_n=\frac{1}{T-\tau}\sum^{T-\tau}_{t=1}\bar x_{nt}$ and $\bar x^{(2)}_n=\frac{1}{T-\tau}\sum^{T-\tau}_{t=1}\bar x_{n,t+\tau}$.
The joint estimator $(\widetilde{\alpha}_{\tau}, \widetilde\kappa_{\tau})$ can then be obtained by
where
and $$\widehat{Q}_{NT}^{(1)}(\alpha,\tau)=\frac{(\widehat q_1^{(1)}(\alpha,\tau)+[N^{2\alpha}]\widehat q_2^{(1)}(\alpha,\tau))^2}{N^{(1)}(\alpha)}.$$ We give the full derivation of ((ref)) in Appendix A.
This joint method estimates $\alpha_0$ and $\kappa_{\tau}$ simultaneously. The above derivations show that it is easy to derive $\widetilde{\alpha}_{\tau}$ and then $\widetilde{\kappa}_{\tau}$. Of course, we can also use some other estimation methods to estimate $\kappa_{\tau}$ and then $\alpha_0$. Notice that we use the weight function $w(n)=n^3$ in each summation part of the objective function $Q_{NT}^{(1)}(\alpha,\kappa,\tau)$ of ((ref)). The involvement of a weight function is due to technical necessity in deriving an asymptotic distribution for $(\widetilde{\alpha}_{\tau},\widetilde{\kappa}_{\tau})$.
In this section, we will establish asymptotic distributions for the proposed joint estimator $(\widetilde{\alpha}_{\tau}, \widetilde\kappa_{\tau})$ and the marginal estimator $\widehat\alpha_{\tau}$, respectively. We assume that $\alpha_{\ell k}=\alpha_0$, $\forall \ell=0,1,\ldots,s$ and $k=1,2,\ldots,m$ for simplicity. The notation $a\asymp b$ denotes that $a=O(b)$ and $b=O(a)$.
For any $1\leq i,j\leq m$ and $0\leq h\leq T-1$, we define
The following theorem establishes an asymptotic distribution for the marginal estimator $\widehat\alpha$. At first we define some notation. $\boldsymbol{\Sigma}_{\tau}=E({\bf F}_t{\bf F}_{t+\tau}^{'})$ and $\boldsymbol{\mu}_v=\mu_v{\bf e}_{m(s+1)}$, in which ${\bf e}_{m(s+1)}$ is an $m(s+1)\times 1$ vector with each element being $1$, $\boldsymbol{\Sigma}_v$ is an $m(s+1)$-dimensional diagonal matrix with each of the diagonal elements being $\sigma_v^2$ and
where
and `vec' means that for a matrix ${\bf X}=({\bf x}_1,\cdots,{\bf x}_n): q\times n$, $vec({\bf X})$ is the $qn\times 1$ vector defined as
Define
where $v_{NT}=\min([N^{\alpha_0}],T-\tau)$.
We are now ready to establish the main results of this paper in the following theorems and propositions.
Under some extra conditions, the conclusion of Theorem 1 can be simplified as given in Proposition 1 below.
From Proposition (ref), one can see that $\widehat\alpha$ is a consistent estimator of $\alpha_0$. Moreover, by a careful inspection on ((ref))--((ref)) in Theorem (ref) one can see that Condition ((ref)) can be replaced by some weak conditions to ensure the consistency of $\widehat\alpha$ under $(N, T)\rightarrow(\infty, \infty)$.
The following theorem establishes an asymptotic distribution for the joint estimator $(\widetilde{\alpha}, \widetilde{\kappa})$.
When the idiosyncratic components are independent, we can just use a finite lag $\tau$ (for example $\tau=1$). In this case, an asymptotic distribution for the estimator $\widehat{\alpha}$ is established in the following theorem.
Theorems (ref)--3 and Proposition (ref) establish some asymptotic properties for the joint estimator $(\widetilde{\alpha}, \widetilde{\kappa})$. Before we will give the proofs of Theorems 1--3 in Appendices B and C below, we have some brief discussion about Condition ((ref)), which is actually equivalent to the following three cases:
(a) $0<\alpha_0\leq\frac{1}{2}, \ [N^{\alpha_0}]<T-\tau, \ \frac{N^{1-3\alpha_0/2}}{(T-\tau)^{1/2}\log N}=o(1)$;
(b) $ \frac{1}{2}<\alpha_0\leq 1, \ [N^{\alpha_0}]<T-\tau; \ \frac{N^{1/2-\alpha_0/2}}{(T-\tau)^{1/2}\log N}=o(1)$;
(c) $\frac{1}{2}<\alpha_0\leq 1, \ [N^{\alpha_0}]\geq T-\tau, \ \frac{N^{1/2-\alpha_0}}{\log N}=o(1)$.
Under these three cases, we can provide some choices for $(N, T)$ as follows:
(d) $0<\alpha_0<\frac{1}{2}, \ [N^{\alpha_0}]<T-\tau; \ T=\tau+[N^{2-3\alpha_0}]$;
(f) $\frac{1}{2}<\alpha_0\leq 1, \ [N^{\alpha_0}]<T-\tau, \ T=\tau+[N^{\alpha_0}]$;
(g) $\frac{1}{2}<\alpha_0\leq 1, \ [N^{\alpha_0}]\geq T-\tau, \ T=\tau+[N^{\alpha_0}/\log(N)]$.
When $\tau\rightarrow\infty$, the term $\kappa_0$ will tend to $0$, because of $\boldsymbol{\Sigma}_{\tau}\rightarrow\textbf{0}$. So, as $\tau$ is very large, the value of $\ln (\kappa_0)$ may be negative in practice. Hence Theorem (ref) provides an alternative form for the asymptotic distribution of $N^{\widehat\alpha-\alpha_0}$ instead of $\widehat\alpha-\alpha_0$, and the case of $\tau$ being fixed is discussed in Theorem (ref).
In this section, we propose an estimator for the parameter $\sigma_\tau^2$ in the asymptotic variance of established theorems above.
Let $n=[N^{\widetilde{\alpha}_{\tau}}]$ and {
}The estimator for the first part of $\sigma_{\tau}^2$ in ((ref)) is
where $\widehat{\sigma}_{\tau,n}^2=\frac{1}{n-1}\sum^{n}_{i=1}\left(\widehat{\sigma}_{i,T}(\tau)-\widehat{\sigma}_{T}(\tau)\right)^2$ with $\widehat{\sigma}_T(\tau)=\frac{1}{n}\sum^{n}_{i=1}\widehat{\sigma}_{i,T}(\tau)$.
For the second term ((ref)), the proposed estimator is
where {
}with
and $\widehat{\sigma}_n(\tau)=\frac{1}{T-\tau}\sum^{T-\tau}_{t=1}\widehat{\sigma}_{n,t}(\tau)$.
Then the estimator for $\sigma^2_{\tau}$ is
The proof of Proposition 3 is provided in Appendix C. Next, we evaluate the finite--sample performance of the proposed estimation methods and the resulting theory in Sections 4 and 5 below.
In this section, we use three data generating processes (DGPs) to illustrate the finite sample performance of the joint and marginal estimators in different scenarios.
Example 1 studies the case of i.i.d error component and AR(1) modelled common factors. Example 2 extends the i.i.d error component in Example 1 to an AR(1) model. In Example 3, we investigate an MA(q) model for the common factors and an AR(1) type error component.
Before our analysis of each example, we provide a method of choosing an optimal value of $\tau$ in the following way.
The idea of this proposed criterion is that we choose a value of $\tau$ to make larger $\widetilde{\kappa}_{\tau}$ and smaller $Q_{NT}^{(1)}\left(\widetilde{\alpha}_{\tau}, \widetilde{\kappa}_{\tau}, \tau\right)$. As $\widetilde{\kappa}_{\tau}$ contains temporal dependence in the common factors, it is reasonable to consider its large value to take into account the information included in it. For $Q_{NT}^{(1)}\left(\widetilde{\alpha}_{\tau}, \widetilde{\kappa}_{\tau}\right)$, it is the objective function for the joint estimator and hence we expect its small value corresponding to an optimal $\tau$.
First, we consider the following two-factor model
The factors are generated by
with $f_{j,-50}=0$ for $j=1,2$ and $\zeta_{jt}\stackrel{i.i.d}{\sim}\mathcal{N}(0,1)$. The idiosyncratic components $u_{it}$ are i.i.d from $\mathcal{N}(0, 1)$ and independent of $\{\zeta_{jt}: t=1, 2, \ldots, T; j=1, 2\}$.
The factor loadings are generated as
where $v_{ir}\stackrel{i.i.d}{\sim}U(0.5, 1.5)$, $M=[N^{\alpha_0}]$ and $\rho=0.8$. Moreover, we set $\mu=1$ and $\rho_j=0.9$ for $j=1,2$.
Under this data generating process, the numerical values of $\widehat{\alpha}$ and $\widetilde{\alpha}$ are reported in Table (ref). The confidence interval for $\alpha_0$ is also calculated, i.e.
which can be derived for the asymptotic distribution of $\widetilde{\alpha}$ in Theorem (ref).
In this table, the number of cross-sections was $N=200$ and the time-length was $T=200$. As $\tau=0$, the estimator $\widehat{\alpha}$ is equivalent to the estimator provided in BKP2016. Moreover, we calculate the marginal and joint estimates when $\tau=1, 2, 3, 4$. From Table (ref), it can be seen that the estimator in BKP2016 behaves well in the case of $\alpha_0=0.8$ while becomes inconsistent as $\alpha_0=0.5$ and $\alpha_0=0.2$. When $\tau>0$, our estimator performs well for all cases including $\alpha_0=0.8, 0.5, 0.2$. However, as $\tau$ increases, the confidence interval will become larger. This phenomenon is consistent with our theoretical result since the confidence interval in ((ref)) depends on $\widetilde{\kappa}_{tau}$ and this parameter decreases for the AR(1) model as $\tau$ increases. Furthermore, compared with the marginal estimation, the joint estimation is a bit worse than marginal estimator as expected due to more information is known in marginal estimation.
From this example, we can see that the estimator with $\tau>0$ is consistent when $\alpha_0<0.5$ at the cost of larger variance. Moreover, the estimator $\widetilde{\alpha}_{\widetilde{\tau}}$, with the choice of $\widetilde{\tau}$, is better than others, although there is a bit deviation. Intuitively, since $\kappa_{\tau}$ will decrease as $\tau$ increases under this example, and the error component has no temporal dependence, the optimal $\tau$ should be $1$ intuitively.
In this part, we also consider the factor model ((ref)). In this model, the common factors and factor loadings follow ((ref)) and ((ref)) respectively. The idiosyncratic error component $u_{it}$ follows an AR(1) model as follows:
where $g_{-50}=0$ and $\eta_i\sim\mathcal{N}(0, 1)$. Moreover, $u_{it}$ are independent of $\zeta_{jt}$ in ((ref)). Here $h=0.2$ and $h=\frac{1}{\sqrt{N}}$ which are smaller than the strength of time serial dependence in the common factors with $\rho=0.9$.
The results when $h=0.2$ or $h=\frac{1}{\sqrt{N}}$ are listed in Tables (ref) and (ref), respectively.
These two tables are derived as $N=400, T=200$. In Table (ref), the values of $\tau$ are comparably large in order to ensure that the common factors and the idiosyncratic error components have different strengths of time-serial dependence. In fact, the autocorrelation of the common factors and the error component with time--lag $\tau$ are of the orders $\rho^{\tau}_j$ and $h^{\tau}$, respectively. When $h$ is constant, the error component has a weak strength order only as $\tau$ tends to infinity. When $h=\frac{1}{\sqrt{N}}$, $h^{\tau}$ tends to zero for any value of $\tau$. This is why $\tau$ in Table (ref) takes relatively small values.
This example is more complicated than Example 1 due to time--serial dependence in the error component. Similar to Example 1, the marginal or the joint estimator with $\tau=0$ performs inconsistently when $\alpha_0=0.5$ and $\alpha_0=0.2$, while the marginal or the joint estimator with $\tau>0$ is consistent for all cases. Compared with Example 1, all the results have relatively larger variances due to the complex structures.
Furthermore, the choice for $\widetilde{\tau}$ in Table (ref) is around 5 while that in Table (ref) is close to $1$. These choices are reasonable due to strong time-serial dependence in the error component for the date generating process in Table (ref).
Now we consider the factor model ((ref)). In this model, factor loadings follow ((ref)), respectively. The common factors $f_{jt}$ and the idiosyncratic error component $u_{it}$ follow MA(2) models as follows:
and
where $Z_{j,s}\sim\mathcal{N}(0, 1)$, $K_s\sim\mathcal{N}(0, 1)$ with $s=-1, 0, 1, 2, \ldots, T$ and $\eta_i\sim\mathcal{N}(0, 1)$. Moreover, $u_{it}$ are independent of $\zeta_{jt}$ in ((ref)). Here $\theta_1=0.8$, $\theta_2=0.6$, $h_1=\frac{1}{\log(N)}$ and $h_2=\frac{1}{\sqrt{N}}$.
Table (ref) presents the results for the case of $N=400$ and $T=200$.
As well known, MA(2) model has zero autocorrelation when $\tau>2$. Hence Table (ref) reports the results with $\tau=0, 1, 2$. We impose a “weaker" MA model for error components in the sense of the coefficients tend to zero. This is to guarantee weaker strength of the time-serial dependence in the error component than that for the common factors.
Except similar observations to the first two examples, the proposed estimator behaves a bit better under MA structure than AR structure in examples 1 and 2. The main reason relies on that MA structure has larger time-serial dependence than that for the AR structure.
Similarly, the value of $\widetilde{\tau}$ is also around $1$. This is because the temporal dependence in the error component is quite weak.
In this section, we show how to obtain an estimate for the exponent of cross--sectional dependence, $\alpha_0$, for each of the following panel data sets: quarterly cross-country data used in global modelling and daily stock returns on the constitutes of Standard and Poor $500$ index.
We provide an estimate for $\alpha_0$ for each of the datasets: Real GDP growth (RGDP), Consumer price index (CPI), Nominal equity price index (NOMEQ), Exchange rate of country $i$ at time $t$ expressed in US dollars (FXdol), Nominal price of oil in US dollars (POILdolL), and Nominal short-term and long-term interest rate per annum (Rshort and Rlong) computed over $33$ countries.\footnote{The datasets are downloaded from http://www-cfap.jbs.cam.ac.uk/research/gvartoolbox/download.html.} The observed cross-country time series, $y_{it}$, over the full sample period, are standardized as $x_{it}=(y_{it}-\bar y_i)/s_i$, where $\bar y_{i}$ is the sample mean and $s_i$ is the corresponding standard deviation for each of the time series. Table (ref) reports the corresponding results.
For the standardized data $x_{it}$, we regress it on the cross-section mean $\bar x_t=\frac{1}{N}\sum^{N}_{i=1}x_{it}$, i.e., $x_{it}=\delta_i\bar x_t+u_{it}$ for $i=1,2,\ldots,N$, where $\delta_i$, $i=1,2,\ldots,N$, are the regression coefficients. With the availability of the OLS estimate $\widehat\delta_i$ for $\delta_i$, we have the estimated versions, $\widehat u_{it}$, of the form: $\widehat u_{it}=x_{it}-\hat\delta_i\bar x_t$.
Since our proposed estimation methods rely on the different extent of serial dependence of the factors and idiosyncratic components, we provide some autocorrelation graphs of $\{\bar x_t=\frac{1}{N}\sum^{N}_{i=1}x_{it}: t=1,2,\ldots,T\}$ and $\{\bar u_t=\frac{1}{N}\sum^{N}_{i=1}u_{it}: t=1,2,\ldots,T\}$ for each group of the real dataset under investigation (see Figures (ref)--(ref)). From these graphs, it is easy to see that CPI, NOMEQ, FXdol and POILdolL have distinctive serial dependences in the factor part $\bar x_{t}$ and idiosyncratic part $\bar u_{t}$. All the observed real data $x_{it}$ are serially dependent.
Due to the existence of serial dependence in the idiosyncratic component, we use the proposed second moment criterion. The marginal estimator $\widehat\alpha$ and the joint estimator $\widetilde{\alpha}$ for these real data are provided in Table (ref). We use $\tau=10$ for two estimators. We can see from Table (ref) that the values of $\widehat\alpha$ and $\widetilde{\alpha}$ are different from the those provided by BKP2016. Some estimated values are not $1$. This phenomenon implies that a factor structure might be a good approximation for modelling global dependency, and the value of $\alpha_0=1$ typically assumed in the empirical factor literature might be exaggerating the importance of the common factors for modelling cross-sectional dependence at the expense of other forms of dependency that originate from trade or financial inter-linkage that are more local or regional rather than global in nature. Furthermore, it is noted that our model is different from that given by BKP2016 and the difference mainly lies on that our model only imposes serial dependence on factor processes and assumes that the idiosyncratic errors are independent. Different models may bring in different exponents.
One of the important considerations in the analysis of financial markets is the extent to which asset returns are interconnected. The classical model is the capital asset pricing model of S1964 and the arbitrage pricing theory of R1976. Both theories have factor representations with at least one strong common factor and an idiosyncratic component that could be weakly cross-sectionally correlated (see C1983). The strength of the factors in these asset pricing models is measured by the exponent of the cross-sectional dependence, $\alpha_0$. When $\alpha_0=1$, as it is typically assumed in the literature, all individual stock returns are significantly affected by the factors, but there is no reason to believe that this will be the case for all assets and at all times. The disconnection between some asset returns and the market factors could occur particularly at times of stock market booms and busts where some asset returns could be driven by some non-fundamentals. Therefore, it would be of interest to investigate possible time variations in the exponent $\alpha_0$ for stock returns.
We base our empirical analysis on daily returns of $96$ stocks in the Standard $\&$ Poor $500$ (S$\&$P500) market during the period of January, 2011-December, 2012. The observations $r_{it}$ are standardized as $x_{it}=(r_{it}-\bar r_i)/s_i$, where $\bar r_i$ is the sample mean of the returns over all the sample and $s_i$ is the corresponding standard deviation. For the standardized data $x_{it}$, we regress it on the cross-section mean $\bar x_t=\frac{1}{N}\sum^{N}_{i=1}x_{it}$, i.e., $x_{it}=\delta_i\bar x_t+u_{it}$ for $i=1,2,\ldots,N$, where $\delta_i$, $i=1,2,\ldots,N$, are the regression coefficients. Based on the OLS estimates: $\widehat\delta_i$ for $\delta_i$, we define $\widehat u_{it}=x_{it}-\hat\delta_i\bar x_t$. The autocorrelation functions (ACFs) of the cross-sectional averages $\bar x_{t}=\frac{1}{N}\sum^{N}_{i=1}x_{it}$ and $\bar u_{t}=\frac{1}{N}\sum^{N}_{i=1}u_{it}$ are presented in Figure (ref).
From Figure (ref), we can see that the serial dependency of the common factor component is stronger than that of the idiosyncratic component. We use the estimates $\widehat\alpha$ and $\widetilde{\alpha}$ to characterize the serial dependences of the common factors and the idiosyncratic component. The estimates $\widehat\alpha$ and $\widetilde{\alpha}$ are calculated with the choice of $\tau=10$. Table (ref) reports the estimates with several different sample sizes. As comparison, the estimates from BKP2016 are also reported. From the table, we can see that their estimation method does not work when $\alpha$ is smaller than $1/2$. The results also show that the cross-sectional exponent of stock returns in S$\&$P500 are smaller than $1$. This indicates the support of using different levels of loadings for the common factor model as assumed in Assumption 2, rather than using the same level of loadings in such scenarios.
Furthermore, Figure (ref) provides the marginal estimate $\widehat\alpha$ and the joint estimate $\widetilde{\alpha}$ for the first $130$ days of all the period. It shows that the estimated values for $\alpha_0$ with the two methods are quite similar. On the other hand, since a $130$-day period is short, meanwhile, it is reasonable that the estimates didn't change very much.
In this paper, we have examined the issue of how to estimate the extent of cross--sectional dependence for large dimensional panel data. The extent of cross--sectional dependence is parameterized as $\alpha_0$, by assuming that only $[N^{\alpha_0}]$ sections are relatively strongly dependent. Compared with the estimation method proposed by BKP2016, we have developed a unified `moment' method to estimate $\alpha_0$. Especially, when stronger serial dependence exists in the factor process than that for the idiosyncratic errors (dynamic principal component analysis), we have recommended the use of the covariance function between the cross-sectional average values of the observed data at different lags to estimate $\alpha_0$. One advantage of this new approach is that it can deal with the case of $0\leq\alpha_0\leq1/2$.
Due to some unknown parameters involved in the panel data model, in addition to the proposed marginal estimation method, we have also constructed a joint estimation method for $\alpha_0$ and the related unknown parameters. The asymptotic properties of the estimators have all been established. The simulation results and an empirical application to two datasets have shown that the proposed estimation methods work well numerically.
Future research includes discussion about how to estimate factors and factor loadings in factor models, and determine the number of factors for the case of $0<\alpha_0<1$. Existing methods available for factor models, such as Baing2002, AH2013, ABC2010, O2010, for the case of $\alpha_0=1$, may not be directly applicable, and should be extended to deal with the case of $0<\alpha_0<1$. Such issues are all left for future work.
The first, the third and fourth authors acknowledge the Australian Research Council Discovery Grants Program for its support under Grant numbers: DP150101012 & DP170104421. Thanks from the second author also go to the Ministry of Education in Singapore for its financial support under Grant $\#$ ARC 14/11.
{
}
{
This material includes two appendices, i.e. Appendices A and B. Appendix A provides the proofs of Theorems 1 and 2 in the main paper. Some lemmas used in the proofs of Theorems 1 and 2 are given in Appendix B. The proof of Theorem 3 in the main paper is omitted since it is similar to that of Theorem 2.
Throughout this material, we use $C$ to denote a constant which may be different from line to line and $||\cdot||$ to denote the spectral norm or the Euclidean norm of a vector. In addition, the notation $a_n\asymp b_n$ means that $a_n=O_P(b_n)$ and $b_n=O_P(a_n)$.
\setcounter{equation}{0}