EconBase
← Back to paper

The Incidental Parameters Problem in Testing for Remaining Cross-section Correlation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,167 characters · 14 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Incidental Parameters Problem in Testing for Remaining Cross-section Correlation

\singlespace \bibpunct{(}{)}{,}{a}{,}

abstractIn this paper we consider the properties of the CDtest2004, doi:10.1080/07474938.2014.956623 CD test for cross-section correlation when applied to residuals obtained from panel data models with many estimated parameters. We show that the presence of period-specific parameters leads the CD test statistic to diverge as the time dimension of the sample grows. This result holds even if cross-section dependence is correctly accounted for and hence constitutes an example of the Incidental Parameters Problem. The relevance of this problem is investigated both for the classical Two-way Fixed Effects estimator as well as the Common Correlated Effects estimator of ECTA:ECTA692. We suggest a weighted CD test statistic which re-establishes standard normal inference under the null hypothesis. Given the widespread use of the CD test statistic to test for remaining cross-section correlation, our results have far reaching implications for empirical researchers. Keywords. Panel Data, Cross-section Dependence, Factor Model, Time Fixed Effects, U-statistic. JEL Classifications. C12, C23, C33.

\onehalfspace

Introduction

Given that economic agents rarely act entirely independently of each other, modeling cross-section dependence plays a prominent role in panel data econometrics. Time fixed effects are probably the simplest way of addressing this issue and allow controlling for a common trend whose effect is homogeneous across cross-section units. During the last decade, interactive fixed effects models have become a popular, more general alternative. They assume the presence of a factor error structure, i.e. a small number of unobserved common trends interacted with entity-specific slope coefficients. Using either of the two modeling possibilities begs the question whether cross-section dependence is adequately accounted for.

In this paper, we show that the application of tests for cross-section dependence to regression residuals obtained from two-way fixed effects models or interactive effects models is problematic. We use the popular CD test statistic of CDtest2004, doi:10.1080/07474938.2014.956623 and show that the inclusion of period-specific parameters introduces a bias term of order $\sqrt{T}$. In order to avoid erroneous rejection of the null hypothesis of no unaccounted cross-section dependence, we suggest a modified test statistic that re-weights cross-section covariances with Rademacher distributed weights. This weighted CD test statistic is asymptotically standard normal and has very good size under appropriate regularity conditions on the chosen weights.

The use of the CD test statistic as a misspecification test for regression models that already account for cross-section dependence has repeatedly been observed in empirical studies (see e.g. HOLLY2010160, JAE:JAE2338, JAE:JAE2468 and EBERHARDT201545 among others). In some applications, the CD test statistic is explicitly used as a model-selection tool, interpreting a reduction in the absolute value of the CD test statistic as an indication of a better model. In other cases, only specifications not rejected by the CD test statistic (given some significance level) were considered. However, to the best of our knowledge, there are currently no theoretical results in the literature that could justify this practice. Furthermore, the scenario of testing for cross-section dependence in cross-sectionally demeaned or defactored data is usually completely ignored in any of the large scale Monte Carlo studies integral to theoretical and empirical papers in the field. Only very recently, mao2018testing reported the size of three tests for cross-section dependence in a subset of his simulation experiments. However, despite clear evidence of excessive over-rejections, these results are neither discussed nor investigated theoretically. Therefore, this study is the first one to investigate the properties of cross-section dependence tests applied to residuals of models that control for sources of common variation across cross-sections. In particular, given its popularity in the applied panel data literature, we restrict our attention to the CD test statistic of CDtest2004, doi:10.1080/07474938.2014.956623. Our interest lies in residuals of models that characterize cross-section dependence as driven by a small number of unobservable factors. This comprises both the two-way fixed effects (2WFE) model as well as models with a multifactor error structure where factors are interacted with unit-specific slope coefficients. The results that we obtain are summarized as follows:

enumerate• The application of the CD test to residuals obtained from a model where common factors enter either as time fixed effects or through a multifactor error structure renders the test statistic biased for any fixed $T$, and divergent as $T\to\infty$. • In addition to the mean of the CD test statistic, even its variance may be affected. This can result in an asymptotically degenerate distribution of the test statistic. • A simple way of eliminating bias is to construct the CD test statistic from specifically weighted cross-section covariances rather than correlations. This leads to a valid test statistic for remaining cross-section correlation with good small sample properties in simulations.

The degeneracy of the CD statistics can be seen as a manifestation of the Incidental Parameters Problem (IPP) of neyman1948. In this respect, this paper contributes to this branch of the literature. So far, major focus in the panel data literature has been related to the IPP stemming from estimated individual specific effects. Our paper is the first one to document an asymptotically non-negligible impact of estimated time specific common parameters which cannot be ruled out by restrictions on the relative expansion rate of $N$ and $T$. Furthermore, since the CD statistic can be seen as a time-series average of second degree (degenerate) U-statistics, our results shed some light on the potential impact of the IPP beyond simple cross-sectional averages. Lastly, while this article only considers linear models, the average correlation approach to testing for cross-section dependence was extended to nonlinear and nonparametric panel data models by OBES:OBES646 and chen_gao_li_2012, respectively. Hence, problems documented in this paper carry over to post-estimation properties of non-linear models discussed in e.g. fernandez_weidner_2014, BonevaLinton2016, or chenvalweidner2014.

The remainder of this paper is structured as follows: Section (ref) introduces the testing problem. In Sections (ref) we present asymptotic results for models with two-way fixed effects and multifactor error structures. In Section (ref) we discuss standard approaches for bias correction and propose a weighted CD test statistic that achieves this goal. Sections (ref) and (ref) illustrate the problem documented in this paper by means of simulated and real data. Section (ref) concludes. Additional technical discussions, Monte Carlo studies, and all proofs are relegated to the Supplementary Appendix.

Notation: $\bm I_m$ denotes an $m \times m$ identity matrix and the subscript is sometimes disregarded from for the sake of simplicity. $\boldsymbol{0}$ denotes a vector of zeros while $\mathbf{O}$ stands for a matrix of zeros. $\bm s_m$ denotes a selection vector all of whose elements are zero except for element $m$ which is one. $\bm \iota$ is a vector entirely consisting of ones. The dimension of these latter vectors and matrices is generally suppressed for the sake of simplicity and needs to be inferred from context. For a generic $m \times n$ matrix $\bm A$, $\bm P_{\bm A} = \bm A(\bm A'\bm A)^{-1}\bm A'$ projects onto the space spanned by the columns of $\bm A$ and $\bm M_{\bm A} = \bm I_m - \bm P_{\bm A}$. Furthermore, $\operatorname{rk}(\bm A)$ denotes the rank of $\bm A$, $\operatorname{tr}(\bm A)$ its trace and $\| \bm A \| =\left( \operatorname{tr} (\bm A'\bm A)\right)^{1/2}$ the Frobenius norm of $\bm A$. For a set of $m \times n$ matrices $\{ \bm A_1, \ldots, \bm A_N \}$, $\overline{\bm A} = N^{-1}\sum_{i=1}^N \bm A_i$. $\delta$ and $M$ stand for a small and large positive real number, respectively. For two real numbers $a$ and $b$, $a \vee b = \max\{ a,b\}$. Lastly, $\operatorname{\mathcal{O}}_P(\cdot)$ and $\operatorname{o\mathcal{}}_P(\cdot)$ express stochastic order relations.

The testing problem

Let $\bm z_{i}$ be a $T$-dimensional data vector observed over $N$ cross-sectional units indexed by $i$. Combining all $\bm z_{i}$ we obtain a two-dimensional data array of panel data (or longitudinal data). In empirical research it is common to investigate whether $\bm z_{i}$ can be regarded as independent over $i$ in order to select a model that can properly characterize the statistical properties of the data. In particular, researchers might be interested in the statistical hypothesis

equation[equation omitted — 83 chars of source]

where we use the $\bot$ notation to denote independence. Most often $\bm z_{1}, \ldots, \bm z_N$ contain residuals obtained from a regression model that does not allow for cross-section dependence, e.g. an entity fixed effects model or linear regression.

By far the most widely used test for cross-section dependence (correlation, to be precise) is the CD test of CDtest2004, doi:10.1080/07474938.2014.956623, which is based on a simple rescaled sum of all pairwise cross-section correlation coefficients, formally denoted

equation[equation omitted — 166 chars of source]

Here,

align[align omitted — 248 chars of source]

is the pairwise sample correlation coefficient between units $i$ and $j$. Obviously, computing the CD test statistic involves obtaining $N(N-1)/2$ parameter estimates, each of which converges to the true parameter value at rate $\sqrt{T}$ only. These circumstances are reminiscent of the panel data setup considered in e.g. phillipsMoon1999, Hahn2002, where estimation of many (incidental) parameters in linear regression models turns out to have distributional effects on the asymptotic properties of common parameters. That is, they cause the incidental parameter problem. By contrast, the asymptotic distribution of the CD test statistic is unaffected by the estimation of all $N(N-1)/2$ cross-section correlation coefficients involved in its construction. In fact, applied to the residuals of a linear regression model with strictly exogenous regressors, and individual specific means, the CD test statistic is asymptotically normal as long as $N,T\to\infty$ (see PesaranBook, Theorem 2), under $H_{0}$ of independence (or even limited local dependence). However, as shown below, this result does not hold when the model specification includes period-specific parameters.

Heuristic discussion of the main result

Consider a linear model where cross-section dependence is due to time fixed effects, so that both unobserved heterogeneity over cross-sections and time enters the model additively. That is the relation between the $T\times 1$ vector $\bm y_i$ and the $T \times m$ matrix $\bm X_i$ is formally denoted

equation[equation omitted — 145 chars of source]

Here, $\mu_i$ and $\bm \tau$ denote an entity-specific intercept and a $T \times 1$ vector of time-specific common parameters $\bm \tau$, respectively. $\bm \varepsilon_i$ is a vector of idiosyncratic error terms, independent across cross-sectional units. In the example of a Difference-in-Differences framework, $\bm \tau$ is the common trend affecting both treated and untreated individuals while the treatment indicator as well as other covariates are contained in $\bm X_{i}$. For the sake of simplicity, we assume $\bm \beta$ and $\mu_i$ to be known, so that $\bm \beta=\boldsymbol{0}$ and $\mu_i = 0 \; \forall i$ can be imposed without loss of generality. This highly restrictive assumption leaves the leading terms in the analysis below unaffected and is hence innocuous for the expository purpose of this section. Model (ref) reduces to

equation[equation omitted — 128 chars of source]

We additionally assume that the variance of $\varepsilon_{i,t}$ is known and fixed to $\sigma_{i}^{2}=1$. In this setup, $y_{i,t}$ are clearly cross-sectionally dependent because of $\tau_{t}$; however we are interested in testing whether $\varepsilon_{i,t}$ are cross-sectionally uncorrelated. Given that $\varepsilon_{i,t}$ are unobserved, the common effects $\tau_{t}$ need to be estimated. The most natural approach is to estimate $\tau_{t}$ by OLS so that $ \widehat{\tau}_{t} = \overline{y}_{t} = \frac{1}{N}\sum_{i=1}^{N}y_{i,t}$ . Using the residuals

equation[equation omitted — 101 chars of source]

the CD test statistic is given by

align[align omitted — 366 chars of source]

Given this definition,

align[align omitted — 109 chars of source]

implying that the first term in (ref) cancels out. The expression $\sum_{t=1}^{T}\sum_{i=1}^{N}\widehat{\varepsilon}_{i,t}^2$ in the second term is nothing more than the Sum of Squared Residuals (SSR) of the estimated model, a statistic that is of order $\operatorname{\mathcal{O}}_P(NT)$. Consequently,

align[align omitted — 58 chars of source]

even though the error-terms $\varepsilon_{i,t}$ are cross-sectionally independent. Hence, the procedure commonly used by practitioners is prone to finding spurious cross-sectional dependence in the data.

The results shown for a model with additive time effects and homoscedastic errors are not coincidental. In fact, as we show later in this paper, they carry over to a more complex model where cross-section dependence is generated by a multifactor error structure and/or the errors are heteroscedastic.

Asymptotic Results

In this section we provide formal asymptotic results for CD statistic based on the residuals obtained from a 2WFE model or a model with multifactor error structure. For the sake of simplicity, we continue to assume that $\bm \beta$ is known throughout this section. If the vector of slope coefficients $\bm \beta$ were unknown, its estimator would converge (under correct model specification) to the true parameter value at the conventional rate $\sqrt{NT}$. Hence, convergence to the true parameter values is fast enough to ensure that any terms present in CD test statistic involving $\widehat{\bm \beta}- \bm \beta_{0}$ are asymptotically negligible. We shall not attempt to prove this claim formally but rely on Monte Carlo results of Section (ref) to support this conjecture. By contrast, the documented effects on the first two moments of the CD test statistic resulted from the fact that estimates of common, period-specific parameters converge at a rate slower than $\sqrt{NT}$.

Two-way error component structure

As in Section (ref) we assume that the true model is given by equation (ref). Additionally, we make the following assumptions on the model errors.

assumption[Errors]\ \begin{enumerate} • Let $\varepsilon_{i,t}=\sigma_{i}\eta_{i,t}$ where $\eta_{i,t}$ is independently and identically distributed across both $i$ and $t$ with $\operatorname{E}\left[\eta_{i,t}\right]=0$, $\operatorname{E}\left[ \eta_{i,t}^2\right]= 1$ and $\operatorname{E}\left[ |\eta_{i,t}|^8\right]< M$ for some $M < \infty$. • $\sigma_{i}$ is defined over an interval $[\delta;M]$ with $0<\delta<M<\infty$. It is independently and identically distributed across $i$, and $\sigma_{i}$ is independent of $\eta_{i,t}$ for all $i$ and $t$. \end{enumerate}

For technical reasons, we assume that all stochastic variables in this paper have finite eighth moments. This is a sufficient condition, which facilitates proving joint convergence of the test statistics considered in this paper, see also JAE:JAE2496. Assumption (ref) is general enough to cover several models of conditional heteroscedasticity in $\varepsilon_{i,t}$. Alternatively this assumption can be formulated in terms of unconditional variances, where $\sigma_{i}$ are fixed numbers. However, in addition to having severe conceptual shortcomings (as discussed in ECTA:ECTA1599) such an approach would lead to incorrect conclusions concerning the power of a CD test statistic for remaining cross-section correlation. See Section (ref) in the Supplementary Appendix for a detailed discussion on power.

For example, natural examples for $\sigma_{i}$ are either a standard exponential skedastic function

equation[equation omitted — 54 chars of source]

or a location-scale model with

equation[equation omitted — 48 chars of source]

Both satisfy the required restrictions, as long as $\mu_{i}$ has a bounded support. Furthermore, we denote different cross-sectional averages of $\sigma_{i}$ by $\overline{\sigma^{k}}=N^{-1}\sum_{i=1}^{N}\sigma_{i}^{k}$ for $k\in \mathbb{Z}$, and the corresponding population quantities by $\operatorname{E}[\sigma_{i}^{k}]$. Assumption (ref) guarantees that these quantities are well defined for all finite $k$.

Using these definitions the CD test statistic obtained from a model with unknown period-specific effects can be characterized as follows:

theoremSuppose that $\bm \beta$ is known. Under Assumption (ref), \begin{align} CD &= \sqrt{\frac{2}{TN(N-1)}}\sum_{i=2}^{N} \sum_{j=1}^{i-1}\sum_{t=1}^{T}\varepsilon_{i,t}\varepsilon_{j,t}\left(\sigma_{i}^{-1}-\overline{\sigma^{-1}}\right)\left(\sigma_{j}^{-1}-\overline{\sigma^{-1}}\right) +\sqrt{T} \Xi \\ &+ \operatorname{\mathcal{O}}_{P}(\sqrt{N}T^{-1})+ \operatorname{\mathcal{O}}_{P}(T^{-1/2})+ \operatorname{\mathcal{O}}_{P}(N^{-1/2})+ \operatorname{\mathcal{O}}_{P}(N^{-1}\sqrt{T}), \notag \end{align} where \begin{align} \Xi &= \sqrt{\frac{N}{2(N-1)}} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \varepsilon_{i,t}^{2}\left(\left(\overline{\sigma^{-1}}\right)^{2} - 2\overline{\sigma^{-1}}\sigma_{i}^{-1} \right). \end{align} Furthermore, let $\Omega = \left(1-2\operatorname{E}[\sigma_{i}]\operatorname{E}[\sigma_{i}^{-1}] + \operatorname{E}[\sigma_{i}^{2}](\operatorname{E}[\sigma_{i}^{-1}])^{2}\right)^2$. Then, \begin{equation} CD - \sqrt{T}\Xi \stackrel{d}{\longrightarrow} N\left(0, \Omega \right) \end{equation} as $N,T\to \infty$ jointly provided that $\sqrt{T}N^{-1} \to 0$ and $\sqrt{N}T^{-1}\to 0$.

Theorem (ref) shows that the inclusion of time fixed effects into the model specification has an asymptotically non-negligible effect on the first two moments of the CD test statistic. Via expansion with appropriate functions of $\{\sigma_i\}_{i=1}^{N}$, it can be shown that $\Xi$ indicates the presence of a deterministic bias of order $\sqrt{T}$. Abstracting from this bias, (ref) is dominated by an expression that reflects the CD test statistic obtained from the true model errors but imposing an incorrect normalization.

Restrictions on the expansion rates of $N$ and $T$, i.e. $\sqrt{N}T^{-1} \to 0$ and$\sqrt{T}N^{-1} \to 0$, are satisfied if one considers diagonal asymptotic expansion schemes as in fernandez_weidner_2014, where $N T^{-1}\to \kappa^{2}$ with $ \kappa\in (0;\infty)$. This way $N/T + T/N$ remains a bounded constant asymptotically. The intuition for the need to restriction both $\sqrt{T}N^{-1} \to 0$ and $\sqrt{N}T^{-1} \to 0$ comes from the fact that the feasible CD statistic is a non-linear function in both $\widehat{\sigma}_{i}$ and $\widehat{\tau}_{t}$. Hence the rate restrictions derived for general non-linear panel data models, as studied by fernandez_weidner_2014, are naturally applicable.

The behavior of the CD test statistic in a model with time fixed effects can be seen as an example of the incidental parameters problem (IPP) of neyman1948, since the bias of the CD test statistic is due to the estimation of $T$ period-specific intercepts $\tau_{t}$. Interestingly, in the context of estimating linear dynamic panel data models, estimation of the time effects $\tau_t$ does not introduce any asymptotic bias into the FE estimator with strictly or weakly exogenous regressors hahnmoon2006. In non-linear models, estimation of the time effects $\tau_{t}$ affects the asymptotic mean of the estimator for slope parameters associated with explanatory variables fernandez_weidner_2014 with the corresponding bias being proportional to $\sqrt{T N^{-1}}$. In this sense, our result adds new insights into the literature in that it highlights a scenario where the inclusion of time fixed effects into a linear model has non-standard implications for the asymptotic distribution of the statistic of interest.

The results of Theorem (ref) suggest that asymptotically standard normal inference can be recovered by bias-correcting and rescaling (ref). Before considering this remedy in Section (ref), it is important to emphasize the following special case:

corollaryUnder Assumption (ref) and given $P(\sigma_{i}=\sigma)=1$, \begin{align} CD &= -\sqrt{\left(T-\frac{T}{N}\right)/2} +\operatorname{\mathcal{O}}_{P}(R_{N,T}), \end{align} where $R_{N,T}=(N^{-1/2}\vee T^{-1/2}\vee N^{-1}\sqrt{T}\vee T^{-1}\sqrt{N})$.

Corollary (ref) provides more intuition about the approximate value of the bias term $\sqrt{T}\Xi$, suggesting that it should be reasonably close to $-\sqrt{T/2}$. More importantly, the result reveals that the leading stochastic component in the CD test statistic, the first term on the right-hand side of equation (ref), cancels out when error variances are homogeneous across $i$. Instead, random variation around the bias term $-\sqrt{T/2}$ is of order $\operatorname{\mathcal{O}}_P(N^{-1/2}\wedge T^{-1/2})$, rendering the distribution of a modified version of the test statistic asymptotically degenerate.

The special case of homogeneous error variances hence entails consequences for the CD test statistic that are qualitatively different from those in the more general case where $\sigma_{i}$ may differ across $i$. Again, it would be possible to allow for asymptotically normal inference by adequately rescaling the modified test statistic. However, the resulting statistic would be of little practical use since the main source of variation is not related to error covariances across cross-sections, but is simply driven by variance of the idiosyncratic components.

Multifactor error structure

Following ECTA:ECTA692 we consider a model where cross-section dependence enters the model via a multifactor error structure. This model is described by the $T\times 1$ vector $\bm y_i$ and the $T\times m$ matrix $\bm X_i$ defined as

align[align omitted — 159 chars of source]

where $\bm F = \left[\bm f_1, \bm f_2, \ldots, \bm f_T \right]'$ denotes a $T\times r$ matrix of unobserved common factors and where $\bm E_i = \left[ \bm e_{i,1}, \bm e_{i,2}, \ldots, \bm e_{i,T} \right]'$ is a $T \times m$ matrix of idiosyncratic variation in $\bm X_{i}$. The most popular estimator designed to estimate the parameter vector $\bm \beta$ in this specific model is the Common Correlated Effects (CCE) estimator of ECTA:ECTA692 which amounts to augmenting a linear regression model with cross-section averages of potentially all variables available to the researcher in order to account for the effect of unobservable common factors. In a model with homogeneous slope coefficients, we have

align[align omitted — 622 chars of source]

Here $\overline{\bm C}$ ($\overline{\bm U}$) are defined implicitly in terms of $\bm \beta$ and the corresponding average factor loadings (average errors). While the CCE estimator is agnostic about the true number of factors that affect the data, it has been shown that consistent estimation of the parameters of interest requires that the number of cross-section averages is at least as large as the true number of factors that drive the data Westerlund2013247.

Given that the factor estimator $\widehat{\bm F}$, defined in Eq. (ref), is constructed based on observed data, restrictions on the DGP of both $y_{i,t}$ and $\bm x_{i,t}$ need to be imposed.

assumption\ \begin{enumerate} • Let $\varepsilon_{i,t}=\sigma_{i}\eta_{i,t}$ where $\eta_{i,t}$ is independently and identically distributed across both $i$ and $t$ with $\operatorname{E}\left[\eta_{i,t}\right]=0$, $\operatorname{E}\left[ \eta_{i,t}^2\right]= 1$ and $\operatorname{E}\left[ |\eta_{i,t}|^8\right]< M$ for some $M < \infty$. • $\sigma_{i}$ is defined over an interval $[\delta;M]$ with $0<\delta<M<\infty$. It is independently and identically distributed across $i$, and $\sigma_{i}$ is independent of $\eta_{i,t}$ for all $i$ and $t$. • The $m\times 1$ random vector $\bm e_{i,t}$ is independently distributed across both $i$ and $t$ with $\operatorname{E}\left[\bm e_{i,t}\right]= \boldsymbol{0}$, $\operatorname{E}\left[ \bm e_{i,t}\bm e_{i,t}'\right]= \bm \varSigma$ with the latter being a positive definite matrix and $\operatorname{E}\left[ ||\bm e_{i,t}||^8\right]< M$. \end{enumerate}
assumption$\bm f_t$ is a covariance stationary $r\times 1$ random vector with positive definite covariance matrix $\bm \varSigma_{\bm F}$, absolutely summable autocovariances and $\operatorname{E}\left[ \| \bm f_t \|^4 \right]< M$ .
assumption$\bm \lambda_{i}$ is iid over $i$ with $\operatorname{E}[\bm \lambda_{i}]=\bm \mu_{\bm \lambda}$ and $\operatorname{E}\left[ || \bm \lambda_{i} ||^4\right]< M$. Furthermore, $\bm \varLambda_i$ is iid over $i$ with $\operatorname{E}\left[\bm \varLambda_i\right] = \bm \mu_{\bm \varLambda}$ and $\operatorname{E}\left[ || \bm \varLambda_i ||^4\right]< M$.
assumption$\bm f_t$, $\{\bm \lambda_{i}, \bm \varLambda_i, \sigma_{i} \}$, $\eta_{i',t'}$ and $\bm e_{i'',t''}$ are mutually independent for all $i,i',i'',t,t'$ and $t''$
assumption$\operatorname{rk}\left(\left[ \bm \mu_{\bm \lambda} , \bm \mu_{\bm \varLambda} \right]\right) = r = m+1$.

This set of assumptions above are a slightly more restrictive version of the framework considered in ECTA:ECTA692. For example, the assumption of common $\bm \varSigma$ can be straightforwardly relaxed. However, unlike $\sigma_{i}$, $\bm \varSigma$ plays no major role for asymptotic results of this paper. Most importantly, we rule out the presence of serial correlation as this is in line with the assumptions made for the CD test to work. Furthermore, we restrict ourselves to a classical panel data regression model instead of considering heterogeneous slope coefficients. Moreover, any dependence between $\varepsilon_{i,t}$ and $\bm e_{i,t}$ is assumed away in order to allow for a tractable proof of the main theoretical result in this paper. Lastly, the fact that we assume $\operatorname{rk}\left( \left[\bm \mu_{\bm \lambda} , \bm \mu_{\bm \varLambda}\right] \right) = r = m+1$ to hold, suggest that our we consider an ideal setup where none of the rank condition related problems documented in KaraReeseWesterlund2015, or JuodisEtal2017 apply.

In analogy with the previous sections, we assume that $\bm \beta_{0}$ is known. Thus we disregard the estimation error $\widehat{\bm \beta}^{CCE} - \bm \beta_{0}$, which is generally of order $\operatorname{\mathcal{O}}_{P}((NT)^{-1/2}) + \operatorname{\mathcal{O}}_P(N^{-1})$, see e.g. ECTA:ECTA692 and JuodisEtal2017 .

We begin our asymptotic analysis, by noting that in the model with assumed (known) homogeneous $\sigma$ the result follows directly as in the model with time effects only. In particular, while it is not generally emphasized in the CCE literature, the residuals from CCE estimation which are formally given by

align[align omitted — 229 chars of source]

satisfy

equation[equation omitted — 60 chars of source]

In that respect, the standard 2WFE estimator is similar to the CCE estimator. More formally we formulate the following result

propositionUnder Assumptions (ref)--(ref) and $P(\sigma_{i}=\sigma)=1$: \begin{align} CD &= -\sqrt{\left(T-\frac{T}{N}\right)/2} +\operatorname{\mathcal{O}}_{P}(R_{N,T}), \end{align} where $R_{N,T}=(N^{-1/2}\vee T^{-1/2}\vee N^{-1}\sqrt{T}\vee T^{-1}\sqrt{N})$.

Proposition (ref) shows that the result we derived previously for the 2WFE estimator in Corollary (ref) continues to hold for models with a factor error structure, as long as $N\to\infty$ and $T\to\infty$.

It is worth mentioning that the above order effect is only valid if $\widehat{\bm F}$ contains cross-sectional averages of the regressand as well as all regressors. If either of those variables is omitted (without affecting the rank condition in Assumption (ref)), the zero mean residual condition in (ref) is violated, and consequently the $CD$ test will have a $\operatorname{\mathcal{O}}_{P}(1)$ term. However, this result is of limited empirical importance as in most cases researchers include all available cross-sectional averages. While this practice ensures that the estimator is invariant to $\bm \beta_{0}$, inclusion of too many cross-sectional averages can potentially have detrimental effects on the asymptotic properties of the estimator, see the corresponding discussion in JuodisEtal2017 and JuodisRCCE2020.

In the homoscedastic case two-way fixed effects and multifactor error models have similar asymptotic effects on CD test statistic. The conclusions of the next theorem, which is the main result of this paper, indicate that this equivalence does not hold in the heteroscedastic case. In order to proceed, we introduce some useful notation. For $t=1,\ldots,T$, let $\bm u_{i,t}=[\varepsilon_{i,t}+\bm \beta'\bm e_{i,t},\bm e_{i,t}']'$, and

equation[equation omitted — 146 chars of source]

where $\bm \psi_{i,t}=\left(\overline{\bm C}'\right)^{-1}\bm u_{i,t}$ is the influence function of the corresponding factor estimator, in this case cross-section averages of $y_{i,t}$ and $\bm x_{i,t}$. Generally, the influence function depends on the joint process $\left[y_{i,t},\bm x_{i,t}\right]'$, as long as all observed variables are used to form cross-sectional averages. Equipped with this notation we formulate the main result of this paper.

theoremSuppose that $\bm \beta$ is known. Under Assumptions (ref)--(ref), \begin{align} CD &= \sqrt{\frac{2}{TN(N-1)}}\sum_{i=2}^{N}\sum_{j=1}^{i-1}\sum_{t=1}^{T}\xi_{iN,t}\xi_{jN,t} + \sqrt{T} \Phi_1 -2\sqrt{T} \Phi_2+ \operatorname{\mathcal{O}}_{P}(R_{N,T}), \end{align} where \begin{equation} \xi_{iN,t}=\sigma_{i}^{-1}\varepsilon_{i,t}-\left(\frac{1}{N}\sum_{i=1}^{N}\sigma_{i}^{-1}\bm \lambda_{i}\right)'\bm \psi_{i,t}, \end{equation} and \begin{align} \Phi_1 &= \sqrt{\frac{N}{2(N-1)}} \left(\frac{1}{N}\sum_{i=1}^{N}\sigma_{i}^{-1}\bm \lambda_{i}\right)' \left( \frac{1}{NT} \sum_{i=1}^{N}\sum_{t=1}^{T}\bm \psi_{i,t}\bm \psi_{i,t}'\right) \left(\frac{1}{N}\sum_{i=1}^{N}\sigma_{i}^{-1}\bm \lambda_{i}\right), \\ \Phi_2 &= \sqrt{\frac{N}{2(N-1)}} \left(\frac{1}{N}\sum_{i=1}^{N}\sigma_{i}^{-1}\bm \lambda_{i}\right)' \left(\frac{1}{NT} \sum_{i=1}^{N}\sum_{t=1}^{T}\bm \psi_{i,t}\sigma_{i}^{-1}\varepsilon_{i,t} \right). \end{align}

Here, as previously, $R_{N,T}=(N^{-1/2}\vee T^{-1/2}\vee N^{-1}\sqrt{T} \vee T^{-1}\sqrt{N})$. In line with all previous results, the CD test statistic in this setup has two diverging components. However, unlike all previous results, in particular Theorem (ref), these bias terms are not solely non-linear functions of $\sigma_{i}$. Instead, they are also influenced by the first (rescaled) moments of factor loadings in $y_{i,t}$ and $\bm x_{i,t}$, as well as corresponding variances of the idiosyncratic components in $\bm x_{i,t}$. Thus the influence function $\bm \psi_{i,t}$ directly alters distributional properties of the CD statistic. The result is qualitatively similar to any parametric two-step estimation procedure with a plug-in first-step estimator. Thus by including cross-sectional averages as factor proxies, one is implicitly testing that both $\varepsilon_{i,t}$ and $\bm \psi_{i,t}$ are jointly cross-sectionally uncorrelated. Notice, that while $\varepsilon_{i,t}$ is usually assumed to be uncorrelated over $t$ (e.g. for models with pre-determined regressors), the same is not true for $\bm u_{i,t}$, e.g. if $x_{i,t}=y_{i,t-1}$. Thus if Assumption (ref) is appropriately relaxed, then $\xi_{iN,t}$ will be serially correlated.

remarkOne can easily see that expressions for $\xi_{iN,t},\Phi_{1},\Phi_{2}$ are the same if one assumes that $\lambda_{i}$ is known . Thus it is only the influence function of common factors, and not those of factor loadings, that has an impact on the asymptotic properties of the test statistic. This conclusion is the same as in the two-way error components model.

As alluded in Proposition (ref), the leading $\operatorname{\mathcal{O}}_{P}(1)$ term is degenerate if $\sigma_{i}=\sigma$, as in this case:

equation[equation omitted — 70 chars of source]

One can easily see that the expressions for $\Phi_{1}$ and $\Phi_{2}$ derived for CCE coincide with the corresponding terms of Theorem (ref) (up to a negligible remainder term), upon setting $\bm \lambda_{i}=1$ (thus $\overline{\bm C}^{-1}=1$), and $\bm \psi_{i,t}=\varepsilon_{i,t}$. This situation is equivalent to using $\widehat{f}_{t}=\overline{y}_{t}-\overline{\bm x}_{t}'\bm \beta_{0}$ as factor proxies, e.g. as suggested by JAE:JAE2535 in the context of predictability testing with cross-section dependence.

Recall that result in Theorem (ref) considers only the CCE setup where the number of observable factor proxies equals the number of the true factors. If Assumption (ref) is relaxed and there are more observables than factors, then following KaraReeseWesterlund2015 and JuodisEtal2017, one can show that the expressions for $\Phi_1$ and $\Phi_2$ will contain additional terms related to this discrepancy. In particular, $\Phi_{1}$, $\Phi_{2}$, and $\xi_{iN,t}$ are functions of an unknown rotation matrix, which cannot be consistently estimated from the data. However, as already emphasized by JuodisEtal2017, this problem can be circumvented via the use of a non-parametric bias-correction method. We will come back to this issue in Section (ref).

Re-establishing standard normal inference

The diverging bias in the CD test statistic applied to residual from models with common, period-specific parameters is fundamentally problematic for its use in the context of testing for remaining cross-section correlation. Still, the literature review in Section (ref) suggests that the underlying question of correct model specification is of high relevance for empirical researchers. For this reason it is relevant not to discard completely the CD test statistic but instead to discuss possible modifications aimed at ensuring an asymptotically standard normal inference under null hypothesis. Thus, methods aimed at addressing the IPP detailed in Sections (ref) and (ref) need to be considered.

Analytic bias correction

Given that parametric expressions for the bias of $CD$ have been derived in Theorems (ref) and (ref), analytic bias correction is feasible. As detailed in Section (ref) in the Supplementary Appendix, estimates of the unknown model components that constitute either $\Xi$ or $\Phi_1$ and $\Phi_2$ can be obtained in order to eliminate the diverging component in $CD$. In addition, Propositions (ref) and (ref) in the Supplementary Appendix suggest that these estimated bias terms would eliminate equivalent terms $\Xi_{\mathbb{H}_1}$ or $\Phi_{\mathbb{H}_1,1}$ and $\Phi_{\mathbb{H}_1,2}$ under the alternative hypothesis without necessarily reducing the rate at which $CD$ diverges.

Still, problems with its implementation in practice lead us to discard analytic bias correction and to consider it merely as a benchmark approach when evaluating our weighted CD test statistic, as introduced below, in Monte Carlo experiments. In particular, the construction of plug-in estimates of $\Phi_1$ and $\Phi_2$ proves to be tedious since both terms are constructed from estimates of $\sigma_{1}^2, \ldots, \sigma_{N}^{2}$, $\bm \lambda_1, \ldots, \bm \lambda_N$ and $\bm \beta$. Additionally, estimation of the cross-section sum $N^{-1} \sum_{i=1}^N \sigma_{i}^{-1} \bm \lambda_{i}$ poses a trade-off between generality and accuracy of the plug-in estimate. To be specific, independence between factor loadings $\bm \lambda_i$ and error variances $\sigma_{i}^2$ needs to be imposed to ensure that estimation error in $\widehat{\bm \lambda}_i$ and $\widehat{\sigma}_{i}^2$ does not dominate the small-sample distribution of the bias-corrected version of $CD$.

The need to impose further assumptions in order to improve the accuracy of the bias estimates is effectively a consequence of the fact that the unknown bias itself diverges at rate $\sqrt{T}$. This means that any error in the estimation of the two aforementioned components is scaled by a factor that increases as $N,T \to \infty$. Still, despite improving the approximation of plug-in estimates for $\Phi_1$ and $\Phi_2$ the assumption of independence between factor loadings and error variances can lead to size distortions in DGPs where it is violated, see Section (ref) for a numerical illustration of the problem.

Weighted cross-section covariances

The method for bias correction favored in this article is to construct a CD test statistic from estimated cross-section covariances that are weighted with individual-specific, random draws $X$ from a Rademacher distribution, i.e. $P(X=1)=P(X=-1)=0.5$. This approach is based on noticing that the cross-section correlation estimator $\widehat{\rho}_{ij}$, as defined in (ref) with $z_{i,t} = \widehat{\varepsilon}_{i,t}$, is merely a unit-specific weighting of $\widehat{\operatorname{cov}} \left[\widehat{\varepsilon}_{i,t}\widehat{\varepsilon}_{j,t}\right]$ with advantageous properties: Under the assumptions made in CDtest2004, it ensures that the CD test statistic has unit variance under the null hypothesis of no cross-section correlation, allowing for standard normal inference without the need to obtain a variance estimate. However, the case is different in models with unknown, period-specific parameters $\tau_{t}$ or $\bm f_{t}$. As shown in Theorems (ref) and (ref), studentization of the model residuals fails to ensure a test statistic with a variance of one, even when the true error variances are known. This justifies the use of an alternative weighting scheme which we construct as a weighted average of individual specific covariances,

align[align omitted — 178 chars of source]

for some set of weights $w_1, \ldots, w_N$. Notice that Eq. (ref) would coincide with the CD test statistic of CDtest2004 for $w_i = \widehat{\sigma}_{i}^{-1} \; \forall i$.

This is not the case we consider here. Instead we make the following additional assumption:

assumption[Weights] $w_{1}, \ldots, w_{N}$ are identically and independently Rademacher distributed. Furthermore, $w_{i}$ is independent of $\bm \lambda_{i}$, $\sigma_{i}$, $\eta_{i,t}$ and $\bm e_{i,t}$ for all $i$ and $t$.

Rademacher distributed weights amount to random sample splitting, a method for breaking dependence that is not new to econometrics. For example, altonji1996 already considered it an old concept. Its effects are most obvious in a simple modification of Theorem (ref) where we replace $\sigma_{i}^{-1}$ with $w_i$. Under Assumption (ref), the expected value of $N^{-1}\sum_{i=1}^N \bm \lambda_{i} w_{i}$ can be conveniently split into the expectations of its two constituents. Given the zero expected value of $w_i$, an appropriate LLN applies and leads the cross-section average of $ \bm \lambda_{i} w_{i}$ to converge to zero. In an even simpler form the same reasoning applies to the cross-section averages involving $w_i$ in an adapted version of Theorem (ref). In both cases, this reduces the order of the leading bias components $\Xi$, $\Phi_{1}$ and $\Phi_{2}$ by a factor of $N$ and entails asymptotic unbiasedness of the weighted CD test statistic, subject to the restriction $\sqrt{T} N^{-1} \to 0$. Additional rescaling with the asymptotic variance of expression (ref) results in a weighted CD test statistic which, as formally stated in Theorem (ref) below, converges to a standard normal distribution.

theoremConsider the weighted CD test statistic \begin{align} CD_W &= \left( \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T \widehat{\varepsilon}_{i,t}^2 w_i^2 \right)^{-1} \left(\sqrt{\frac{2}{TN(N-1)}}\sum_{t=1}^{T}\sum_{i=2}^{N} \sum_{j=1}^{i-1}w_{i}\widehat{\varepsilon}_{i,t} w_{j}\widehat{\varepsilon}_{j,t} \right) \end{align} and assume that either of the following two points hold. \begin{enumerate} • The data are generated by the time fixed effects model (ref) such that Assumption (ref) holds. $\widehat{\varepsilon}_{i,t}$ is given by (ref). • The data are generated by the latent common factor model (ref) such that Assumptions (ref)-(ref) hold. $\widehat{\varepsilon}_{i,t}$ is defined by (ref). \end{enumerate} Under either of the two sets of assumptions above as well as Assumption (ref), it holds that \begin{equation} CD_{W}\stackrel{d}{\longrightarrow} N(0,1), \end{equation} as $N,T\to \infty$ jointly subject to the restriction $\sqrt{T}N^{-1} \to 0$.

As stated by Theorem (ref), the use of independent Rademacher distributed weights, analogous to weights drawn from many other distributions with zero mean, re-establishes asymptotic standard normal inference under the null hypothesis of the CD test statistic. However, asymptotic unbiasedness of $CD_W$ comes at the cost of power. More specifically, our approach to bias correction centers the leading components of $CD_W$ around zero, irrespective of whether cross-section correlation in the data is completely controlled for or not. As a consequence, only increases in the variance of $CD_W$ under its alternative hypothesis lead to power against the null hypothesis of this test. We further discuss this point in Section (ref) of the Supplementary Appendix.

To improve the power properties of $CD_W$ we suggest a refinement of this test which follows the power enhancement approach of fan2015. The authors suggest improving the power of high-dimensional cross-sectional tests by adding to the test statistic of interest a screening statistic. This screening statistic is equal to zero with probability approaching one under the null hypothesis of the test, but diverges at a fast rate under the alternative. In our case we choose the absolute sum of thresholded cross-section correlation coefficients. This results in a power-enhanced weighted CD test statistic which is defined as

align[align omitted — 193 chars of source]

where $\widehat{\rho}_{ij}$ is as in (ref) with $z_{i,t}=\widehat{\varepsilon}_{i,t}$ and where $\mathbf{1}(A)$ is the indicator function for event $A$. The screening statistic on the right-hand side of (ref) has an asymptotically negligible effect on the size of $CD_{W+}$ because individual cross-section correlation coefficients are still consistent for their true value under $\mathbb{H}_0$. For additional discussion on $CD_{W+}$ we refer to Section (ref) of the Supplementary Appendix.

remarkAn approach to reducing the dependence of $CD_W$ on a specific set of random weights would consist of averaging several weighted CD test statistics. Denote by $CD_W^{(g)}$ the CD test statistic obtained for a given draw of $N$ Rademacher distributed weights, the latter being indexed by $g$. For a total of $G$ different draws, an averaged weighted CD test can be constructed as \begin{align} \overline{CD}_W = \frac{1}{\sqrt{G}} \sum_{g=1}^{G} CD_W^{(g)}, \end{align} where the total number of draws $G$ should be chosen sufficiently small (e.g. $G=30$) to avoid size distortions that may arise from scaling up lower-order terms in $CD_W^{(g)}$.
remarkAn initially appealing alternative to external random numbers as weights for a bias-corrected CD test statistic would be a set of $N$ statistics derived from the dataset available to the researcher. When considering such internal (sample-specific) weights, it is desirable to opt for functions of the data that are optimal in the sense that they maximize the rate at which a weighted CD test statistic diverges for a wide number of alternatives. In Section (ref) of the Supplementary Appendix we argue that such improved weights suggest the use of a different test statistic for each alternative. Hence, as far as a bias-corrected version of $CD$ is preferred to the use of a different test, obtaining the former by weighting cross-section covariances with external random weights is the best possible alternative.

Monte Carlo study

We investigate the properties of the CD test as well as its weighted alternatives in a small set of simulation experiments. We consider the common factor model

eqnarray[eqnarray omitted — 224 chars of source]

We consider two alternative specifications for the factor loadings on the regressand $y_{i,t}$. We define $\bm \lambda_{i} = \bm \iota_r + \tilde{\bm \lambda}_{i}$ where the $r$ elements of $\tilde{\bm \lambda}_{i}$ are drawn (I) from $U(-0.75, 0.75)$ or (II) from a standardized $\chi ^{2}\left(2\right) $ distribution that has zero mean and a variance of 1/6. The latter case is designed to match the first two moments of $\tilde{\bm \lambda}_{i}$ in case (I). Concerning the factor loadings on the regressors $x_{i,t}$ we draw element $\Lambda^{(r')}$ of the $r\times 1$ vector $\bm \varLambda_{i}$ from $U(-0.5,0.5)$ if $r'=1$ and from $U(0.5,1,5)$ otherwise. Concerning the latent factors, we let $\bm f_{t}\sim N\left( \boldsymbol{0}_{r}, \bm I_r\right) $. When considering the two-way fixed effects model instead of the common latent factor model, we set $r=2$ and define $\bm f_t = \left[f_t^{(1)}, 1 \right]'$ as well as $\bm \lambda_{i}= \left[1, \lambda_{i}^{(2)} \right]'$ and $\bm \varLambda_{i}= \left[1, \Lambda_{i}^{(2)} \right]'$ with the construction of $f_t^{(1)}, \lambda_{i}^{(2)}$ and $\Lambda_{i}^{(2)}$ being unchanged.

Analogous to our treatment of factor loadings we consider two cases for $\varepsilon_{i,t}$, namely that they are drawn independently over $i$ and $t$ from (i) a standard normal distribution or (ii) a standardized $\chi ^{2}\left(2\right) $ distribution that has zero mean and unit variance. The scalar random variable $e_{i,t}$ is generated in the same way as $\varepsilon_{i,t}$. Lastly, error variances are obtained in two different fashions: We set (a) $\sigma _{i}^{2}=c_{\sigma }\left( \varsigma _{i}^{2}-2\right) / \sqrt{24} + 1$ where $\varsigma _{i}^{2}\sim\chi ^{2}\left( 2\right)$ or (b) $\sigma _{i}^{2} = 0.5 + d_{\sigma} T^{-1} \sum_{t=1}^T (\bm \lambda_{i}' \bm f_t)^2$. The scaling $c_{\sigma }$, which is only included in setup (a), is chosen manually as discussed below. This construction entails that $\operatorname{E}\left[ \sigma_{i}^{2}\right] =1$ and $\operatorname{var}\left[ \sigma _{i}^{2}\right] =c_{\sigma }^{2}/6$. The normalizing constant $d_{\sigma}$, which appears in case (b) only, is $d_{\sigma} = \sqrt{2}$ in model error case (i) and $d_{\sigma} = (1/3)^{-1/2}$ in case (b). This scaling ensures that the variance of $\sigma _{i}^{2}$ across cross-sections is approximately of the same magnitude as in case (a) with $c_{\sigma}=1$.

In all following simulations we use feasible CD statistics, where, unlike the theory part of this paper, we treat $\beta$ as an unknown parameter.

Standard CD statistic

In a first instance, we consider the properties of the original CD test statistic when applied to either 2WFE or CCE residuals in a scenario where all sources of cross-section correlation are adequately accounted for. We do this for the setup (i)-(a)-(I) as described above, meaning that errors are normally distributed, error variances independent of factor loadings (or individual fixed effects) and that the latter are symmetrically distributed around their mean. Results for a model with errors from a standardized $\chi^2$ distribution are almost identical and reported in Section (ref) of the Supplementary Appendix. Furthermore, we consider different degrees of heterogeneity among the individual-specific error variances by reporting results for the values $c_{\sigma }\in \left\{ 0.1,0.5,1,1.5\right\}$.

Table (ref) reports the first two moments of $CD$ when applied either to 2WFE residuals when the true model is one with two-way fixed effects or to CCE residuals in a model with multifactor error structure and two factors. The results for both cases are identical and show that $CD$ has a bias that diverges towards $-\infty $ as $T\rightarrow \infty $ and a variance that is considerably below $1$. The bias term is reasonably close to the benchmark value of $-\sqrt{T/2}$ indicated by Proposition (ref) but tends towards zero as heterogeneity among the individual-specific error variances is amplified.

table[table omitted — 2,968 chars of source]

Given that the theoretical results of Sections (ref) and (ref) are confirmed by Table (ref), we turn to the properties of $CD$ under its alternative hypothesis. For this purpose, we simulate a model with multifactor error structure and three factors, implying that neither 2WFE nor CCE estimation completely accounts for all sources of cross-section correlation in the simulated data. Heterogeneity among unit-specific error variances is kept constant by only considering $c_{\sigma}=1$. Instead we consider cases (I) and (II) for factor loadings as well as cases (a) and (b) for error variances. Again, results are presented only for normally distributed model errors. The moments of $CD$ are very similar when model errors are drawn from a standardized $\chi^2(2)$ distribution and corresponding tables can be found in Section (ref) of the Supplementary Appendix.

Tables (ref) and (ref) report the mean and variance of $CD$ in the presence of remaining cross-section correlation. Both are very similar to the numbers reported in Table (ref) as long as factor loadings are symmetrically distributed, an observation which is in line with the theoretical results provided in Section (ref) of the Supplementary Appendix. Since symmetric loadings center the leading stochastic component of $CD$ around zero, the latter is dominated by a bias equivalent that diverges towards $-\infty$ as $T\to \infty$. Qualitatively very different results can be observed when factor loadings are drawn from a skewed distribution, particularly when applying the CD test statistic to 2WFE residuals. As one can observe in columns two and four of Table (ref), the strong negative divergence of $CD$ as $T$ increases is mitigated and it is reasonable to assume that it will be turned into positive divergence for $N>200$. This pattern, which is amplified if error variances are a function of factor loadings, reflects the presence of a nonzero mean in the leading stochastic component of $CD$ that diverges to $+\infty$ at rate $N\sqrt{T}$. In the case of 2WFE residuals, its magnitude is sufficiently large to counteract the negative bias equivalent already for $N=200$. By contrast, columns two and four in Table (ref) suggest that this term is considerably smaller when the CD test is applied to CCE residuals since the mean of $CD$ continues to diverge towards $-\infty$ for $N=200$. We conjecture that the same sign reversal will eventually be obtained even if this case, but that the required cross-section dimension for this to happen is higher.

table[table omitted — 3,203 chars of source]
table[table omitted — 2,434 chars of source]

Weighted CD statistic

Next, we proceed with investigating the properties of our weighted CD test statistic. As previously, we keep heterogeneity across unit-specific error variances fixed at $c_{\sigma}=1$ and consider two different specifications each for factor loadings and error variances. We report size and power for our weighted test statistic $CD_W$ as well as the power-enhanced refinement $CD_{W+}$. As a benchmark test, we include a CD test statistic with analytic bias correction which corrects $CD$ with sample equivalents of the asymptotic bias terms in Theorems (ref) and (ref). Details on its implementation can be found in Section (ref) of the Supplementary Appendix. The original CD test is left out for the sake of saving space and in particular since rejection rates are $100\%$ for most cases.

Table (ref) reports the size and power of all three bias corrected test statistic when applied to 2WFE residuals. The size of $CD_W$ and $CD_{W+}$ is very close to the nominal level of $5\%$ as long as $N\geq T$. Size distortions are given in cases where $T$ is considerably large than $N$ and the effect of power enhancement on size is generally negligible. The analytically bias-corrected CD test statistic $CD_{BC}$ exhibits hardly any tendency to over-reject but is rather conservative, in particular when $T$ is small.

Panel B in Table (ref) report results on power. Here, we see that the power of $CD_W$ is in general low and increases only in $T$. This is improved upon considerably by power enhancement, leading the refined test statistic $CD_{W+}$ to reliably reject when it should as long as the number of time periods is large enough. The performance of $CD_{W+}$ is considerably above that of the benchmark statistic $CD_{BC}$ when factor loadings are drawn from a symmetric distribution and is on par in most other cases. An exception is the case of skewed loadings with small $T$ and large $N$ where $CD_{BC}$ has markedly higher rejection rates. It can also be noted that the performance of $CD_{BC}$ largely depends on whether factor loadings are drawn from a symmetric distribution or not. If this is the case, analytic bias correction leads to rejection rates under the alternative hypothesis that are worst among all three tests considered.

When testing for remaining cross-section correlation in CCE residuals we observe similar results as in Table (ref). Interestingly, it can be observed that the benchmark statistic $CD_{BC}$ exhibits size distortions in samples with large $T$ when loadings are drawn from a skewed distribution and if factor loadings and error variances are dependent. This results from the nature of our bias correction which assumes independence between $\sigma^{2}_i$ and $\bm \lambda{i}$ to considerably improve the accuracy of the bias estimate.

The power properties of $CD_W$ and $CD_{W+}$ when either of these tests is applied to CCE residuals mirror those seen in the 2WFE case, even though rejection rates are generally somewhat lower. However, the power-enhanced statistic $CD_{W+}$ now performs best in all cases without ever being inferior to $CD_{BC}$.

table[table omitted — 6,124 chars of source]
table[table omitted — 5,970 chars of source]

Empirical illustration: R& D investments

In this section we illustrate the applicability of the standard and the weighted CD statistics using the R&D investments data of doi:10.1162/REST00272. Information on up to twelve manufacturing industries in ten countries (Denmark, Finland, Germany, Italy, Japan, The Netherlands, Portugal, Sweden, United Kingdom, and The US) over a time period from 1980 to 2005 was used to construct the dataset. After some minor modification to the original dataset (see Section (ref) of the Supplementary Appendix) we are left with with a panel dataset of $N=82$ and $T=25$.

In this application serial correlation is important from an economic point of view. In Section (ref) of the Supplementary Appendix we outline how the CD statistic can be modified to account for serial correlation, given a set of known weights $w_{i}$.

doi:10.1162/REST00272 question whether R&D can be estimated in a standard Griliches-type “knowledge production function” framework ignoring the potential presence of knowledge spillovers between cross-sectional units as well as other cross-section dependencies. Among other things they document a “[\ldots] strong evidence for cross-sectional dependence and the presence of a common factor structure in the data, which we interpret as indicative for the presence of knowledge spillovers and additional unobserved cross-sectional dependencies” doi:10.1162/REST00272. Cross-sectional dependence was measured by means of a CD test.

In this section we will primarily revisit some of the results in Table 5 of doi:10.1162/REST00272 for pooled (static) production function estimates. Our goal is to investigate how the divergent properties of the CD test statistic might have influenced the choice between the First Difference (FD) estimator with yearly dummies and Pooled CCE (CCEP) estimators, as presented in columns 3 and 4, respectively, of the table mentioned above. Which of the two models is considered to be correctly specified has important consequences for the conclusions that can be drawn from the entire table. With regard to the coefficient of private R&D investments, doi:10.1162/REST00272 report that a significance test for the corresponding slope coefficient cannot reject the null hypothesis in the FD model while it can in the CCE model.

table[table omitted — 1,142 chars of source]

As the results for $CD_{W}$ statistic allowing for serial-correlation correction, are similar to those without serial correlation, we only focus on the latter option. First of all, as we can see from Table (ref) the conclusions that we can draw from the original CD test are almost identical to those doi:10.1162/REST00272, despite adjustments made in terms of the sample size. In particular, while the value of CD statistic based on CCE residuals imply rejection of the null hypothesis, no such conclusion is implied by FD (at least at a $5\%$ significance level). However, as given that the original test statistic suffers from the IPP, we further investigate if this conclusion also holds after proper bias-corrections. Motivated by finite sample evidence in Section (ref), we consider only the $CD_{W}$ statistic.

Taking our attention to the proposed $CD_{W}$ statistic based on random Rademacher weights, we can see that the corresponding values of $\overline{CD}_{W}$ indicate that both the FD and CCEP models generate residuals that are cross-sectionally uncorrelated. As we discussed in Section (ref), this conclusion might be partially attributed to the fact that for small values of $T$, the proposed statistic might lack power. For this reason, we also make use of the power-enhanced version $\overline{CD}_{W+}$ of our test statistic. Notice that for this dataset $2\sqrt{\ln(N)/T}\approx 0.84$, thus power-enhancement will pick up only very large values of the correlation coefficients that otherwise might be averaged away by the original CD statistic. As we can see from the corresponding column in Table (ref), power enhancement does not alter the conclusion we draw for the FD residuals, as there are only 2 coefficients $\widehat{\rho}_{ij}$ above the threshold. However, for CCEP residuals the effect of power enhancement is non-negligible as there are 10 values of $\widehat{\rho}_{ij}$ above the threshold.

These results indicate that after appropriate adjustments, the empirical evidence presented in doi:10.1162/REST00272 that favor simple First-difference estimator, as opposed to the CCEP estimator, remain in place. As we can see from Table (ref), this conclusion is unaltered after appropriate serial-correlation adjustments, thus serial-correlation cannot be the main factor affecting this ranking (unlike the discussion in doi:10.1162/REST00272). Instead, our results might indicate that the underlying CCE restrictions (rank condition and regressors with finite factor structure) might be violated, as these restrictions are irrelevant for the FD estimator with time-effects only.

Conclusion

This article documents how the estimation of common time-specific parameters using panel data causes the CD test of CDtest2004, doi:10.1080/07474938.2014.956623 to break down. Using popular additively and multiplicatively interacted specifications for individual- and time-specific components in the model errors, we show that the CD test statistic applied to residuals of correctly specified regression models is divergent under null hypothesis of cross-section independence. We can find an equivalent term under the alternative hypothesis which may balance out the leading diverging component of $CD$ and can lead to low power in small samples. The results documented in this article are interpreted as a manifestation of the incidental parameter problem (IPP) since they ultimately follow from the estimation of $\operatorname{\mathcal{O}}(T)$ period-specific parameters. Our main theorems illustrate the pervasive nature of the IPP in this setup, given that the consequences of estimating time specific parameters do not disappear as the sample size increases.

Our proposed weighted CD test statistic achieves our primary goal of re-establishing asymptotic standard normal inference and hence constitutes an alternative to popular bias correction methods which circumvents the problems these approaches have in the present context. Our results have far reaching implications for empirical panel data analysis, where CD test has been widely used as a model selection/diagnostic tool. An illustration of how our theoretical results translate into applications is given via simulations and real datasets.

Finally, in this paper we assumed that the parameters in the linear model are estimated using a least squares objective function. If one deviates from this setup, and instead uses an (over-identified) GMM criterion function to estimate parameters, the usual GMM J-statistic is readily available for the purpose of testing residual cross-sectional correlation. Examples in a fixed-$T$ framework are given by Sarafidis2009, AhnEtal2013, and JuodisPseudo2014 among others. In the large $N,T$ setup average J-statistic as a model specification tool was explicitly used e.g. by JAE:JAE2338 for the CCE-GMM estimator. Hence given the scale of the problems with CD statistic documented in this paper, these alternative procedures (if applicable) are more appropriate.

thebibliography{61} \expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi \bibitem[\citeauthoryear{Ahn, Lee, and Schmidt}{Ahn et al.}{2013}]{AhnEtal2013} Ahn, S. C., Y. H. Lee, and P. Schmidt (2013): “Panel Data Models with Multiple Time-varying Individual Effects,” Journal of Econometrics, 174, 1--14. \bibitem[\citeauthoryear{Altonji and Segal}{Altonji and Segal}{1996}]{altonji1996} Altonji, J. G., and L. M. Segal (1996): “Small-sample bias in GMM estimation of covariance structures,” Journal of Business & Economic Statistics, 14, 353--366. \bibitem[\citeauthoryear{Bai}{Bai}{2009}]{ECTA:ECTA954} Bai, J.(2009): “Panel Data Models With Interactive Fixed Effects,” Econometrica, 77, 1229--1279. \bibitem[\citeauthoryear{Bai and Ng}{Bai and Ng}{2004}]{baing2004panic} \textsc{Bai, J. and S. Ng} (2004): “A PANIC Attack on Unit Roots and Cointegration,” \emph{Econometrica}, 72, 1127--1177. \bibitem[\citeauthoryear{Bailey, Holly, and Pesaran}{Bailey et al.}{2016}]{JAE:JAE2468} \textsc{Bailey, N., S. Holly, and M. H. Pesaran} (2016): “A Two-Stage Approach to Spatio-Temporal Analysis with Strong and Weak Cross-Sectional Dependence,” \emph{Journal of Applied Econometrics}, 31, 249--280. \bibitem[\citeauthoryear{Boneva and Linton}{Boneva and Linton}{2017}]{BonevaLinton2016} \textsc{Boneva, L. and O. Linton} (2017): “A Discrete Choice Model for Large Heterogeneous Panels with Interactive Fixed Effects with an Application to the Determinants of Corporate Bond Issuance,” \emph{Journal of Applied Econometrics}, 32, 1226--1243. \bibitem[\citeauthoryear{Chen, Gao, and Li}{Chen et al.}{2012}]{chen_gao_li_2012} \textsc{Chen, J., J. Gao, and D. Li} (2012): “A New Diagnostic Test for Cross-section Uncorrelatedness in Nonparametric Panel Data Models,” \emph{Econometric Theory}, 28, 1144–--1163. \bibitem[\citeauthoryear{Chen, Fern\'{a}ndez-Val, and Weidner}{Chen et al.}{2020}]{chenvalweidner2014} \textsc{Chen, M., I. Fern\'{a}ndez-Val, and M. Weidner} (2020): “Nonlinear Panel Models with Interactive Effects,” Forthcoming \emph{in Journal of Econometrics}. \bibitem[\citeauthoryear{Chudik, Mohaddes, Pesaran, and Raissi}{Chudik et al.}{2017}]{PesaranRestat2017} \textsc{Chudik, A., K. Mohaddes, M. H. Pesaran, and M. Raissi} (2017): “Is There a Debt-Threshold Effect on Output Growth?” \emph{The Review of Economics and Statistics}, 99, 135--150. \bibitem[\citeauthoryear{Demetrescu and Homm}{Demetrescu and Homm}{2016}]{JAE:JAE2496} \textsc{Demetrescu, M. and U. Homm} (2016): “Directed Tests of No Cross-Sectional Correlation in Large-N Panel Data Models,” \emph{Journal of Applied Econometrics}, 31, 4--31. \bibitem[\citeauthoryear{Eberhardt, Helmers, and Strauss}{Eberhardt et al.}{2013}]{doi:10.1162/REST00272} \textsc{Eberhardt, M., C. Helmers, and H. Strauss} (2013): “Do Spillovers Matter When Estimating Private Returns to R&D?” \emph{The Review of Economics and Statistics}, 95, 436--448. \bibitem[\citeauthoryear{Eberhardt and Presbitero}{Eberhardt and Presbitero}{2015}]{EBERHARDT201545} \textsc{Eberhardt, M. and A. F. Presbitero} (2015): “Public Debt and Growth: Heterogeneity and Non-linearity,” \emph{Journal of International Economics}, 97, 45 -- 58. \bibitem[\citeauthoryear{Ertur and Musolesi}{Ertur and Musolesi}{2017}]{Ertur2017} \textsc{Ertur, C. and A. Musolesi} (2017): “Weak and Strong Cross-Sectional Dependence: A Panel Data Analysis of International Technology Diffusion,” \emph{Journal of Applied Econometrics}, 32, 477--503. \bibitem[\citeauthoryear{Everaert and Pozzi}{Everaert and Pozzi}{2014}]{JAE:JAE2338} \textsc{Everaert, G. and L. Pozzi} (2014): “The Predictability of Aggregate Consumption Growth in OECD Countries: A Panel Data Analysis,” \emph{Journal of Applied Econometrics}, 29, 431--453. \bibitem[\citeauthoryear{Fan, Liao and Yao}{Fan et al.}{2015}]{fan2015} \textsc{Fan, J., Y. Liao and J. Yao, } (2015): “Power Enhancement in High-Dimensional Cross-Sectional Tests,” \emph{Econometrica}, 83, 1497--1541. \bibitem[\citeauthoryear{Fern\'{a}ndez-Val and Weidner}{Fern\'{a}ndez-Val and Weidner}{2016}]{fernandez_weidner_2014} \textsc{Fern\'{a}ndez-Val, I. and M. Weidner} (2016): “Individual and Time Effects in Nonlinear Panel Models with Large N, T,” \emph{Journal of Econometrics}, 192, 291 -- 312. \bibitem[\citeauthoryear{Gagliardini, Ossola, and Scaillet}{Gagliardini et al.}{2016}]{ECTA:ECTA1599} \textsc{Gagliardini, P., E. Ossola, and O. Scaillet} (2016): “Time-Varying Risk Premium in Large Cross-Sectional Equity Data Sets,” \emph{Econometrica}, 84, 985--1046. \bibitem[\citeauthoryear{Hahn and Kuersteiner}{Hahn and Kuersteiner}{2002}]{Hahn2002} \textsc{Hahn, J. and G. Kuersteiner} (2002): “Asymptotically Unbiased Inference for a Dynamic Panel Model with Fixed Effects When Both N and T are Large,” \emph{Econometrica}, 70(4), 1639--1657. \bibitem[\citeauthoryear{Hahn and Moon}{Hahn and Moon}{2006}]{hahnmoon2006} \textsc{Hahn, J. and H. R. Moon} (2006): “Reducing Bias of MLE in A Dynamic Panel Model,” \emph{Econometric Theory}, 22, 499--512. \bibitem[\citeauthoryear{Hahn and Newey}{Hahn and Newey}{2004}]{HN2004} \textsc{Hahn, J. and W. Newey} (2004): “Jackknife and Analytical Bias Reduction for Nonlinear Panel Models,” \emph{Econometrica}, 72(4), 1295--1319. \bibitem[\citeauthoryear{Hall and Heyde}{Hall and Heyde}{1980}]{HallHeyde1980} \textsc{Hall, P. and C. C. Heyde} (1980): \emph{Martingale Limit Theory and Its Application}, Probability and Mathematical Statistics, Academic Press. \bibitem[\citeauthoryear{Holly, Pesaran, and Yamagata}{Holly et al.}{2010}]{HOLLY2010160} \textsc{Holly, S., M. H. Pesaran, and T. Yamagata} (2010): “A Spatio-temporal Model of House Prices in the USA,” \emph{Journal of Econometrics}, 158, 160 -- 173. \bibitem[\citeauthoryear{Hsiao, Pesaran, and Pick}{Hsiao et al.}{2012{\natexlab{a}}}]{OBES:OBES646} \textsc{Hsiao, C., M. H. Pesaran, and A. Pick} (2012{\natexlab{a}}): “Diagnostic Tests of Cross-section Independence for Limited Dependent Variable Panel Data Models,” \emph{Oxford Bulletin of Economics and Statistics}, 74, 253--277. \bibitem[\citeauthoryear{Juodis}{Juodis}{2018}]{JuodisPseudo2014} \textsc{Juodis, A.} (2018): “Pseudo Panel Data Models with Cohort Interactive Effects,” \emph{Journal of Business & Economic Statistics}, 36, 47--61. \bibitem[\citeauthoryear{Juodis}{Juodis}{2020}]{JuodisRCCE2020} --------- (2020): “A Regularization Approach to Cross-section Average-Augmented Panel Data Models,” Mimeo. \bibitem[\citeauthoryear{Juodis, Karab{\i}y{\i}k, and Westerlund}{Juodis et al.}{2020}]{JuodisEtal2017} \textsc{Juodis, A., H. Karab{\i}y{\i}k, and J. Westerlund} (2020): “On the Robustness of the Pooled CCE Estimator,” forthcoming in \emph{Journal of Econometrics}. \bibitem[\citeauthoryear{Karab{\i}y{\i}k, Reese, and Westerlund}{Karab{\i}y{\i}k et al.}{2017}]{KaraReeseWesterlund2015} \textsc{Karab{\i}y{\i}k, H., S. Reese, and J. Westerlund} (2017): “On The Role of The Rank Condition in CCE Estimation of Factor-augmented Panel Regressions,” \emph{Journal of Econometrics}, 197, 60--64. \bibitem[\citeauthoryear{Mao}{Mao}{2018}]{mao2018testing} \textsc{Mao, G.} (2018): “Testing for sphericity in a two-way error components panel data model,” \emph{Econometric Reviews}, 37, 491--506. \bibitem[\citeauthoryear{Mastromarco, Serlenga, and Shin}{Mastromarco et al.}{2016}]{JAE:JAE2439} \textsc{Mastromarco, C., L. Serlenga, and Y. Shin} (2016): “Modelling Technical Efficiency in Cross Sectionally Dependent Stochastic Frontier Panels,” \emph{Journal of Applied Econometrics}, 31, 281--297. \bibitem[\citeauthoryear{Neyman and Scott}{Neyman and Scott}{1948}]{neyman1948} \textsc{Neyman, J. and E. L. Scott} (1948): “Consistent Estimation from Partially Consistent Observations,” \emph{Econometrica}, 16, 1--32. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2004}]{CDtest2004} \textsc{Pesaran, M. H.} (2004): “General Diagnostic Tests for Cross Section Dependence in Panels,” CESifo Working Paper No. 1229. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2006}]{ECTA:ECTA692} --------- (2006): “Estimation and Inference in Large Heterogeneous Panels with a Multifactor Error Structure,” \emph{Econometrica}, 74, 967--1012. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2015{\natexlab{a}}}]{doi:10.1080/07474938.2014.956623} --------- (2015{\natexlab{a}}): “Testing Weak Cross-Sectional Dependence in Large Panels,” \emph{Econometric Reviews}, 34, 1089--1117. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2015{\natexlab{b}}}]{PesaranBook} --------- (2015{\natexlab{b}}): \emph{Time Series and Panel Data Econometrics}, Oxford University Press. \bibitem[\citeauthoryear{Phillips and Moon}{Phillips and Moon}{1999}]{phillipsMoon1999} \textsc{Phillips, P. C. B. and H. R. Moon} (1999): “Linear Regression Limit Theory for Nonstationary Panel Data,” \emph{Econometrica}, 67, 1057--1111. \bibitem[\citeauthoryear{Reese and Westerlund}{Reese and Westerlund}{2016}]{reese2016panicca} \textsc{Reese, S. and J. Westerlund} (2016): “Panicca: Panic on Cross-Section Averages,” \emph{Journal of Applied Econometrics}, 31, 961--981. \bibitem[\citeauthoryear{Sarafidis, Yamagata, and Robertson}{Sarafidis et al.}{2009}]{Sarafidis2009} \textsc{Sarafidis, V., T. Yamagata, and D. Robertson} (2009): “A Test of Cross Section Dependence for a Linear Dynamic Panel Model with Regressors,” \emph{Journal of Econometrics}, 148, 149--161. \bibitem[\citeauthoryear{Westerlund, Karabiyik, and Narayan}{Westerlund et al.}{2017}]{JAE:JAE2535} \textsc{Westerlund, J., H. Karabiyik, and P. Narayan} (2017): “Testing for Predictability in Panels with General Predictors,” \emph{Journal of Applied Econometrics}, 32, 554--574. \bibitem[\citeauthoryear{Westerlund and Urbain}{Westerlund and Urbain}{2013}]{Westerlund2013247} \textsc{Westerlund, J. and J.-P. Urbain} (2013): “On the Estimation and Inference in Factor-augmented Panel Regressions with Correlated Loadings,” \emph{Economics Letters}, 119, 247 -- 250. \bibitem[\citeauthoryear{Westerlund and Urbain}{Westerlund and Urbain}{2015}]{Westerlund2015372} --------- (2015): “Cross-sectional Averages Versus Principal Components,” \emph{Journal of Econometrics}, 185, 372 -- 377.

\makeatletter\@input{auxsupp.tex}\makeatother