Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,167 characters · 14 sections · 57 citation commands
The Incidental Parameters Problem in Testing for Remaining Cross-section Correlation
\singlespace \bibpunct{(}{)}{,}{a}{,}
\onehalfspace
Given that economic agents rarely act entirely independently of each other, modeling cross-section dependence plays a prominent role in panel data econometrics. Time fixed effects are probably the simplest way of addressing this issue and allow controlling for a common trend whose effect is homogeneous across cross-section units. During the last decade, interactive fixed effects models have become a popular, more general alternative. They assume the presence of a factor error structure, i.e. a small number of unobserved common trends interacted with entity-specific slope coefficients. Using either of the two modeling possibilities begs the question whether cross-section dependence is adequately accounted for.
In this paper, we show that the application of tests for cross-section dependence to regression residuals obtained from two-way fixed effects models or interactive effects models is problematic. We use the popular CD test statistic of CDtest2004, doi:10.1080/07474938.2014.956623 and show that the inclusion of period-specific parameters introduces a bias term of order $\sqrt{T}$. In order to avoid erroneous rejection of the null hypothesis of no unaccounted cross-section dependence, we suggest a modified test statistic that re-weights cross-section covariances with Rademacher distributed weights. This weighted CD test statistic is asymptotically standard normal and has very good size under appropriate regularity conditions on the chosen weights.
The use of the CD test statistic as a misspecification test for regression models that already account for cross-section dependence has repeatedly been observed in empirical studies (see e.g. HOLLY2010160, JAE:JAE2338, JAE:JAE2468 and EBERHARDT201545 among others). In some applications, the CD test statistic is explicitly used as a model-selection tool, interpreting a reduction in the absolute value of the CD test statistic as an indication of a better model. In other cases, only specifications not rejected by the CD test statistic (given some significance level) were considered. However, to the best of our knowledge, there are currently no theoretical results in the literature that could justify this practice. Furthermore, the scenario of testing for cross-section dependence in cross-sectionally demeaned or defactored data is usually completely ignored in any of the large scale Monte Carlo studies integral to theoretical and empirical papers in the field. Only very recently, mao2018testing reported the size of three tests for cross-section dependence in a subset of his simulation experiments. However, despite clear evidence of excessive over-rejections, these results are neither discussed nor investigated theoretically. Therefore, this study is the first one to investigate the properties of cross-section dependence tests applied to residuals of models that control for sources of common variation across cross-sections. In particular, given its popularity in the applied panel data literature, we restrict our attention to the CD test statistic of CDtest2004, doi:10.1080/07474938.2014.956623. Our interest lies in residuals of models that characterize cross-section dependence as driven by a small number of unobservable factors. This comprises both the two-way fixed effects (2WFE) model as well as models with a multifactor error structure where factors are interacted with unit-specific slope coefficients. The results that we obtain are summarized as follows:
The degeneracy of the CD statistics can be seen as a manifestation of the Incidental Parameters Problem (IPP) of neyman1948. In this respect, this paper contributes to this branch of the literature. So far, major focus in the panel data literature has been related to the IPP stemming from estimated individual specific effects. Our paper is the first one to document an asymptotically non-negligible impact of estimated time specific common parameters which cannot be ruled out by restrictions on the relative expansion rate of $N$ and $T$. Furthermore, since the CD statistic can be seen as a time-series average of second degree (degenerate) U-statistics, our results shed some light on the potential impact of the IPP beyond simple cross-sectional averages. Lastly, while this article only considers linear models, the average correlation approach to testing for cross-section dependence was extended to nonlinear and nonparametric panel data models by OBES:OBES646 and chen_gao_li_2012, respectively. Hence, problems documented in this paper carry over to post-estimation properties of non-linear models discussed in e.g. fernandez_weidner_2014, BonevaLinton2016, or chenvalweidner2014.
The remainder of this paper is structured as follows: Section (ref) introduces the testing problem. In Sections (ref) we present asymptotic results for models with two-way fixed effects and multifactor error structures. In Section (ref) we discuss standard approaches for bias correction and propose a weighted CD test statistic that achieves this goal. Sections (ref) and (ref) illustrate the problem documented in this paper by means of simulated and real data. Section (ref) concludes. Additional technical discussions, Monte Carlo studies, and all proofs are relegated to the Supplementary Appendix.
Notation: $\bm I_m$ denotes an $m \times m$ identity matrix and the subscript is sometimes disregarded from for the sake of simplicity. $\boldsymbol{0}$ denotes a vector of zeros while $\mathbf{O}$ stands for a matrix of zeros. $\bm s_m$ denotes a selection vector all of whose elements are zero except for element $m$ which is one. $\bm \iota$ is a vector entirely consisting of ones. The dimension of these latter vectors and matrices is generally suppressed for the sake of simplicity and needs to be inferred from context. For a generic $m \times n$ matrix $\bm A$, $\bm P_{\bm A} = \bm A(\bm A'\bm A)^{-1}\bm A'$ projects onto the space spanned by the columns of $\bm A$ and $\bm M_{\bm A} = \bm I_m - \bm P_{\bm A}$. Furthermore, $\operatorname{rk}(\bm A)$ denotes the rank of $\bm A$, $\operatorname{tr}(\bm A)$ its trace and $\| \bm A \| =\left( \operatorname{tr} (\bm A'\bm A)\right)^{1/2}$ the Frobenius norm of $\bm A$. For a set of $m \times n$ matrices $\{ \bm A_1, \ldots, \bm A_N \}$, $\overline{\bm A} = N^{-1}\sum_{i=1}^N \bm A_i$. $\delta$ and $M$ stand for a small and large positive real number, respectively. For two real numbers $a$ and $b$, $a \vee b = \max\{ a,b\}$. Lastly, $\operatorname{\mathcal{O}}_P(\cdot)$ and $\operatorname{o\mathcal{}}_P(\cdot)$ express stochastic order relations.
Let $\bm z_{i}$ be a $T$-dimensional data vector observed over $N$ cross-sectional units indexed by $i$. Combining all $\bm z_{i}$ we obtain a two-dimensional data array of panel data (or longitudinal data). In empirical research it is common to investigate whether $\bm z_{i}$ can be regarded as independent over $i$ in order to select a model that can properly characterize the statistical properties of the data. In particular, researchers might be interested in the statistical hypothesis
where we use the $\bot$ notation to denote independence. Most often $\bm z_{1}, \ldots, \bm z_N$ contain residuals obtained from a regression model that does not allow for cross-section dependence, e.g. an entity fixed effects model or linear regression.
By far the most widely used test for cross-section dependence (correlation, to be precise) is the CD test of CDtest2004, doi:10.1080/07474938.2014.956623, which is based on a simple rescaled sum of all pairwise cross-section correlation coefficients, formally denoted
Here,
is the pairwise sample correlation coefficient between units $i$ and $j$. Obviously, computing the CD test statistic involves obtaining $N(N-1)/2$ parameter estimates, each of which converges to the true parameter value at rate $\sqrt{T}$ only. These circumstances are reminiscent of the panel data setup considered in e.g. phillipsMoon1999, Hahn2002, where estimation of many (incidental) parameters in linear regression models turns out to have distributional effects on the asymptotic properties of common parameters. That is, they cause the incidental parameter problem. By contrast, the asymptotic distribution of the CD test statistic is unaffected by the estimation of all $N(N-1)/2$ cross-section correlation coefficients involved in its construction. In fact, applied to the residuals of a linear regression model with strictly exogenous regressors, and individual specific means, the CD test statistic is asymptotically normal as long as $N,T\to\infty$ (see PesaranBook, Theorem 2), under $H_{0}$ of independence (or even limited local dependence). However, as shown below, this result does not hold when the model specification includes period-specific parameters.
Consider a linear model where cross-section dependence is due to time fixed effects, so that both unobserved heterogeneity over cross-sections and time enters the model additively. That is the relation between the $T\times 1$ vector $\bm y_i$ and the $T \times m$ matrix $\bm X_i$ is formally denoted
Here, $\mu_i$ and $\bm \tau$ denote an entity-specific intercept and a $T \times 1$ vector of time-specific common parameters $\bm \tau$, respectively. $\bm \varepsilon_i$ is a vector of idiosyncratic error terms, independent across cross-sectional units. In the example of a Difference-in-Differences framework, $\bm \tau$ is the common trend affecting both treated and untreated individuals while the treatment indicator as well as other covariates are contained in $\bm X_{i}$. For the sake of simplicity, we assume $\bm \beta$ and $\mu_i$ to be known, so that $\bm \beta=\boldsymbol{0}$ and $\mu_i = 0 \; \forall i$ can be imposed without loss of generality. This highly restrictive assumption leaves the leading terms in the analysis below unaffected and is hence innocuous for the expository purpose of this section. Model (ref) reduces to
We additionally assume that the variance of $\varepsilon_{i,t}$ is known and fixed to $\sigma_{i}^{2}=1$. In this setup, $y_{i,t}$ are clearly cross-sectionally dependent because of $\tau_{t}$; however we are interested in testing whether $\varepsilon_{i,t}$ are cross-sectionally uncorrelated. Given that $\varepsilon_{i,t}$ are unobserved, the common effects $\tau_{t}$ need to be estimated. The most natural approach is to estimate $\tau_{t}$ by OLS so that $ \widehat{\tau}_{t} = \overline{y}_{t} = \frac{1}{N}\sum_{i=1}^{N}y_{i,t}$ . Using the residuals
the CD test statistic is given by
Given this definition,
implying that the first term in (ref) cancels out. The expression $\sum_{t=1}^{T}\sum_{i=1}^{N}\widehat{\varepsilon}_{i,t}^2$ in the second term is nothing more than the Sum of Squared Residuals (SSR) of the estimated model, a statistic that is of order $\operatorname{\mathcal{O}}_P(NT)$. Consequently,
even though the error-terms $\varepsilon_{i,t}$ are cross-sectionally independent. Hence, the procedure commonly used by practitioners is prone to finding spurious cross-sectional dependence in the data.
The results shown for a model with additive time effects and homoscedastic errors are not coincidental. In fact, as we show later in this paper, they carry over to a more complex model where cross-section dependence is generated by a multifactor error structure and/or the errors are heteroscedastic.
In this section we provide formal asymptotic results for CD statistic based on the residuals obtained from a 2WFE model or a model with multifactor error structure. For the sake of simplicity, we continue to assume that $\bm \beta$ is known throughout this section. If the vector of slope coefficients $\bm \beta$ were unknown, its estimator would converge (under correct model specification) to the true parameter value at the conventional rate $\sqrt{NT}$. Hence, convergence to the true parameter values is fast enough to ensure that any terms present in CD test statistic involving $\widehat{\bm \beta}- \bm \beta_{0}$ are asymptotically negligible. We shall not attempt to prove this claim formally but rely on Monte Carlo results of Section (ref) to support this conjecture. By contrast, the documented effects on the first two moments of the CD test statistic resulted from the fact that estimates of common, period-specific parameters converge at a rate slower than $\sqrt{NT}$.
As in Section (ref) we assume that the true model is given by equation (ref). Additionally, we make the following assumptions on the model errors.
For technical reasons, we assume that all stochastic variables in this paper have finite eighth moments. This is a sufficient condition, which facilitates proving joint convergence of the test statistics considered in this paper, see also JAE:JAE2496. Assumption (ref) is general enough to cover several models of conditional heteroscedasticity in $\varepsilon_{i,t}$. Alternatively this assumption can be formulated in terms of unconditional variances, where $\sigma_{i}$ are fixed numbers. However, in addition to having severe conceptual shortcomings (as discussed in ECTA:ECTA1599) such an approach would lead to incorrect conclusions concerning the power of a CD test statistic for remaining cross-section correlation. See Section (ref) in the Supplementary Appendix for a detailed discussion on power.
For example, natural examples for $\sigma_{i}$ are either a standard exponential skedastic function
or a location-scale model with
Both satisfy the required restrictions, as long as $\mu_{i}$ has a bounded support. Furthermore, we denote different cross-sectional averages of $\sigma_{i}$ by $\overline{\sigma^{k}}=N^{-1}\sum_{i=1}^{N}\sigma_{i}^{k}$ for $k\in \mathbb{Z}$, and the corresponding population quantities by $\operatorname{E}[\sigma_{i}^{k}]$. Assumption (ref) guarantees that these quantities are well defined for all finite $k$.
Using these definitions the CD test statistic obtained from a model with unknown period-specific effects can be characterized as follows:
Theorem (ref) shows that the inclusion of time fixed effects into the model specification has an asymptotically non-negligible effect on the first two moments of the CD test statistic. Via expansion with appropriate functions of $\{\sigma_i\}_{i=1}^{N}$, it can be shown that $\Xi$ indicates the presence of a deterministic bias of order $\sqrt{T}$. Abstracting from this bias, (ref) is dominated by an expression that reflects the CD test statistic obtained from the true model errors but imposing an incorrect normalization.
Restrictions on the expansion rates of $N$ and $T$, i.e. $\sqrt{N}T^{-1} \to 0$ and$\sqrt{T}N^{-1} \to 0$, are satisfied if one considers diagonal asymptotic expansion schemes as in fernandez_weidner_2014, where $N T^{-1}\to \kappa^{2}$ with $ \kappa\in (0;\infty)$. This way $N/T + T/N$ remains a bounded constant asymptotically. The intuition for the need to restriction both $\sqrt{T}N^{-1} \to 0$ and $\sqrt{N}T^{-1} \to 0$ comes from the fact that the feasible CD statistic is a non-linear function in both $\widehat{\sigma}_{i}$ and $\widehat{\tau}_{t}$. Hence the rate restrictions derived for general non-linear panel data models, as studied by fernandez_weidner_2014, are naturally applicable.
The behavior of the CD test statistic in a model with time fixed effects can be seen as an example of the incidental parameters problem (IPP) of neyman1948, since the bias of the CD test statistic is due to the estimation of $T$ period-specific intercepts $\tau_{t}$. Interestingly, in the context of estimating linear dynamic panel data models, estimation of the time effects $\tau_t$ does not introduce any asymptotic bias into the FE estimator with strictly or weakly exogenous regressors hahnmoon2006. In non-linear models, estimation of the time effects $\tau_{t}$ affects the asymptotic mean of the estimator for slope parameters associated with explanatory variables fernandez_weidner_2014 with the corresponding bias being proportional to $\sqrt{T N^{-1}}$. In this sense, our result adds new insights into the literature in that it highlights a scenario where the inclusion of time fixed effects into a linear model has non-standard implications for the asymptotic distribution of the statistic of interest.
The results of Theorem (ref) suggest that asymptotically standard normal inference can be recovered by bias-correcting and rescaling (ref). Before considering this remedy in Section (ref), it is important to emphasize the following special case:
Corollary (ref) provides more intuition about the approximate value of the bias term $\sqrt{T}\Xi$, suggesting that it should be reasonably close to $-\sqrt{T/2}$. More importantly, the result reveals that the leading stochastic component in the CD test statistic, the first term on the right-hand side of equation (ref), cancels out when error variances are homogeneous across $i$. Instead, random variation around the bias term $-\sqrt{T/2}$ is of order $\operatorname{\mathcal{O}}_P(N^{-1/2}\wedge T^{-1/2})$, rendering the distribution of a modified version of the test statistic asymptotically degenerate.
The special case of homogeneous error variances hence entails consequences for the CD test statistic that are qualitatively different from those in the more general case where $\sigma_{i}$ may differ across $i$. Again, it would be possible to allow for asymptotically normal inference by adequately rescaling the modified test statistic. However, the resulting statistic would be of little practical use since the main source of variation is not related to error covariances across cross-sections, but is simply driven by variance of the idiosyncratic components.
Following ECTA:ECTA692 we consider a model where cross-section dependence enters the model via a multifactor error structure. This model is described by the $T\times 1$ vector $\bm y_i$ and the $T\times m$ matrix $\bm X_i$ defined as
where $\bm F = \left[\bm f_1, \bm f_2, \ldots, \bm f_T \right]'$ denotes a $T\times r$ matrix of unobserved common factors and where $\bm E_i = \left[ \bm e_{i,1}, \bm e_{i,2}, \ldots, \bm e_{i,T} \right]'$ is a $T \times m$ matrix of idiosyncratic variation in $\bm X_{i}$. The most popular estimator designed to estimate the parameter vector $\bm \beta$ in this specific model is the Common Correlated Effects (CCE) estimator of ECTA:ECTA692 which amounts to augmenting a linear regression model with cross-section averages of potentially all variables available to the researcher in order to account for the effect of unobservable common factors. In a model with homogeneous slope coefficients, we have
Here $\overline{\bm C}$ ($\overline{\bm U}$) are defined implicitly in terms of $\bm \beta$ and the corresponding average factor loadings (average errors). While the CCE estimator is agnostic about the true number of factors that affect the data, it has been shown that consistent estimation of the parameters of interest requires that the number of cross-section averages is at least as large as the true number of factors that drive the data Westerlund2013247.
Given that the factor estimator $\widehat{\bm F}$, defined in Eq. (ref), is constructed based on observed data, restrictions on the DGP of both $y_{i,t}$ and $\bm x_{i,t}$ need to be imposed.
This set of assumptions above are a slightly more restrictive version of the framework considered in ECTA:ECTA692. For example, the assumption of common $\bm \varSigma$ can be straightforwardly relaxed. However, unlike $\sigma_{i}$, $\bm \varSigma$ plays no major role for asymptotic results of this paper. Most importantly, we rule out the presence of serial correlation as this is in line with the assumptions made for the CD test to work. Furthermore, we restrict ourselves to a classical panel data regression model instead of considering heterogeneous slope coefficients. Moreover, any dependence between $\varepsilon_{i,t}$ and $\bm e_{i,t}$ is assumed away in order to allow for a tractable proof of the main theoretical result in this paper. Lastly, the fact that we assume $\operatorname{rk}\left( \left[\bm \mu_{\bm \lambda} , \bm \mu_{\bm \varLambda}\right] \right) = r = m+1$ to hold, suggest that our we consider an ideal setup where none of the rank condition related problems documented in KaraReeseWesterlund2015, or JuodisEtal2017 apply.
In analogy with the previous sections, we assume that $\bm \beta_{0}$ is known. Thus we disregard the estimation error $\widehat{\bm \beta}^{CCE} - \bm \beta_{0}$, which is generally of order $\operatorname{\mathcal{O}}_{P}((NT)^{-1/2}) + \operatorname{\mathcal{O}}_P(N^{-1})$, see e.g. ECTA:ECTA692 and JuodisEtal2017 .
We begin our asymptotic analysis, by noting that in the model with assumed (known) homogeneous $\sigma$ the result follows directly as in the model with time effects only. In particular, while it is not generally emphasized in the CCE literature, the residuals from CCE estimation which are formally given by
satisfy
In that respect, the standard 2WFE estimator is similar to the CCE estimator. More formally we formulate the following result
Proposition (ref) shows that the result we derived previously for the 2WFE estimator in Corollary (ref) continues to hold for models with a factor error structure, as long as $N\to\infty$ and $T\to\infty$.
It is worth mentioning that the above order effect is only valid if $\widehat{\bm F}$ contains cross-sectional averages of the regressand as well as all regressors. If either of those variables is omitted (without affecting the rank condition in Assumption (ref)), the zero mean residual condition in (ref) is violated, and consequently the $CD$ test will have a $\operatorname{\mathcal{O}}_{P}(1)$ term. However, this result is of limited empirical importance as in most cases researchers include all available cross-sectional averages. While this practice ensures that the estimator is invariant to $\bm \beta_{0}$, inclusion of too many cross-sectional averages can potentially have detrimental effects on the asymptotic properties of the estimator, see the corresponding discussion in JuodisEtal2017 and JuodisRCCE2020.
In the homoscedastic case two-way fixed effects and multifactor error models have similar asymptotic effects on CD test statistic. The conclusions of the next theorem, which is the main result of this paper, indicate that this equivalence does not hold in the heteroscedastic case. In order to proceed, we introduce some useful notation. For $t=1,\ldots,T$, let $\bm u_{i,t}=[\varepsilon_{i,t}+\bm \beta'\bm e_{i,t},\bm e_{i,t}']'$, and
where $\bm \psi_{i,t}=\left(\overline{\bm C}'\right)^{-1}\bm u_{i,t}$ is the influence function of the corresponding factor estimator, in this case cross-section averages of $y_{i,t}$ and $\bm x_{i,t}$. Generally, the influence function depends on the joint process $\left[y_{i,t},\bm x_{i,t}\right]'$, as long as all observed variables are used to form cross-sectional averages. Equipped with this notation we formulate the main result of this paper.
Here, as previously, $R_{N,T}=(N^{-1/2}\vee T^{-1/2}\vee N^{-1}\sqrt{T} \vee T^{-1}\sqrt{N})$. In line with all previous results, the CD test statistic in this setup has two diverging components. However, unlike all previous results, in particular Theorem (ref), these bias terms are not solely non-linear functions of $\sigma_{i}$. Instead, they are also influenced by the first (rescaled) moments of factor loadings in $y_{i,t}$ and $\bm x_{i,t}$, as well as corresponding variances of the idiosyncratic components in $\bm x_{i,t}$. Thus the influence function $\bm \psi_{i,t}$ directly alters distributional properties of the CD statistic. The result is qualitatively similar to any parametric two-step estimation procedure with a plug-in first-step estimator. Thus by including cross-sectional averages as factor proxies, one is implicitly testing that both $\varepsilon_{i,t}$ and $\bm \psi_{i,t}$ are jointly cross-sectionally uncorrelated. Notice, that while $\varepsilon_{i,t}$ is usually assumed to be uncorrelated over $t$ (e.g. for models with pre-determined regressors), the same is not true for $\bm u_{i,t}$, e.g. if $x_{i,t}=y_{i,t-1}$. Thus if Assumption (ref) is appropriately relaxed, then $\xi_{iN,t}$ will be serially correlated.
As alluded in Proposition (ref), the leading $\operatorname{\mathcal{O}}_{P}(1)$ term is degenerate if $\sigma_{i}=\sigma$, as in this case:
One can easily see that the expressions for $\Phi_{1}$ and $\Phi_{2}$ derived for CCE coincide with the corresponding terms of Theorem (ref) (up to a negligible remainder term), upon setting $\bm \lambda_{i}=1$ (thus $\overline{\bm C}^{-1}=1$), and $\bm \psi_{i,t}=\varepsilon_{i,t}$. This situation is equivalent to using $\widehat{f}_{t}=\overline{y}_{t}-\overline{\bm x}_{t}'\bm \beta_{0}$ as factor proxies, e.g. as suggested by JAE:JAE2535 in the context of predictability testing with cross-section dependence.
Recall that result in Theorem (ref) considers only the CCE setup where the number of observable factor proxies equals the number of the true factors. If Assumption (ref) is relaxed and there are more observables than factors, then following KaraReeseWesterlund2015 and JuodisEtal2017, one can show that the expressions for $\Phi_1$ and $\Phi_2$ will contain additional terms related to this discrepancy. In particular, $\Phi_{1}$, $\Phi_{2}$, and $\xi_{iN,t}$ are functions of an unknown rotation matrix, which cannot be consistently estimated from the data. However, as already emphasized by JuodisEtal2017, this problem can be circumvented via the use of a non-parametric bias-correction method. We will come back to this issue in Section (ref).
The diverging bias in the CD test statistic applied to residual from models with common, period-specific parameters is fundamentally problematic for its use in the context of testing for remaining cross-section correlation. Still, the literature review in Section (ref) suggests that the underlying question of correct model specification is of high relevance for empirical researchers. For this reason it is relevant not to discard completely the CD test statistic but instead to discuss possible modifications aimed at ensuring an asymptotically standard normal inference under null hypothesis. Thus, methods aimed at addressing the IPP detailed in Sections (ref) and (ref) need to be considered.
Given that parametric expressions for the bias of $CD$ have been derived in Theorems (ref) and (ref), analytic bias correction is feasible. As detailed in Section (ref) in the Supplementary Appendix, estimates of the unknown model components that constitute either $\Xi$ or $\Phi_1$ and $\Phi_2$ can be obtained in order to eliminate the diverging component in $CD$. In addition, Propositions (ref) and (ref) in the Supplementary Appendix suggest that these estimated bias terms would eliminate equivalent terms $\Xi_{\mathbb{H}_1}$ or $\Phi_{\mathbb{H}_1,1}$ and $\Phi_{\mathbb{H}_1,2}$ under the alternative hypothesis without necessarily reducing the rate at which $CD$ diverges.
Still, problems with its implementation in practice lead us to discard analytic bias correction and to consider it merely as a benchmark approach when evaluating our weighted CD test statistic, as introduced below, in Monte Carlo experiments. In particular, the construction of plug-in estimates of $\Phi_1$ and $\Phi_2$ proves to be tedious since both terms are constructed from estimates of $\sigma_{1}^2, \ldots, \sigma_{N}^{2}$, $\bm \lambda_1, \ldots, \bm \lambda_N$ and $\bm \beta$. Additionally, estimation of the cross-section sum $N^{-1} \sum_{i=1}^N \sigma_{i}^{-1} \bm \lambda_{i}$ poses a trade-off between generality and accuracy of the plug-in estimate. To be specific, independence between factor loadings $\bm \lambda_i$ and error variances $\sigma_{i}^2$ needs to be imposed to ensure that estimation error in $\widehat{\bm \lambda}_i$ and $\widehat{\sigma}_{i}^2$ does not dominate the small-sample distribution of the bias-corrected version of $CD$.
The need to impose further assumptions in order to improve the accuracy of the bias estimates is effectively a consequence of the fact that the unknown bias itself diverges at rate $\sqrt{T}$. This means that any error in the estimation of the two aforementioned components is scaled by a factor that increases as $N,T \to \infty$. Still, despite improving the approximation of plug-in estimates for $\Phi_1$ and $\Phi_2$ the assumption of independence between factor loadings and error variances can lead to size distortions in DGPs where it is violated, see Section (ref) for a numerical illustration of the problem.
The method for bias correction favored in this article is to construct a CD test statistic from estimated cross-section covariances that are weighted with individual-specific, random draws $X$ from a Rademacher distribution, i.e. $P(X=1)=P(X=-1)=0.5$. This approach is based on noticing that the cross-section correlation estimator $\widehat{\rho}_{ij}$, as defined in (ref) with $z_{i,t} = \widehat{\varepsilon}_{i,t}$, is merely a unit-specific weighting of $\widehat{\operatorname{cov}} \left[\widehat{\varepsilon}_{i,t}\widehat{\varepsilon}_{j,t}\right]$ with advantageous properties: Under the assumptions made in CDtest2004, it ensures that the CD test statistic has unit variance under the null hypothesis of no cross-section correlation, allowing for standard normal inference without the need to obtain a variance estimate. However, the case is different in models with unknown, period-specific parameters $\tau_{t}$ or $\bm f_{t}$. As shown in Theorems (ref) and (ref), studentization of the model residuals fails to ensure a test statistic with a variance of one, even when the true error variances are known. This justifies the use of an alternative weighting scheme which we construct as a weighted average of individual specific covariances,
for some set of weights $w_1, \ldots, w_N$. Notice that Eq. (ref) would coincide with the CD test statistic of CDtest2004 for $w_i = \widehat{\sigma}_{i}^{-1} \; \forall i$.
This is not the case we consider here. Instead we make the following additional assumption:
Rademacher distributed weights amount to random sample splitting, a method for breaking dependence that is not new to econometrics. For example, altonji1996 already considered it an old concept. Its effects are most obvious in a simple modification of Theorem (ref) where we replace $\sigma_{i}^{-1}$ with $w_i$. Under Assumption (ref), the expected value of $N^{-1}\sum_{i=1}^N \bm \lambda_{i} w_{i}$ can be conveniently split into the expectations of its two constituents. Given the zero expected value of $w_i$, an appropriate LLN applies and leads the cross-section average of $ \bm \lambda_{i} w_{i}$ to converge to zero. In an even simpler form the same reasoning applies to the cross-section averages involving $w_i$ in an adapted version of Theorem (ref). In both cases, this reduces the order of the leading bias components $\Xi$, $\Phi_{1}$ and $\Phi_{2}$ by a factor of $N$ and entails asymptotic unbiasedness of the weighted CD test statistic, subject to the restriction $\sqrt{T} N^{-1} \to 0$. Additional rescaling with the asymptotic variance of expression (ref) results in a weighted CD test statistic which, as formally stated in Theorem (ref) below, converges to a standard normal distribution.
As stated by Theorem (ref), the use of independent Rademacher distributed weights, analogous to weights drawn from many other distributions with zero mean, re-establishes asymptotic standard normal inference under the null hypothesis of the CD test statistic. However, asymptotic unbiasedness of $CD_W$ comes at the cost of power. More specifically, our approach to bias correction centers the leading components of $CD_W$ around zero, irrespective of whether cross-section correlation in the data is completely controlled for or not. As a consequence, only increases in the variance of $CD_W$ under its alternative hypothesis lead to power against the null hypothesis of this test. We further discuss this point in Section (ref) of the Supplementary Appendix.
To improve the power properties of $CD_W$ we suggest a refinement of this test which follows the power enhancement approach of fan2015. The authors suggest improving the power of high-dimensional cross-sectional tests by adding to the test statistic of interest a screening statistic. This screening statistic is equal to zero with probability approaching one under the null hypothesis of the test, but diverges at a fast rate under the alternative. In our case we choose the absolute sum of thresholded cross-section correlation coefficients. This results in a power-enhanced weighted CD test statistic which is defined as
where $\widehat{\rho}_{ij}$ is as in (ref) with $z_{i,t}=\widehat{\varepsilon}_{i,t}$ and where $\mathbf{1}(A)$ is the indicator function for event $A$. The screening statistic on the right-hand side of (ref) has an asymptotically negligible effect on the size of $CD_{W+}$ because individual cross-section correlation coefficients are still consistent for their true value under $\mathbb{H}_0$. For additional discussion on $CD_{W+}$ we refer to Section (ref) of the Supplementary Appendix.
We investigate the properties of the CD test as well as its weighted alternatives in a small set of simulation experiments. We consider the common factor model
We consider two alternative specifications for the factor loadings on the regressand $y_{i,t}$. We define $\bm \lambda_{i} = \bm \iota_r + \tilde{\bm \lambda}_{i}$ where the $r$ elements of $\tilde{\bm \lambda}_{i}$ are drawn (I) from $U(-0.75, 0.75)$ or (II) from a standardized $\chi ^{2}\left(2\right) $ distribution that has zero mean and a variance of 1/6. The latter case is designed to match the first two moments of $\tilde{\bm \lambda}_{i}$ in case (I). Concerning the factor loadings on the regressors $x_{i,t}$ we draw element $\Lambda^{(r')}$ of the $r\times 1$ vector $\bm \varLambda_{i}$ from $U(-0.5,0.5)$ if $r'=1$ and from $U(0.5,1,5)$ otherwise. Concerning the latent factors, we let $\bm f_{t}\sim N\left( \boldsymbol{0}_{r}, \bm I_r\right) $. When considering the two-way fixed effects model instead of the common latent factor model, we set $r=2$ and define $\bm f_t = \left[f_t^{(1)}, 1 \right]'$ as well as $\bm \lambda_{i}= \left[1, \lambda_{i}^{(2)} \right]'$ and $\bm \varLambda_{i}= \left[1, \Lambda_{i}^{(2)} \right]'$ with the construction of $f_t^{(1)}, \lambda_{i}^{(2)}$ and $\Lambda_{i}^{(2)}$ being unchanged.
Analogous to our treatment of factor loadings we consider two cases for $\varepsilon_{i,t}$, namely that they are drawn independently over $i$ and $t$ from (i) a standard normal distribution or (ii) a standardized $\chi ^{2}\left(2\right) $ distribution that has zero mean and unit variance. The scalar random variable $e_{i,t}$ is generated in the same way as $\varepsilon_{i,t}$. Lastly, error variances are obtained in two different fashions: We set (a) $\sigma _{i}^{2}=c_{\sigma }\left( \varsigma _{i}^{2}-2\right) / \sqrt{24} + 1$ where $\varsigma _{i}^{2}\sim\chi ^{2}\left( 2\right)$ or (b) $\sigma _{i}^{2} = 0.5 + d_{\sigma} T^{-1} \sum_{t=1}^T (\bm \lambda_{i}' \bm f_t)^2$. The scaling $c_{\sigma }$, which is only included in setup (a), is chosen manually as discussed below. This construction entails that $\operatorname{E}\left[ \sigma_{i}^{2}\right] =1$ and $\operatorname{var}\left[ \sigma _{i}^{2}\right] =c_{\sigma }^{2}/6$. The normalizing constant $d_{\sigma}$, which appears in case (b) only, is $d_{\sigma} = \sqrt{2}$ in model error case (i) and $d_{\sigma} = (1/3)^{-1/2}$ in case (b). This scaling ensures that the variance of $\sigma _{i}^{2}$ across cross-sections is approximately of the same magnitude as in case (a) with $c_{\sigma}=1$.
In all following simulations we use feasible CD statistics, where, unlike the theory part of this paper, we treat $\beta$ as an unknown parameter.
In a first instance, we consider the properties of the original CD test statistic when applied to either 2WFE or CCE residuals in a scenario where all sources of cross-section correlation are adequately accounted for. We do this for the setup (i)-(a)-(I) as described above, meaning that errors are normally distributed, error variances independent of factor loadings (or individual fixed effects) and that the latter are symmetrically distributed around their mean. Results for a model with errors from a standardized $\chi^2$ distribution are almost identical and reported in Section (ref) of the Supplementary Appendix. Furthermore, we consider different degrees of heterogeneity among the individual-specific error variances by reporting results for the values $c_{\sigma }\in \left\{ 0.1,0.5,1,1.5\right\}$.
Table (ref) reports the first two moments of $CD$ when applied either to 2WFE residuals when the true model is one with two-way fixed effects or to CCE residuals in a model with multifactor error structure and two factors. The results for both cases are identical and show that $CD$ has a bias that diverges towards $-\infty $ as $T\rightarrow \infty $ and a variance that is considerably below $1$. The bias term is reasonably close to the benchmark value of $-\sqrt{T/2}$ indicated by Proposition (ref) but tends towards zero as heterogeneity among the individual-specific error variances is amplified.
Given that the theoretical results of Sections (ref) and (ref) are confirmed by Table (ref), we turn to the properties of $CD$ under its alternative hypothesis. For this purpose, we simulate a model with multifactor error structure and three factors, implying that neither 2WFE nor CCE estimation completely accounts for all sources of cross-section correlation in the simulated data. Heterogeneity among unit-specific error variances is kept constant by only considering $c_{\sigma}=1$. Instead we consider cases (I) and (II) for factor loadings as well as cases (a) and (b) for error variances. Again, results are presented only for normally distributed model errors. The moments of $CD$ are very similar when model errors are drawn from a standardized $\chi^2(2)$ distribution and corresponding tables can be found in Section (ref) of the Supplementary Appendix.
Tables (ref) and (ref) report the mean and variance of $CD$ in the presence of remaining cross-section correlation. Both are very similar to the numbers reported in Table (ref) as long as factor loadings are symmetrically distributed, an observation which is in line with the theoretical results provided in Section (ref) of the Supplementary Appendix. Since symmetric loadings center the leading stochastic component of $CD$ around zero, the latter is dominated by a bias equivalent that diverges towards $-\infty$ as $T\to \infty$. Qualitatively very different results can be observed when factor loadings are drawn from a skewed distribution, particularly when applying the CD test statistic to 2WFE residuals. As one can observe in columns two and four of Table (ref), the strong negative divergence of $CD$ as $T$ increases is mitigated and it is reasonable to assume that it will be turned into positive divergence for $N>200$. This pattern, which is amplified if error variances are a function of factor loadings, reflects the presence of a nonzero mean in the leading stochastic component of $CD$ that diverges to $+\infty$ at rate $N\sqrt{T}$. In the case of 2WFE residuals, its magnitude is sufficiently large to counteract the negative bias equivalent already for $N=200$. By contrast, columns two and four in Table (ref) suggest that this term is considerably smaller when the CD test is applied to CCE residuals since the mean of $CD$ continues to diverge towards $-\infty$ for $N=200$. We conjecture that the same sign reversal will eventually be obtained even if this case, but that the required cross-section dimension for this to happen is higher.
Next, we proceed with investigating the properties of our weighted CD test statistic. As previously, we keep heterogeneity across unit-specific error variances fixed at $c_{\sigma}=1$ and consider two different specifications each for factor loadings and error variances. We report size and power for our weighted test statistic $CD_W$ as well as the power-enhanced refinement $CD_{W+}$. As a benchmark test, we include a CD test statistic with analytic bias correction which corrects $CD$ with sample equivalents of the asymptotic bias terms in Theorems (ref) and (ref). Details on its implementation can be found in Section (ref) of the Supplementary Appendix. The original CD test is left out for the sake of saving space and in particular since rejection rates are $100\%$ for most cases.
Table (ref) reports the size and power of all three bias corrected test statistic when applied to 2WFE residuals. The size of $CD_W$ and $CD_{W+}$ is very close to the nominal level of $5\%$ as long as $N\geq T$. Size distortions are given in cases where $T$ is considerably large than $N$ and the effect of power enhancement on size is generally negligible. The analytically bias-corrected CD test statistic $CD_{BC}$ exhibits hardly any tendency to over-reject but is rather conservative, in particular when $T$ is small.
Panel B in Table (ref) report results on power. Here, we see that the power of $CD_W$ is in general low and increases only in $T$. This is improved upon considerably by power enhancement, leading the refined test statistic $CD_{W+}$ to reliably reject when it should as long as the number of time periods is large enough. The performance of $CD_{W+}$ is considerably above that of the benchmark statistic $CD_{BC}$ when factor loadings are drawn from a symmetric distribution and is on par in most other cases. An exception is the case of skewed loadings with small $T$ and large $N$ where $CD_{BC}$ has markedly higher rejection rates. It can also be noted that the performance of $CD_{BC}$ largely depends on whether factor loadings are drawn from a symmetric distribution or not. If this is the case, analytic bias correction leads to rejection rates under the alternative hypothesis that are worst among all three tests considered.
When testing for remaining cross-section correlation in CCE residuals we observe similar results as in Table (ref). Interestingly, it can be observed that the benchmark statistic $CD_{BC}$ exhibits size distortions in samples with large $T$ when loadings are drawn from a skewed distribution and if factor loadings and error variances are dependent. This results from the nature of our bias correction which assumes independence between $\sigma^{2}_i$ and $\bm \lambda{i}$ to considerably improve the accuracy of the bias estimate.
The power properties of $CD_W$ and $CD_{W+}$ when either of these tests is applied to CCE residuals mirror those seen in the 2WFE case, even though rejection rates are generally somewhat lower. However, the power-enhanced statistic $CD_{W+}$ now performs best in all cases without ever being inferior to $CD_{BC}$.
In this section we illustrate the applicability of the standard and the weighted CD statistics using the R&D investments data of doi:10.1162/REST00272. Information on up to twelve manufacturing industries in ten countries (Denmark, Finland, Germany, Italy, Japan, The Netherlands, Portugal, Sweden, United Kingdom, and The US) over a time period from 1980 to 2005 was used to construct the dataset. After some minor modification to the original dataset (see Section (ref) of the Supplementary Appendix) we are left with with a panel dataset of $N=82$ and $T=25$.
In this application serial correlation is important from an economic point of view. In Section (ref) of the Supplementary Appendix we outline how the CD statistic can be modified to account for serial correlation, given a set of known weights $w_{i}$.
doi:10.1162/REST00272 question whether R&D can be estimated in a standard Griliches-type “knowledge production function” framework ignoring the potential presence of knowledge spillovers between cross-sectional units as well as other cross-section dependencies. Among other things they document a “[\ldots] strong evidence for cross-sectional dependence and the presence of a common factor structure in the data, which we interpret as indicative for the presence of knowledge spillovers and additional unobserved cross-sectional dependencies” doi:10.1162/REST00272. Cross-sectional dependence was measured by means of a CD test.
In this section we will primarily revisit some of the results in Table 5 of doi:10.1162/REST00272 for pooled (static) production function estimates. Our goal is to investigate how the divergent properties of the CD test statistic might have influenced the choice between the First Difference (FD) estimator with yearly dummies and Pooled CCE (CCEP) estimators, as presented in columns 3 and 4, respectively, of the table mentioned above. Which of the two models is considered to be correctly specified has important consequences for the conclusions that can be drawn from the entire table. With regard to the coefficient of private R&D investments, doi:10.1162/REST00272 report that a significance test for the corresponding slope coefficient cannot reject the null hypothesis in the FD model while it can in the CCE model.
As the results for $CD_{W}$ statistic allowing for serial-correlation correction, are similar to those without serial correlation, we only focus on the latter option. First of all, as we can see from Table (ref) the conclusions that we can draw from the original CD test are almost identical to those doi:10.1162/REST00272, despite adjustments made in terms of the sample size. In particular, while the value of CD statistic based on CCE residuals imply rejection of the null hypothesis, no such conclusion is implied by FD (at least at a $5\%$ significance level). However, as given that the original test statistic suffers from the IPP, we further investigate if this conclusion also holds after proper bias-corrections. Motivated by finite sample evidence in Section (ref), we consider only the $CD_{W}$ statistic.
Taking our attention to the proposed $CD_{W}$ statistic based on random Rademacher weights, we can see that the corresponding values of $\overline{CD}_{W}$ indicate that both the FD and CCEP models generate residuals that are cross-sectionally uncorrelated. As we discussed in Section (ref), this conclusion might be partially attributed to the fact that for small values of $T$, the proposed statistic might lack power. For this reason, we also make use of the power-enhanced version $\overline{CD}_{W+}$ of our test statistic. Notice that for this dataset $2\sqrt{\ln(N)/T}\approx 0.84$, thus power-enhancement will pick up only very large values of the correlation coefficients that otherwise might be averaged away by the original CD statistic. As we can see from the corresponding column in Table (ref), power enhancement does not alter the conclusion we draw for the FD residuals, as there are only 2 coefficients $\widehat{\rho}_{ij}$ above the threshold. However, for CCEP residuals the effect of power enhancement is non-negligible as there are 10 values of $\widehat{\rho}_{ij}$ above the threshold.
These results indicate that after appropriate adjustments, the empirical evidence presented in doi:10.1162/REST00272 that favor simple First-difference estimator, as opposed to the CCEP estimator, remain in place. As we can see from Table (ref), this conclusion is unaltered after appropriate serial-correlation adjustments, thus serial-correlation cannot be the main factor affecting this ranking (unlike the discussion in doi:10.1162/REST00272). Instead, our results might indicate that the underlying CCE restrictions (rank condition and regressors with finite factor structure) might be violated, as these restrictions are irrelevant for the FD estimator with time-effects only.
This article documents how the estimation of common time-specific parameters using panel data causes the CD test of CDtest2004, doi:10.1080/07474938.2014.956623 to break down. Using popular additively and multiplicatively interacted specifications for individual- and time-specific components in the model errors, we show that the CD test statistic applied to residuals of correctly specified regression models is divergent under null hypothesis of cross-section independence. We can find an equivalent term under the alternative hypothesis which may balance out the leading diverging component of $CD$ and can lead to low power in small samples. The results documented in this article are interpreted as a manifestation of the incidental parameter problem (IPP) since they ultimately follow from the estimation of $\operatorname{\mathcal{O}}(T)$ period-specific parameters. Our main theorems illustrate the pervasive nature of the IPP in this setup, given that the consequences of estimating time specific parameters do not disappear as the sample size increases.
Our proposed weighted CD test statistic achieves our primary goal of re-establishing asymptotic standard normal inference and hence constitutes an alternative to popular bias correction methods which circumvents the problems these approaches have in the present context. Our results have far reaching implications for empirical panel data analysis, where CD test has been widely used as a model selection/diagnostic tool. An illustration of how our theoretical results translate into applications is given via simulations and real datasets.
Finally, in this paper we assumed that the parameters in the linear model are estimated using a least squares objective function. If one deviates from this setup, and instead uses an (over-identified) GMM criterion function to estimate parameters, the usual GMM J-statistic is readily available for the purpose of testing residual cross-sectional correlation. Examples in a fixed-$T$ framework are given by Sarafidis2009, AhnEtal2013, and JuodisPseudo2014 among others. In the large $N,T$ setup average J-statistic as a model specification tool was explicitly used e.g. by JAE:JAE2338 for the CCE-GMM estimator. Hence given the scale of the problems with CD statistic documented in this paper, these alternative procedures (if applicable) are more appropriate.
\makeatletter\@input{auxsupp.tex}\makeatother