Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,820 characters · 4 sections · 61 citation commands
On Selection of Cross-Section Averages in Non-stationary Environments
{0.5cm} { \affil[1]{ Free University of Bozen-Bolzano} \affil[2]{BI Norwegian Business School, Department of Economics} }
Keywords: Information criteria, factors, CCE, cross-section averages, non-stationary data\\ MOS subject classification: 91B84
Consider a set of $K$ variables (stacked over time $t=1,\ldots,T$) that admits a factor structure for cross-sectional units $i=1,\ldots,N$:
where $\*Z_{i}=[\*z_{i,1},\ldots,\*z_{i,T}]'\in \mathbb{R}^{T\times K}$, $\*U_i$ is an error term, while $\*F\in\mathbb{R}^{T\times m}$ are latent common factors, where $m$ denotes the number of factors, and $\*C_i\in \mathbb{R}^{m\times K }$ is the loading matrix. ((ref)) nests several settings. The most popular one is that of interactive effects (see bai2009panel). Then $\*Z_i=[\*y_i, \*X_i]\in \mathbb{R}^{T\times (k+1)}$, where $\*y_i=\*X_i\+\beta+\*F\+\gamma_i+\+\varepsilon_i$ gives a model with unobserved heterogeneity. If $\*X_i=\*F\+\Gamma_i+\*V_i\in \mathbb{R}^{T\times k}$, we have the Common Correlated Effects (CCE) setting of Pesaran2006 with $\*V_i, \+\varepsilon_i$ ($\+\Gamma_i, \+\gamma_i$) representing idiosyncratics (loadings). ((ref)) goes beyond interactive effects and CCE. stauskas2025handling consider Distinct Correlated Effects (DCE), where $\*y_i$ and $\*X_i$ load on distinct sets of factors. massacci2024forecasting use $\*Z_i=\*X_i$ to model a set of co-moving regressors in a forecasting equation. \\
It is possible to estimate unobserved factors in several ways, including the Principal Components (PC) method (see bai2006confidence) or diversified projections (see fan2022learning). Pesaran2006 suggests a simple and elegant way to proxy $\*F$ up to a linear transformation by taking cross-section averages (CAs): $\overline{\*Z}=\frac{1}{N}\sum_{i=1}^N\*Z_i=\widehat{\*F}=\*F\overline{\*C}+o_p(1)$, because CAs of the idiosyncratics are negligible under a wide range of empirically relevant assumptions (see e.g. Pesaran2011a). This is enough to estimate $\*F$ in DCE and forecasting settings. In interactive effects, the CCE estimator of $\+\beta$ is then simply the least squares (LS) estimator augmented with $\overline{\*Z}=\widehat{\*F}$ as an additional regressor. \\
A substantial advantage of CAs is their ability to estimate $\*F$ irrespective of their time series properties. Indeed, the CCE estimator enjoys popularity, because it is consistent and asymptotically normal under very general $\*F$ in large $N,T$ settings without any modifications (see e.g. Kapetanios2011, westerlund2018cce, westerlund2018asymptotic, or stauskas2022complete). This property is valid if the rank of $\overline{\*C}$ is equal to $m$, which means that the number of factors cannot exceed the effective number of CAs. However, there are reasons to ensure that the number of CAs matches $m$. In CCE, if $m<k+1$ and $TN^{-1}$ is bounded, CCE produces an asymptotic bias whose analytical correction is infeasible (see Karabiyik2017). This situation is common in macro datasets (see de2024cross). In general, too many CAs might even lead to a model that over-fits the data. In contrast, inclusion of too few generally leads to inconsistent estimates of the model parameters (see Juodis2022CCER, for a discussion). Beyond CCE, stauskas2025handling show that DCE only uses $\overline{\*Z}=\overline{\*X}$, but $m=k$ must hold. In forecasting, the same is needed to avoid only conservative confidence intervals around factor-augmented forecasts (see karabiyik2021forecasting). In response to this, margaritella2023using (MW) were the first to demonstrate that the information criteria (ICs) of Bai2002 and bai2009_supp from the PC literature can be applied in the pure CCE setting under stationary factors to identify an optimal set of CAs from $\overline{\*Z}$. de2024cross (DVS) show similar results when CAs are selected from $\overline{\*X}$, which is more applicable in the DCE and forecasting settings, but the procedure has a similar theoretical basis.\footnote{Even in basic CCE setting, it is advised to omit $\overline{\*y}$ and use $\overline{\*X}$ only to avoid computational issues (see karavias2023structural).}$^,$\footnote{PC techniques are limited in CAs setting, as they focus on $m$ by detecting the largest eigenvalues of the data matrix. They cannot detect sets of CAs as they do not have a natural ordering. With IC, we learn $m$ from the cardinality of the selected set. } \\
MW stress that an assumption of stationary $\*F$ is required only to simplify the proofs, hinting at a much greater generality of IC. As CAs proxy a general factor structure, it is natural to evaluate this statement and re-visit the ICs proposed in both MW and DVS due to their wide applicability and similar theoretical foundation. The subtlety arises here because IC is minimized by grid-search at different combinations of CAs. Inevitably, there is a need to understand the asymptotic behavior of IC evaluated at a combination inconsistent for $\*F$ that is non-stationary, which is a new undertaking in the CAs literature. Therefore, this study can be seen as the CAs counterpart of bai2004estimating, where classical PC results of Bai2002 are evaluated against pure unit root factors. Instead, we use a mildly integrated process by magdalinos2009limit to model $\*F$, which allows us to experiment with varying degrees of factor persistence. We demonstrate formally and in simulations that, while ICs remain consistent as $(N,T)\to \infty$, highly persistent factors negatively affect their small sample performance. We discuss differences between our results and those in bai2004estimating, and also explore penalty functions adapted to non-stationary $\*F$, as suggested in the latter study. Since they make the performance of our ICs even worse, we argue that practitioners should not mechanically take recommendations from the PC literature if the goal is to select CAs. More importantly, the ICs of DVS and MW (with the original penalties) should be applied in relatively large samples if the presence of non-stationary factors is suspected.
We focus on the IC of DVS to analyze our theoretical results, however, the conclusions apply to MW, as well. While we provide numerical experiments on both DVS and MW for comparison, theoretical arguments regarding the latter are similar, and we relegate them to Section 3.3 of the Supplementary material. To operationalize our analysis, we introduce $M$ which is a set of column indices of $\overline{\*X}$, and $\*q_{M}\in \mathbb{R}^{k\times g}$ picks the corresponding $g$ averages in practice. That is, $\overline{\*X}\*q_M=\widehat{\mathbf{F}}_M$ defines a selection of $g$ out of $k$ CAs. Consequently, let $M_{0}$ denote the true set of averages from $\overline{\*X}$ such that $\mathrm{rank}(\overline{\+\Gamma} \*q_{M_0})=m$, and $|M_0|=m$, where $|M|$ denotes the cardinality of an arbitrary set $M$. IC under consideration is given by
where $\ln(.)$ denotes the natural logarithm and $\mathrm{det}(\*A)$ is a determinant of any square matrix $\*A$. $\overline{\*Q}_M=\frac{1}{NT}\sum_{i=1}^N\*X_i'\*M_{\widehat{\*F}_M}\*X_i$, $\*M_\*A=\*I-\*A(\*A'\*A)^+\*A'$ is a projection matrix, $\*A^+$ is the Moore-Penrose inverse, and $p_{N,T}$ is a penalty term.\footnote{In case of MW, $\overline{\*Q}_M= \frac{1}{NT} \sum_{i=1}^{N} \widehat{\+\nu}_i' \mathbf{M}_{\widehat{\*F}_M} \widehat{\+\nu}_i $ (scalar), where $\widehat{\+\nu}_i=\*y_i-\*X_i\widehat{\+\beta}$ and $\widehat{\+\beta}$ is obtained using the CCE estimator under all available $k+1$ CAs. Then $\mathrm{IC}(M)=\ln(\overline{\*Q}_M)+g\cdot p_{N,T}$.} Examples of feasible and most popular penalties are given by
for $C_{N,T}=\mathrm{min}(\sqrt{N},\sqrt{T})$ (see more in Section 5 of Bai2002). Note that ((ref)) is function of (a version of) the denominator of the CCE estimator, which is robust to general unknown factors as long as CAs are (rotationally) consistent for $\*F$ (see Theorem 1 in westerlund2018cce). \\
Let $\overline{M}$ be the set of indices of all available CAs that, so that $|\overline{M}|=k$. Then
where $|\widehat{M}|=g$ provides the estimator of the number of factors as a consequence. In DVS (and MW), it is demonstrated that $\mathbb{P}\left(\mathrm{IC}(M)-\mathrm{IC}(M_0)<0\right)\to0$ (therefore, $\mathbb{P}(\widehat{M}=M_0)\to1$) as $(N,T)\to \infty$. This result means that some other $M$ does not minimize IC asymptotically and theoretically justifies its use in the case of CAs. These findings are based on stationary $\*F$, which has been considerably relaxed in the CCE/CAs literature. Therefore, our natural goal is to examine ((ref)) in a way that departs from this restrictive assumption. Throughout our analysis, we employ the following set of assumptions. \\
Assumption 1. $\{\*f_t \}$ is a mildly integrated process as defined in magdalinos2009limit, such that
where $\*u_{f,t}$ is a zero-mean linear process.
Assumption 2. Let $\*e_{i,t}=(\varepsilon_{i,t}, \*v_{i,t}')'\in \mathbb{R}^{k+1}$. Then
Assumption 3. $\mathbf{f}_t$ and $\mathbf{e}_{i,s}$ are independent for all $t$, $s$ and $i$.
Assumption 4. $\*C_i$ is a deterministic matrix, such that $\|\*C_i\|<\infty$ and $\frac{1}{N}\sum_{i=1}^N\*C_i\*C_i'\to \+\Sigma_\*C$ positive definite. Also, $\overline{\*C}=[\overline{\*C}_m, \overline{\*C}_{-m}]$, where $\overline{\*C}_{-m}\in \mathbb{R}^{m\times (k+1-m)}$ and $\overline{\*C}_m=\overline{\*C}\*q_{M_0}\in\mathbb{R}^{m \times m}$ for a unique $M_0$ is full rank for all $N$, including $N\to \infty$. If $m=k+1$, then $\overline{\*C}=\overline{\*C}_m$.
Assumption 1 (a) treats the factors very flexibly and offers comparative statics, as $\tau \to 1$ increases persistence of the process. We are not interested in a specific model for $\*F$, but magdalinos2009limit allow us to vary $\tau$ and give theoretical guarantees, such as $\frac{1}{T^{1+\tau}}\*F'\*F\to_p\+\Sigma_{\*F}$, which is a constant positive definite matrix. The current specification ensures that we return to the usual stationarity conditions under $\tau=0$, as the eigenvalues of $\*I_m-\*G$ lie within a unit circle, so that the process can be inverted and admit a representation of $MA(\infty)$. Assumption 2 (a) is split into two parts. If the idiosyncratics $\{\*e_{i,t}\}$ are correlated over time, for $\*E_i=[\*e_{i,1},\ldots, \*e_{i,T}]'$, we have that $\left\|\frac{1}{T^{1+\tau}}\*F'\*E_i\right\|=O_p(T^{-\tau/2})$ (see Lemma 3.1 in magdalinos2009limit), which is too slow to demonstrate that the factor estimation error is negligible. In order not to obscure the very effect of factor persistence on IC, we restrict the correlation of idiosyncratics and show that $\left\|\frac{1}{T^{1+\tau}}\*F'\*E_i\right\|=O_p(T^{-(1+\tau)/2})$ as needed (see our auxiliary results in the Supplementary material, and a comparative rate requirement when $\tau=0$ in e.g. Karabiyik2017). We present simulations under serial correlation in the Supplement, which do not show any negative effect on IC. The rest of the assumptions are standard, as they ensure weak cross-section dependence of the idiosyncratics (Assumption 2 (b)), mutual independence of factors and the error terms (Assumption 3; see also Pesaran2006), and informativeness of the loadings (rank condition in Assumption 4). The loadings are deterministic, but they can also admit a random coefficient model (see e.g. de2021bias). In our assumptions, we cover $k+1$ variables in order to accommodate MW, as well. \\
Let $\mathrm{d}\overline{\*Q}_{M,M_0}=\ln\mathrm{det}(\overline{\*Q}_M)-\ln\mathrm{det}(\overline{\*Q}_{M_0})=\ln\det\left(\*I_k+T^{\tau}\+\Omega_{M,M_0}\right)$, where $\+\Omega_{M,M_0}:=T^{-\tau}[\overline{\*Q}_M-\overline{\*Q}_{M_0}]\overline{\*Q}_{M_0}^{-1}$. $\mathrm{d}\overline{\*Q}_{M,M_0}$ is important since it helps describing the behavior of $\mathrm{IC}(M)$ in the neighborhood of $\mathrm{IC}(M_0)$ under $M_0\subset M$ (over-specification) and $M\subset M_0$ (under-specification). In addition, it prescribes properties of $p_{N,T}$, which can be seen from
To illustrate, bai2004estimating (in PC setting) shows that under $M_0\subset M$, (an equivalent of) $\mathrm{d}\overline{\*Q}_{M,M_0}$ is $O_p(1)$ when factors are non-stationary, and so the contribution of the extra estimates to the sum of squared residuals is non-negligible. This leads to $p_{N,T}\to \infty$ to heavily penalize redundant estimates so that $\mathrm{IC}(M)-\mathrm{IC}(M_0)>0$. However, when factors are stationary, Bai2002 (and DVS) detect $\mathrm{d}\overline{\*Q}_{M,M_0}=O_p(C_{N,T}^{-2})$, which means that $p_{N,T}\to 0$ at a slower rate to penalize lightly to ensure $\mathrm{IC}(M)-\mathrm{IC}(M_0)>0$ asymptotically. Proposition 1 below can be seen as the CAs equivalent of Lemma A3 and A4 of bai2004estimating.\footnote{ Strictly speaking, under-specification happens when $M\subset M_0$, $M_0 \cap M\neq \emptyset$ but neither is a weak subset of each other, and when $M\cap M_0=\emptyset$. We only analyze the case of $M\subset M_0$ for brevity. As margaritella2023using pointed out, analysis would lead to the same conclusion in all cases.}\\
Proposition 1. Under Assumptions 1-4 as $(N,T)\to \infty$, we have that:
where $\+\Omega_{M.M_0}^0=\operatorname*{plim}_{(N,T)\to\infty}\+\Omega_{M.M_0}$ (positive semi-definite, non-zero), $|R_{N,T}|=o_p(1)$ is the remainder, $\lambda_j(\*A)$ is the $j$-th eigenvalue of $\*A$, while $b$ is a number of strictly positive eigenvalues.}
Proof: Section 3 of the Supplementary material. \\
In (a), the CAs are able to approximate all $m$ factors under Assumption 4. Then $\overline{\*Q}_M$ is exactly (log-determinant of) the denominator of the CCE estimator. It uses the fact that under $M_0\subset M$ we have $\*F=\overline{\*X}(\overline{\+\Gamma}\*q_{M})^++o_p(1)$ for fixed $T$, and by inserting this into $\*X_i$, we can show that $\ln\mathrm{det}(\overline{\*Q}_M)=\ln\mathrm{det}\left(\+\Sigma_\*v\right)+o_p(1)$, whose argument is a positive definite matrix under general factors (see also westerlund2018cce). Given that $\ln\mathrm{det}(\overline{\*Q}_{M_0})$ admits the same asymptotic representation, it is natural that $\mathrm{d}\overline{\*Q}_{M,M_0}$ is negligible. Moreover, the rate is already detected by DVS, MW and Bai2002 (for PC) under stationary $\*F$, and it remains identical here. This rate is instructive and determines the properties of penalty $p_{N,T}$, which is the key motivation behind the functional forms in ((ref)). Specifically, because redundant $g-m$ CAs have an asymptotically negligible contribution to the sum of squared residuals, over-specification does not need to be heavily penalized, and so ((ref)) becomes
as $(N, T)\to \infty$ if $p_{N,T}\to 0$ and $p_{N,T}C_{N,T}^2\to \infty$. Note that this stands in sharp contrast to findings in Lemma A4 of bai2004estimating, where excess factor estimates contribute non-trivially. Overall, the usual consistency of CAs under general unknown factors prevails when $M_0\subset M$ (see also westerlund2018cce, or stauskas2022complete). However, the result is more nuanced under $M\subset M_0$. \\
In (b), the rank condition in Assumption 4 is not satisfied. This means that CAs are inconsistent for the $m-g$ factors. Note that $\overline{\*Q}_M$ is a quadratic form in $\*F$, and we know that $\left\|\*F'\*F\right\|=O_p(T^{1+\tau})$ as implied by Assumption 1. However, the objective function offers normalization by $T$ only as the integration order is unknown. Exactly this imbalance causes the divergence of $\mathrm{d}\overline{\*Q}_{M,M_0}$. A similar situation arises in Lemma A3 of bai2004estimating, where the central step is to detect the sign of divergence, and we follow this route. In ((ref)), $b$ is finite ($b< k$), so $\mathrm{d}\overline{\*Q}_{M,M_0}\to_p + \infty$ and therefore
as $(N,T)\to \infty$, as required. Two comments are in order. Firstly, neither $\sum_{j:\lambda>0}\ln\left(\lambda_j(\+\Omega_{M.M_0}^0)\right)$ nor $R_{N,T}$ are guaranteed to be positive. In addition, in the Supplement, we show that $|R_{N,T}|=O_p(N^{-1})+O_p(N^{-1/2}T^{(\tau-1)/2})$, which means that for $\tau \approx 1$ and a small $N$, it vanishes slowly as $T\to \infty$. This can induce $\mathrm{IC}(M)-\mathrm{IC}(M_0)<0$ in small samples. Hence, it may be necessary to rely on large $N,T$ combinations so that $\tau b\ln(T)$ dominates to detect the minimum of IC outside of $M\subset M_0$ region if $\*F$ is very persistent. We explore and confirm such risks in Monte Carlo simulations in Section 3, where we see that both MW and DVS misselect (in fact, underselect), unless both $N$ and $T$ are large. Secondly, while this result is similar to Lemma A3 in bai2004estimating, there the divergence rate is $O_p(T/\ln\ln(T))$, since the law of iterated logarithm is used to arrive at the expression analogous to ((ref)) when factors have unit root. Theorem 1 leads to our consistency result.\\
Theorem 1. Under conditions of Proposition 1 with $p_{N,T}\to 0$ and $p_{N,T}C_{N,T}^2\to \infty$, we have as $(N,T)\to \infty$
} Proof. Follows from Proposition 1, because $\mathbb{P}\left(\mathrm{IC}(M)-\mathrm{IC}(M_0)<0\right)=\mathbb{P}\left(\left(\mathrm{IC}(M)-\mathrm{IC}(M_0)\right)/p_{N,T}<0\right)\to0$ under (a) and a similar statement holds under (b).\\
The key message of Theorem 1 is that the IC of DVS (and MW) is consistent, but exhibits hybrid asymptotic properties characteristic to stationary and non-stationary factors. Consequently, two forces are in effect: $p_{N,T}$ should still be negligible, but the factor estimation error can be “large” if $M\subset M_0$ as indicated by the remainder in ((ref)). To balance these properties in practice, it may be tempting to experiment further with the penalty. Indeed, the change of penalty is the key message of Theorem 1 in bai2004estimating in the PC setting, where two conditions are satisfied: 1) $p_{N,T}\to \infty$, and 2) $p_{N,T}/\ln(T)\to0$ which is given in footnote 3 of the study to adapt the penalty to the logarithmic sum of squared residuals, which is exactly our setting. Importantly, this condition does not interfere with consistency in our Theorem 1, because clearly $p_{N,T}=o(\ln(T))$ holds. Therefore, we compare both types of penalties in the simulations to illustrate the outcome if practitioners simply extrapolate recommendations from the PC literature. We let $\widetilde{p}_{N,T}=\ln(T)p_{N,T}$, which follows the suggestion in (12) in bai2004estimating, and obeys $\widetilde{p}_{N,T}\to \infty$ and $\widetilde{p}_{N,T}/\ln(T)\to0$. \\
Remark 1. Note that under stationarity, $\mathrm{d}\overline{\*Q}_{M,M_0}\to_pc>0$ when $M\subset M_0$, according to the results in DVS (and similarly in MW). Our Proposition 1 naturally accommodates this result under $\tau=0$, because, by following the same approximation steps, ((ref)) becomes
as $(N,T)\to \infty$ because each summand is greater than 1 w.p. 1. Here, $\+\Omega_{M.M_0}^0=\operatorname*{plim}_{(N,T)\to \infty}[\overline{\*Q}_M-\overline{\*Q}_{M_0}]\overline{\*Q}_{M_0}^{-1}$ (positive semi-definite, non-zero). Also, $|R_{N,T}|=O_p(N^{-1})+O_p((NT)^{-1/2})$, which is the same as in DVS (and MW) under $\tau=0$, as expected from the discussion of ((ref)).}\\
Remark 2. Throughout, we assume that $M_0$ is unique. However, because CAs do not have a natural ordering, the problem is only set-identified. It can be shown that both ICs select the set that minimizes the asymptotic mean squared error out of all the sets that satisfy Assumption 4 (see e.g. Corollary 3.1 in margaritella2023using). In this study, the first order issue is ability of IC to select any $\widehat{M}$ with the cardinality of $m$. \\
Remark 3. Juodis2022CCER considers eigenvalue-based method of ahn2013eigenvalue under stationarity, where the focus is learning $m$, but not the set of CAs. We leave analysis of this approach for the future research, but provide comparative simulation evidence under mildly integrated $\*F$ in the Supplementary material.
The data generating process of the simulation follows MW:
where \(\mathbf{X}_i = \left[\mathbf{x}_{i,1},...,\mathbf{x}_{i,T}\right]'\) is a $T\times k$ matrix of observable regressors, \(\boldsymbol{\beta} = \iota' 0.5\) a $k\times 1$ vector of unobserved parameters. The factor loadings are generated as $\boldsymbol{\Gamma}_i = [\*I_{m}, \*0_{m \times (k-m)}]\psi_i$, that is, a matrix $m \times k$, with $\psi_i\sim N(1,1)$ and the elements of the vector $1\times m $ $\boldsymbol{\gamma}_i$ are drawn from $N(1,1)$. In the case of errors independent over time and cross-sectionally, $\mathbf{e}_{i,t}$ is drawn from $N(\*0_{k+1},\*I_{k+1})$. In the case of errors correlated weakly across time and/or cross-sections we follow Bai2002,margaritella2023using:
where \( \boldsymbol{\epsilon}_t\) is a $N\times 1$ stack of $\epsilon_{i,t} \sim N(0,1)$ over $i$, and $\*w_i$ is the $i$-th row ($1\times N$) of $\*W$, a weight matrix with the $J$-th off diagonal elements equal to zero. $\rho$ and $\rho_v$ control the correlation over time in $\boldsymbol{\varepsilon}_t$ and $\mathbf{V}_i$, respectively. For this simulation exercise, we set them to zero. In the Supplement we explore settings with $\rho=\rho_v=0.5$, an extension to our theory which prescribes uncorrelated innovations. Across all specifications, we allow for weak correlation across units with $\kappa = \kappa_v = 0.2$, $J=J_v=5$ and $\*w_i=\tilde{\*w}_i$. $\mathbf{\Upsilon}_t$ is a $N \times k$ matrix, which stacks $\+\upsilon_{i,t}\sim N(\*0_k,\*I_k)$ over $i$, and it is independent of $\+\epsilon_t$. The vector $\+\iota_{N,i}\in \mathbb{R}^{N\times 1}$ contains 1 in the $i$-th coordinate, and zeros elsewhere. The factors are drawn as $\mathbf{f}_t = \*R_{f,T} \mathbf{f}_{t-1} + \*u_{f,t}$, where
In the stationary factors case, the factors are drawn as a stable VAR with $\tau=0$. Two cases of non-stationary factors are considered: the non-stationary case of $\tau = 0.4$, and the case close to local-to-unity process $\tau = 0.9$. All simulations have 1000 repetitions, and we compare DVS and MW. In the Supplement, we provide additional simulation evidence of the eigenvalue growth ratio (ER) based on Juodis2022CCER for comparison. The focus of this method is the number of factors ($m$), and it does not select the sufficient set of CAs. Nevertheless, its performance crumbles with $\tau\to 1$ and does not improve as $(N,T)\to \infty$. \\
Table (ref) below presents the frequency of the correctly selected number of CAs ($g = 4$) with default penalties $p_{N,T,1}$ and $p_{N,T,2}$ from ((ref)), and the penalty from bai2004estimating, $\widetilde{p}_{N,T}$, for the case of stationary factors ($\tau = 0$), the mildly non-stationary case ($\tau = 0.4$) and the close to local-to-unity process. The idiosyncratic terms $\*v_{i,t}$ and $\varepsilon_{i,t}$ are uncorrelated over time, but weakly correlated over units. In what follows, the low selection frequency reported is dominated by underslection. We start with the case established in the literature, which is stationary factors and default penalties. For small $N$ and $T$ ($T=N=50$) and stationary factors, we observe a frequency between $16\%$ and $36\%$, a pattern which resembles the results in MW.\footnote{In comparison to margaritella2023using we set the number of factors to 4 rather than 2 and hence underestimate the number of CAs more severely in the case of $T=20$ and $N>20$. If we reduce the number of factors to 2, we identify a similar number of factors as in the case of 4 factors. } As expected, consistency, illustrated here by a selection frequency of 100%, is quickly achieved when $N$ and $T$ increase. We now turn to the case of non-stationary factors. For small $N$ and $T$, the selection frequency decreases to zero, which is driven by underselection of the number of CAs, as we show in Section 4 in the Supplement. To gain consistency, much larger combinations of $(N,T)$ are required. For the extreme case of a process close to local-to-unity, combinations of $(N,T)>300$ are integral. The behavior of both criteria with respect to increases in $N$, $T$, and $\tau$ is almost identical. bai2004estimating proposes to adjust the penalty by $\ln(T)$ in the presence of non-stationary factors. The lower panel of Table (ref) presents the frequency of the correctly selected number of CAs with a penalty of $\widetilde{p}_{N,T} = \ln(T)\frac{N+T}{NT}\ln(\frac{NT}{N+T})$ for $\mathrm{IC}_1$ and $\widetilde{p}_{N,T} = \ln(T)\frac{N+T}{NT}\ln(C^2_{N,T})$ for $\mathrm{IC}_2$. We observe a selection frequency of 0 for all four ICs and for almost all combinations of $N$ and $T$ and $\tau$. Only in the large sample cases of $N=500$ and $T>300$ and only under stationary factors the selection frequency reaches 100%. This strongly suggests that prescriptions from the PC literature do not directly apply to CAs, as ICs perform even worse, and selecting more CAs should not be heavily penalized. \\
In the Supplement, we also investigate the share of misselected number of cross-sectional averages when increasing $\tau$ using $IC^{MW}_1$ and $IC^{DVS}_1$ for a small ($N=T=100$) and large sample ($N=T=500$). For the small sample, the share remains relatively flat until $\tau$ reaches levels of around $0.5$. As expected, this is much less of a problem for the large sample and we observe under-selection of CA only for a high degree of non-stationarity.
This study is the first to examine the selection of an optimal set of CAs in CCE and related settings by ICs inspired by Bai2002 when latent factors are non-stationary. In particular, we use mild integration to explore varying degrees of factor persistence and demonstrate that ICs remain consistent without any modifications. However, the more persistent common factors are, the worse their small sample performance becomes, and ICs regain selection consistency only in very large samples (i.e. $(N,T)>300$, according to our experiments). Importantly, a divergent penalty suggested by bai2004estimating to address non-stationary factors in PC makes our ICs perform even worse in the case of CAs. Therefore, our recommendation for CCE/CAs practitioners is not to automatically take PC literature prescriptions and interpret IC results with caution in the presence of highly persistent data, unless $N,T$ are substantial.