Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,081 characters · 7 sections · 37 citation commands
How well can we learn large factor models without assuming strong factors?
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
Models with hidden factors have been a popular tool for analyzing economic data. These models provide a convenient framework in describing datasets that are big in both cross-sectional and time series dimensions. The basic formulation is
where $M\in\mathbb{R}^{n\times T}$ is a low-rank matrix and $u_{i,t}\sim N(0,1)$ is i.i.d across $(i,t)$. The low rank property of $M$ is another way of describing the factor structure: if $M=LF'$ with $L\in\mathbb{R}^{n\times k}$ and $F\in\mathbb{R}^{T\times k}$, then ${\rm rank}\, M\leq k$. For this reason, we view factor models as a low-rank matrix plus an error matrix. Throughout the paper, all the idiosyncratic terms are assumed to be i.i.d and have a normal distribution.
Most of the results in the literature on factor models assume the strong factor condition or spiked eigenvalue condition, see Bai2003a,fan2011high,wang2017asymptotics among many others. This condition states that the largest $k$ singular values of $M$ grow at the rate $\sqrt{nT}$ or at least faster than $\sqrt{n+T}$. Besides the obvious question of whether such an assumption can be checked in the data, perhaps a more relevant question is whether we can assess the factor strength well enough for the purpose of estimation and inference. It is important to know how well we can possibly learn the factor models when the factor strength is also learned from the data. In this paper, we can answer this question in pure factor models and panel regression with interactive fixed effects. We study the problem of learning scalar (low-dimensional) components, e.g., entries of large matrices or regression coefficients.
The conceptual tools we use are minimax rates and adaptivity. We mainly examine two issues.
In a pure factor model, the task is to learn entries in $M$, say $M_{1,1}$, where $X_{1,1}$ is potentially missing. We characterize the impact of factor strength on estimation and inference. The main findings are summarized as follows.
When the number of factors is known, the minimax rate for estimation of $M_{1,1}$ depends on the factor strength (singular values of $M$). If the singular values of $M$ does not grow faster than $\sqrt{n+T}$, then it is impossible to achieve consistency in estimating $M_{1,1}$. Moreover, the rate of convergence for $M_{1,1}$ can be learned from the data: one can construct confidence intervals that are valid for arbitrary factor strength and automatically achieve the optimal rate in their widths.
When the number of factor is not given a priori, there is a tradeoff between validity and efficiency. Any confidence interval that is valid for arbitrary factor strength cannot have widths that shrink to zero even when applied to strong factor settings. Conversely, if a confidence interval learns the number of factors from the data and has widths shrinking to zero when the all the factors are strong, then this confidence interval cannot have uniform coverage for all factor strengths and the worst-case coverage probability is below $1/2$. This tradeoff only applies to inference. Achieving the minimax rate via a data-driven estimator is entirely possible. The key insight is that although ignoring weak factors does not reduce the rate in estimation, it is quite damaging for inference valdity.
In a panel regression with interactive fixed effects, the situation is much better. The fixed effects in these models have a factor structure and the task is to learn the regression coefficient. We show that even if the exact number of factors is unknown, it is possible to learn the regression coefficient at the rate $(nT)^{-1/2}$, regardless of the factor strengths in the fixed effects. However, when factors are allowed to be weak, uncertainty in the number of factors can cause a great loss of efficiency. This is in drastic contrast with existing results that under the strong factor condition, it is not important to know the exact number of factors, see moon2014linear.
Our work is related to the literature on weak factors. Problems of weak factors have been documented in simulations boivin2006more,bai2008large that study the performance of standard asymptotic theories. Theoretically, the inconsistency of PCA (principal component analysis) has been pointed out by johnstone2009consistency and onatski2012asymptotics among others. onatski2012asymptotics also derived detailed asymptotic distribution of PCA when the factors are weak. Inconsistency under weak factors is one of the main motivations driving developments in the sparse PCA literature. There sparsity is imposed on factor loadings to achieve consistent estimation of the factor structure, see amini2008high,berthet2013optimal,cai2013sparse,birnbaum2013minimax among many others. It turns out that even when the factors are observed, estimation and inference could encounter non-standard difficulties if the factors are only weakly influential, kleibergen2009tests,gospodinov2017spurious,anatolyev1807.04094. The situation is more complicated if there are potentially missing factors. We will now proceed to analysis of different models. Additional discussion on related literature will be presented in Sections (ref) and (ref) regarding pure factor models (with missing values) and panel regressions, respectively.
Notations. For a matrix $A$, $\sigma_{1}(A)\geq\sigma_{2}(A)\geq\cdots$ denote the singular values of $A$ in decreasing order. For matrix $A$, $\|A\|_{\infty}=\max_{i,j}|A_{i,j}|$, $\|A\|=\sigma_{1}(A)$ denotes the spectral norm, $\|A\|_{F}$ denotes the Frobenius norm and $\|A\|_{*}$ denotes the nuclear norm (sum of all singular values of $A$). For a vector $a$, $\|a\|_{2}$ denotes the Euclidean norm. In model ((ref)), the distribution of the data is indexed by $M$; the probability measure and expectation under $M$ are denoted by $\mathbb{P}_{M}$ and $\mathbb{E}_{M}$, respectively. For two positive sequences $X_{n},Y_{n}$, we use $X_{n}\lesssim Y_{n}$ to denote $X_{n}\leq CY_{n}$ for a constant. We use $X_{n}\gtrsim Y_{n}$ to denote $Y_{n}\lesssim X_{n}$ and $X_{n}\asymp Y_{n}$ means that $X_{n}\lesssim Y_{n}$ and $X_{n}\gtrsim Y_{n}$. We use $\mathbf{1}\{\cdot\}$ for the indicator function. For a matrix $A\in\mathbb{R}^{n\times T}$, $A_{-1,-1}$ denotes the matrix $A$ with its (1,1) entry replaced by zero. For a vector $a=(a_{1},...,a_{k})'$, we denote $a_{-1}=(a_{2},...,a_{k})'$.
In this section, we work with the following the parameter space: \[ \mathcal{M}=\left\{ A\in\mathbb{R}^{n\times T}:\ \|A\|_{\infty}\leq\kappa\quad\text{and}\quad{\rm rank}\, A\leq2\right\} , \] where $\kappa>0$ is a constant. The set $\mathcal{M}$ corresponds to factor models with at most two factors. We assume that entries of the factor structure are bounded so the factor strength would be a meaningful quantity. This is related to the incoherence condition; see Bai1910.06677 for an excellent discussion on the relation between incoherence and factor strength. Here, we impose bounded $\|A\|_{\infty}$ simply to rule out the situations in which the entire factor structure concentrates on a few entries. For $M\in\mathbb{R}^{n\times T}$ with $M_{i,t}=nT\mathbf{1}\{(i,t)=(1,1)\}$, the strong factor condition holds $\sigma_{1}(M)\asymp\sqrt{nT}$ but clearly the standard asymptotic theory for principal component analysis (PCA) in Bai2003a does not hold even if all entries are observed; entries in $X_{-1,-1}$ obviously have no information on $M_{1,1}$.
We define the set of one-factor models with factor strength $\tau$: \[ \mathcal{M}(\tau)=\left\{ A\in\mathcal{M}:\ \sigma_{1}(A)\geq\tau\quad\text{and}\quad\sigma_{2}(A)=0\right\} . \]
Our analysis of the case with known number of factors will deal with $\mathcal{M}(\tau)$ and later discussion on unknown number of factors mainly focuses on whether the number of factors is one or two. Restricting the analysis to at most two factors is for simplicity and the results can be extended to any fixed number of factors with extra (but perhaps unnecessary) complications in notations and proofs. In this section, we assume that the idiosyncratic terms are i.i.d standard normal variables for simplicity.\footnote{The assumption of unit variance is not too restrictive since the quantity $(nT)^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}u_{i,t}^{2}$ can always be consistently estimated at the rate $\max\{n^{-1},T^{-1}\}$, regardless of the factor strength. To see this, consider the PCA estimator $\hat{M}$ with rank $\bar{k}$, assuming $\bar{k}\geq k$. Since $\|X-\hat{M}\|\leq\|u\|$, we have $\|\hat{M}-M\|\leq2\|u\|$. Since ${\rm rank}\,(\hat{M}-M)\leq2\bar{k}$, we have $\|\hat{M}-M\|_{F}\leq2\sqrt{2\bar{k}}\|u\|=O_{P}(\sqrt{n+T})$ due to results from random matrix theory.}
One example that involves missing data is related to the synthetic control problem. This is a fast growing literature since the seminal papers by abadie2003economic,abadie2015comparative and has attracted much attention in applied research. One of the leading models for synthetic control is to use a factor model for the counterfactuals, see abadie2015comparative,gobillon2016regional,li2017statistical,athey2017matrix,xu2017generalized,chernozhukov1712.09089,li2018inference among many others. We consider a simple formulation.
Like other causal inference problems, learning $\gamma$ is primarily the task of learning the counterfactual $Y_{1,T}^{N}$, which is unobserved. If $Y_{1,T}^{N}$ were known, then the obvious 95% confidence interval for $\gamma$ is $[Y_{1,T}-Y_{1,T}^{N}-1.96,Y_{1,T}-Y_{1,T}^{N}+1.96]$. Since $Y_{1,T}^{N}$ is unobserved, we can replace it with a consistent estimate. However, there is no guarantee that such an estimate exists when the factors are not strong. If consistent estimation for $Y_{1,T}^{N}$ is impossible, then a 95% confidence interval for $\gamma$ needs to have a width larger than $1.96\times2$. To formally investigate the estimation and inference for $Y_{1,T}^{N}$, we consider the following problem.
The main feature of Example (ref) is that the missing pattern is non-random. Since there is only one entry missing, we essentially observe the entire data. For this reason, the problem is very closely related to the following one.
Example (ref) corresponds to the problem of estimation and inference in standard large factor models. Many classical results are developed by Bai2003a and the references therein. In fact, the problem missing data in factor models also has a long history dating back to at least stock2002macroeconomic. Very recently, works by su2019factor,Bai1910.06677,xiong1910.08273 provided extensive asymptotic theories for various missing data problems under the strong factor condition. However, when the factors are not assumed to be strong, it is still unclear what can be done and what is impossible.
Example (ref) is also called the matrix completion problem. This literature aims to estimate $M$ from $(\tilde{X},\Xi)$ and most of the results are stated in terms of the overall risk in Frobenius norm: construct an estimate $\hat{M}$ and provide bound on $\|\hat{M}-M\|_{F}$, see candes2009exact,candes2010power,recht2010guaranteed,keshavan2010matrix,candes2011tight,rohde2011estimation,koltchinskii2011nuclear among others.
These results might seem quite encouraging in that they derive the bounds without the strong factor condition. In particular, one can construct an estimate $\hat{M}$ and guarantee \[ (nT)^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}(\hat{M}_{i,t}-M_{i,t})^{2}=o_{P}(1) \] without any assumption on the factor strength. Since the average (across entries) squared errors tends to zero, one might expect the estimation error for a typical entry to be small without the strong factor condition, e.g., $\hat{M}_{1,1}-M_{1,1}=o_{P}(1)$. We now show that this is not true.
We focus on the analysis of Example (ref). When the number of factors is known to be one, we consider the problem of how the strength of this factor affects the efficiency of estimation and inference for $M_{1,1}$. The basic problem of our analysis is to distinguish between \[ \mathcal{M}^{{\rm null}}(\tau):=\left\{ M\in\mathcal{M}(\tau):\ M_{1,1}=0\right\} \] and \[ \mathcal{M}_{\rho}(\tau):=\left\{ M\in\mathcal{M}(\tau):\ |M_{1,1}|\geq\rho\right\} \] for $\rho>0$. When the factors are strong, i.e., $\tau\gtrsim\sqrt{nT}$, results in Bai1910.06677 suggest that non-trivial power can be achieved when $\rho\gtrsim\min\{n^{-1/2},T^{-1/2}\}$. Our goal in this subsection is to answer the following questions:
In this section, we fix $\alpha\in(0,1)$. We first give an impossibility result concerning the first question.
In Theorem (ref), $\psi$ is a test for $H_{0}:\ M\in\mathcal{M}^{{\rm null}}(\tau)$ versus $H_{1}:\ M\in\mathcal{M}_{\rho}(\tau)$. A test is a measurable function that maps the data $X_{-1,-1}$ to $\{0,1\}$, where 1 denotes rejection of $H_{0}$. We allow for random tests by extending $\{0,1\}$ to the interval $[0,1]$. The requirement in ((ref)) ensures size control uniformly in the null space $\mathcal{M}^{{\rm null}}(\tau)$. The conclusion of Theorem (ref) states that when $\rho\lesssim\min\{1,\sqrt{n+T}/\tau\}$, no test can guarantee non-trivial power in distinguishing $\mathcal{M}^{{\rm null}}(\tau)$ and $\mathcal{M}_{\rho}(\tau)$.
As a consequence, it is impossible to achieve consistency in estimating $M_{1,1}$ if $\tau\lesssim\sqrt{n+T}$. To see this, notice that in this case, there exists a constant $\rho_{0}>0$ such that no test can have high power in testing $H_{0}:\ M_{1,1}=0$ against $H_{1}:\ |M_{1,1}|\geq\rho_{0}$. Hence, if a consistent estimator $\hat{M}_{1,1}$ exists, then the simple test of $\mathbf{1}\{|\hat{M}_{1,1}|\leq\rho_{0}/2\}$ would have both Type I error and Type II error tending to zero for $H_{0}:\ M_{1,1}=0$ against $H_{1}:\ |M_{1,1}|\geq\rho_{0}$. However, since such a test is impossible (due to Theorem (ref)), a consistent estimator for $M_{1,1}$ does not exist.
Another consequence of Theorem (ref) is that the factor strength has a direct impact on the rate of learning $M_{1,1}$. In order to obtain the usual rate of $\min\{n^{-1/2},T^{-1/2}\}$, the strong factor assumption of $\tau\asymp\sqrt{nT}$ is necessary. If $\tau\ll\sqrt{nT}$, the rate for $M_{1,1}$ is strictly worse than $\min\{n^{-1/2},T^{-1/2}\}$.
We now present an implication of Theorem (ref) on confidence intervals. Although this implication is immediate, the notations we introduce will be very useful in discussing later issues, especially with unknown number of factors. For a given parameter space $\mathcal{M}^{(1)}\subseteq\mathcal{M}$, we define the set of valid $1-\alpha$ confidence intervals as measurable functions that map the data $X_{-1,-1}$ to an interval such that this interval covers $M_{1,1}$ with probability at least $1-\alpha$ for all parameters in $\mathcal{M}^{(1)}$. In other words, we define \[ \Phi(\mathcal{M}^{(1)})=\left\{ CI(\cdot)=[l(\cdot),u(\cdot)]:\ \inf_{M\in\mathcal{M}^{(1)}}\mathbb{P}_{M}\left(M_{1,1}\in CI(X_{-1,-1})\right)\geq1-\alpha\right\} . \]
The minimax expected length of confidence intervals over $\mathcal{M}^{(1)}$ is \[ \mathcal{L}(\mathcal{M}^{(1)})=\inf_{CI\in\Phi(\mathcal{M}^{(1)})}\sup_{M\in\mathcal{M}^{(1)}}\mathbb{E}_{M}|CI(X_{-1,-1})|, \] where $|CI|$ denotes the length of the confidence interval $CI$, i.e., $|CI(\cdot)|=u(\cdot)-l(\cdot)$ for $CI(\cdot)=[l(\cdot),u(\cdot)]$. Minimax rate for the length of confidence intervals is a common way of examining the efficiency for inference in high-dimensional models\footnote{Compared to stating results in terms of size and power, discussing length of confidence intervals brings simplicity mainly because we do not need a notation on the hypothesized value in the null.}, cai2017confidence,bradic1802.09117. Theorem (ref) has the following simple implication; we omit the proof for brevity.
Corollary (ref) states a lower bound for $\mathcal{L}(\mathcal{M}(\tau))$. We now show that this lower bound is also optimal. Since consistency is impossible when $\tau_{n}\lesssim\sqrt{n+T}$, we focus on $\tau_{n}\gg\sqrt{n+T}$. (If $\tau\lesssim\sqrt{n+T}$, we can simply use the simple confidence interval $[-\kappa,\kappa]$ assuming $\kappa$ is known.) We first construct an estimator that has the rate of convergence $\sqrt{n+T}/\tau$ when $\tau_{n}\gtrsim\sqrt{n+T}$. The idea is based on the following the observation: \[ M_{1,1}=\frac{L_{1}L_{-1}'M_{-1,1}}{\|L_{-1}\|_{2}^{2}}, \] where $M_{-1,1}=(M_{2,1},...,M_{n,1})'\in\mathbb{R}^{n-1}$. Hence, a plug-in estimator would be
where $X_{-1,1}=(X_{2,1},...,X_{n,1})'\in\mathbb{R}^{n-1}$. Here, $\hat{L}=(\hat{L}_{1},\hat{L}_{-1}')'\in\mathbb{R}^{n}$ is the PCA estimator for $L$ using data $X_{,-1}=\{X_{i,t}\}_{1\leq i\leq n,\ 2\leq t\leq T}$. This estimator can be viewed as a version of the tall-wide estimator by Bai1910.06677. The proof in our case is significantly complicated by the fact that the factor strengths might not be strong. This leads to difficulties in bounding certain terms; we use a novel construction that allows us to apply a decoupling argument, see Lemma (ref) in the appendix. The following result establishes its rate of convergence under any given factor strength.
Fortunately, distinguishing between $\tau\gg\sqrt{n+T}$ and $\tau\lesssim\sqrt{n+T}$ is not difficult, thanks to bounds in random matrix theory. Since $\|M_{-1,-1}\|_{F}^{2}=\|M\|_{F}^{2}-M_{1,1}^{2}\geq\tau^{2}-\kappa^{2}$, $\|M_{-1,-1}\|_{F}$ and $\tau$ have the same rate. Since $M_{-1,-1}$ has rank at most 2, $\|M_{-1,-1}\|_{F}\geq\|M_{-1,-1}\|\geq\|M_{-1,-1}\|_{F}/\sqrt{2}$. Therefore, $\|X_{-1,-1}\|\gtrsim\|M_{-1,-1}\|-\|u_{-1,-1}\|=O_{P}(|\tau-\sqrt{n+T}|)$ by Corollary 5.35 of vershynin2010introduction. Hence, the following estimator always has the optimal rate of convergence:
where $\bar{\kappa}>0$ is a constant that upper bounds $\kappa$. The idea is that if $\|X_{-1,-1}\|$ is smaller than the above threshold, we can safely conclude $\tau\lesssim\sqrt{n+T}$ and thus no estimator is consistent for $M_{1,1}$, which means that the estimator zero would have the optimal “rate” too (since $M_{1,1}$ is bounded). Therefore, we have the following result.
Notice that the estimator in ((ref)) is adaptive in that it automatically achieves the optimal rate $\min\{1,\sqrt{n+T}/\tau\}$ without prior knowledge of $\tau$. Since we can learn the rate for $\tau$ from $\|X_{-1,-1}\|$, we can estimate the rate $\min\{1,\sqrt{n+T}/\tau\}$ from the data and thereby construct a confidence interval centered around the estimator in ((ref)): simply choose a width of $C_{0}\min\{\sqrt{n+T}/\|X_{-1,-1}\|,1\}$ for a universal constant $C_{0}$ (which can be explicitly determined). Therefore, the length of the confidence interval is also adaptive since it automatically adjusts to have the optimal rate. Unfortunately, this turns out to be true only when the number of factors is known. We will now see that when we do not know the exact number of factors (including weak ones), such adaptivity is impossible.
The previous analysis already establishes the optimal rate under knowledge of the exact number of factors. Now we aim to answer the following questions:
These questions arise naturally in practice as the researcher is typically not given the exact number of factors. In these cases, one often needs to estimate the number of factors from the data and then proceeds using this estimate.
Although such a strategy makes intuitive sense, it is still far from clear whether lack of knowledge on the number of factors has any cost. This problem is particularly tricky for inference since uniform size control (or coverage) is a concern; we do not have a notion of uniform validity for estimation. Suppose that we know there are either one or two factors. Ideally, we would like to construct a confidence interval that has $1-\alpha$ coverage probability uniformly over all one-factor models and all two-factor models, regardless of factor strengths. Procedures that involve a pre-estimation of the number of factors are intended to have such robustness. If there is one factor, then hopefully the procedure will find out that there is one factor and constructs the confidence interval accordingly; if there are two factors, then the procedure is supposed to detect that and constructs a valid confidence interval based on that finding. This leads to a confidence interval summarized in Figure (ref). Even for methods, such as nuclear-norm-regularized approaches, which do not need the number of factors as an input, we still hope that there is a data-driven procedure that would provide valid inference no matter what the true number of factors is. In light of this robustness requirement, we are interested in finding out whether these uniformly valid procedures lose any efficiency, compared to the situation with known number of factors.
We would not expect any loss of efficiency if the number of factors can be consistently estimated. Unfortunately, determining the number of weak factors is quite difficult (if not impossible) since typical methods bai2002determining,onatski2009testing,ahn2013eigenvalue only promise success in detecting strong factors. If we run these pre-tests and ignore potential weak factors, do we pay a price in terms of validity or efficiency?
To formally answer this question, we introduce the notion of inference adaptivity. Consider the class of one-factor models with factor strength $\tau_{0}$: \[ \mathcal{M}^{(1)}=\mathcal{M}(\tau_{0}). \]
From the previous analysis, we know that a confidence interval for $M_{1,1}$ with validity over $\mathcal{M}^{(1)}$ can have shrinking length only if $\tau_{0}\gg\sqrt{n+T}$. Since we do not know for sure that there is only one factor, we would like to consider a confidence interval that has validity over a larger space. For example, given $\tau_{1}\geq\tau_{2}>0$, we define \[ \mathcal{M}^{(2)}=\mathcal{M}(\tau_{1},\tau_{2}):=\left\{ A\in\mathcal{M}:\ \sigma_{1}(A)\geq\tau_{1},\ \sigma_{2}(A)\leq\tau_{2}\right\} , \] where $\tau_{1}\leq\tau_{0}$. Clearly, $\mathcal{M}^{(1)}\subset\mathcal{M}^{(2)}$.
The primary setting is $\tau_{1},\tau_{0}\gg\sqrt{n+T}$ and $\tau_{2}\lesssim\sqrt{n+T}$. In this setting, $\mathcal{M}^{(2)}$ allows for a second weak factor. The exercise is to consider all the confidence intervals that have uniform coverage over $\mathcal{M}^{(2)}$ and then choose the one that has the best efficiency over $\mathcal{M}^{(1)}$ (in terms of expected length). The questions listed in the two bullet points above now become an issue of adaptivity. If there is a confidence interval that is valid over $\mathcal{M}^{(2)}$ and has length automatically achieves the rate of $\sqrt{n+T}/\tau_{0}$ over $\mathcal{M}^{(1)}$, then (at least for the rate) lack of knowledge on the number of factors does not have a cost and in particular we do not pay a price in terms of size and power for failing to detect weak factors. If any confidence interval that has validity over $\mathcal{M}^{(2)}$ necessarily has a length greater than $O(\sqrt{n+T}/\tau_{0})$ over $\mathcal{M}^{(1)}$, then allowing for the extra validity on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$ causes an efficiency loss. In particular, we are interested in the following quantity \[ \mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})=\inf_{CI\in\Phi(\mathcal{M}^{(2)})}\sup_{M\in\mathcal{M}^{(1)}}\mathbb{E}_{M}|CI(Y_{-1,-1})|. \]
The above quantity is the best guaranteed expected length over $\mathcal{M}^{(1)}$ among all the confidence intervals with uniform validity on $\mathcal{M}^{(2)}$. The minimax rate $\mathcal{L}(\mathcal{M}^{(1)})$ defined before is a special case since $\mathcal{L}(\mathcal{M}^{(1)})=\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(1)})$. Notice that by construction, we have $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\geq\mathcal{L}(\mathcal{M}^{(1)})$. If $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\gg\mathcal{L}(\mathcal{M}^{(1)})$, then the additional validity requirement on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$ has a negative impact on the efficiency on $\mathcal{M}^{(1)}$. If $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\asymp\mathcal{L}(\mathcal{M}^{(1)})$, then we say that it is possible to achieve adaptive inference; the procedure can automatically recognize that the parameter is in the “fast-rate-region” $\mathcal{M}^{(1)}$ and constructs the confidence interval accordingly.
We now show that such adaptivity is unfortunately impossible for factor models when the number of factors is unknown.
In this paper, confidence intervals with validity over $\mathcal{M}^{(2)}$ will be referred to as robust confidence intervals (or weak-factor-robust confidence intervals). Theorem (ref) has a striking implication when $\tau_{1},\tau_{0}\asymp\sqrt{nT}$ and $\tau_{2}$ is bounded away from zero. In this case, no robust confidence intervals can guarantee to have shrinking width when there is actually only one factor and this factor is strong; recall that in this case, $\mathcal{L}(\mathcal{M}^{(1)})\asymp\sqrt{n+T}/\tau_{0}\asymp\min\{n^{-1/2},T^{-1/2}\}$. Hence, $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\gg\mathcal{L}(\mathcal{M}^{(1)})$. In other words, robustness necessarily causes efficiency loss. Unlike the adaptivity in Remark (ref) for linear IV models, no confidence intervals valid over $\mathcal{M}^{(2)}$ can adapt its efficiency on $\mathcal{M}^{(1)}$. When $\tau_{0},\tau_{1},\tau_{2}\asymp\sqrt{nT}$, robust confidence intervals will have lengths tending to zero only if there are two factors and both are strong, but in this case we are back to the situation with known number of factors.
The other side of this coin is the observation that any confidence interval that has shrinking width and ignores weak factor necessarily lacks uniform coverage. We state an explicit result below.
Theorem (ref) reveals the danger of some popular methods in practice. Consider the case with $\tau_{0},\tau_{1},\tau_{2}\asymp\sqrt{nT}$ (so $\tau_{0},\tau_{2}\gtrsim1$ is satisfied). Recall the procedure in Figure (ref) under the simplified assumption that the number of factors is either one or two. By classical results, the PCA estimate is asymptotically normal over $\mathcal{M}^{(1)}$ with a standard error shrinking to zero. Hence, the overall procedure in Figure (ref), which uses these standard errors, produces a random interval with shrinking length on $\mathcal{M}^{(1)}$. However, Theorem (ref) indicates that precisely due to its shrinking width on $\mathcal{M}^{(1)}$, the procedure in Figure (ref) cannot have uniform coverage over $\mathcal{M}^{(2)}$ unless the number of factors (including weak ones) can be consistently determined. Since this result holds regardless of which pre-test is used, developing better tests for the number of factors might not significantly improve inference quality. Moreover, using methods that do not require the number of factors as an input (e.g., nuclear-norm penalized methods) does not solve the problem either since Theorem (ref) does not assume a specific form for $\mathcal{I}_{*}$. In this regard, Theorem (ref) also serves as a simple check on the uniform validity of a given procedure: if there exists two different numbers $k_{1}$ and $k_{2}$ such that the procedure under consideration gives shrinking confidence intervals in both cases ((1) when there are $k_{1}$ strong factors and no weak factors and (2) when there are $k_{2}$ strong factors and no weak factors), then this procedure does not have uniform validity.
Theorems (ref) and (ref) indicate an inherent tradeoff between efficiency and validity. On one hand, Theorem (ref) states that procedures with validity over $\mathcal{M}^{(2)}$ cannot provide accurate inference on $\mathcal{M}^{(1)}$. On the other hand, Theorem (ref) makes the equivalent claim that procedures that provide accurate inference on $\mathcal{M}^{(1)}$ cannot guarantee uniform validity on $\mathcal{M}^{(2)}$. Therefore, requiring validity on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$ necessarily reduces the inference efficiency.
Since Theorem (ref) only involves the worst-case width (or power) on $\mathcal{M}^{(1)}$. One might wonder whether robust procedures could still have decent efficiency on average over $\mathcal{M}^{(1)}$ despite the bad worst-case performance. We now show that this is not the case: lack of efficiency occurs at many points in $\mathcal{M}^{(1)}$.
Since $\eta$ can be chosen arbitrarily, $\mathcal{M}_{*}^{(1)}$ represents quite many (if not most) points in $\mathcal{M}^{(1)}$. Theorem (ref) states that the efficiency is bad at every point in $\mathcal{M}_{*}^{(1)}$ for any robust confidence interval. Therefore, the impossibility result in Theorem (ref) is not driven by only a few unlucky points in $\mathcal{M}^{(1)}$. This means that there is fundamental difficulty in entry-wise learning when the number of factors is unknown. Hence, in order to achieve efficient inference, the number of factors needs to be given a priori.
The panel data model with interactive fixed effects pesaran2006cross,bai2009panel assumes \[ Y_{i,t}=L_{i}'F_{t}+X_{i,t}\beta+\varepsilon_{i,t}, \] where the observed data is $\{(Y_{i,t},X_{i,t})\}_{1\leq i\leq n,\ 1\leq t\leq T}$. Here, $\varepsilon_{i,t}$ is i.i.d across $(i,t)$ following $N(0,\sigma_{\varepsilon}^{2})$ and the factor structure $L_{i}'F_{t}$ is the fixed effect with $L_{i},F_{t}\in\mathbb{R}^{r_{0}}$. For simplicity, here $X_{i,t}$ and $\beta$ are scalars. The main requirement of $X_{i,t}$ is that it cannot be absorbed by the fixed effects. Since the fixed effects have a factor structure, we assume that $X_{i,t}$ is a factor structure plus non-negligible noise: \[ X_{i,t}=\alpha_{i}'g_{t}+u_{i,t}, \] where $\alpha_{i},g_{t}\in\mathbb{R}^{r_{1}}$ and $u_{i,t}$ is i.i.d across $(i,t)$ following $N(0,\sigma_{u}^{2})$. We assume that $u$ and $\varepsilon$ are mutually independent. Factor structures in the regressor $X_{i,t}$ have been a common assumption pesaran2006cross,moon2014linear,zhu2017high,chernozhukvo2018panel. Under these assumptions, we can write the model as
where $M,D,\varepsilon,u\in\mathbb{R}^{n\times T}$, ${\rm rank}\, M\leq r_{0}$ and ${\rm rank}\, D\leq r_{1}$. The distribution of the data $(Y,X)$ is thus indexed by $\theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta)$. We consider the following parameter space: \[ \Theta=\left\{ \theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta):\ \ {\rm rank}\, M\leq r_{0},\ {\rm rank}\, D\leq r_{1},\ \{\sigma_{\varepsilon},\sigma_{u}\}\subset[\kappa^{-1},\kappa],\ |\beta|\leq\kappa\right\} , \] where $\kappa>0$ is a constant and $r_{1},r_{2}>0$ are fixed. The requirement of $\{\sigma_{\varepsilon},\sigma_{u}\}\subset[\kappa^{-1},\kappa]$ is merely saying that $\sigma_{\varepsilon}$ and $\sigma_{u}$ are bounded away from zero and infinity. For simplicity, we shall also require that $\beta$ be bounded. Notice that we do not require boundedness of $\|M\|_{\infty}$ and $\|D\|_{\infty}$; it turns out that such a requirement will not be needed. We note that $r_{0}$ and $r_{1}$ are only upper bounds on the number of factors, instead of the exact number of factors. Moreover, there is no requirement on the relative magnitude between $n$ and $T$.
The most important feature of $\Theta$ is that there is no assumption at all regarding the strength of the factors. For estimating $\beta$, bai2009panel showed that when all the factors are strong and the number of factors is known (or consistently estimable), one can achieve the rate $(nT)^{-1/2}$. moon2014linear showed that when all the factors are strong, overstating the number of factors does not have any impact on the rate for estimating $\beta$. These results still leave two open questions that are quite relevant in practice:
We now provide an estimator that achieves the rate $(nT)^{-1/2}$ uniformly over $\Theta$. As a result, lack of knowledge on the number of factors and potentially weak factors do not create any problem for the rate on learning $\beta$. Notice that the rate $(nT)^{-1/2}$ is minimax optimal by a simple two-point argument.
The basic idea of our estimator is as follows. The assumption in ((ref)) implies a factor structure in $Y$: \[ Y=(M+D\beta)+V, \] where ${\rm rank}\,(M+D\beta)\leq r_{0}+r_{1}$ and $V=u\beta+\varepsilon$. Thus, we can identify $\beta$ as $\beta=\mathbb{E}{\rm trace\,}(V'u)/(nT\sigma_{u}^{2})$. Since both $V$ and $u$ are idiosyncratic parts in $Y$ and $X$, respectively, we can estimate them by the typical low-rank estimation strategies.
Without loss of generality, we assume that $T\geq n$; if $T<n$, then we flip our data from $(Y,X)$ to $(Y',X')$. For any matrix $A\in\mathbb{R}^{n\times r}$, we define $\Pi_{A}=I_{n}-P_{A}$ and $P_{A}=A(A'A)^{\dagger}A'$, where $^{\dagger}$ denoting the Moore-Penrose pseudo-inverse. Let $\hat{\alpha}\in\arg\min_{A\in\mathbb{R}^{n\times r_{1}}}{\rm trace\,}(X'\Pi_{A}X)$ and $\hat{\Lambda}\in\arg\min_{A\in\mathbb{R}^{n\times k}}{\rm trace\,}(Y'\Pi_{A}Y)$ with $k=r_{0}+r_{1}$, Notice that $\hat{\alpha}$ and $\hat{\Lambda}$ are simply the eigenvectors corresponding to the largest $r_{1}$ and $k$ eigenvalues of $XX'$ and $YY'$, respectively. Then the estimator is defined as
where $\hat{r}=k+r_{1}-{\rm trace\,}(P_{\hat{\Lambda}}P_{\hat{\alpha}})$.
We make three comments on the estimator. First, due to the identification condition of $\beta=\mathbb{E}{\rm trace\,}(V'u)/(nT\sigma_{u}^{2})$, we can view the estimation problem as learning the expected conditional covariance newey_fast_rate1801.09138,chernozhukov2018learning,chernozhukov1802.08667 and thus construct a solution in a similar spirit. The means of $X$ and $Y$ are high-dimensional structures: $A$ and $M+D\beta$. Since the trace operation defines an inner product, ${\rm trace\,}(\mathbb{E} V'u)$ is a linear functional of the covariance $\mathbb{E}[Y-(M+D\beta)]'[X-A]$. The simple plug-in approach is to construct an estimate for $M+D\beta$ and $A$ and replace $\mathbb{E}(\cdot)$ with $(nT)^{-1}{\rm trace\,}(\cdot)$. Under PCA, the estimate for $Y-(M+D\beta)$ and $X-D$ are $\Pi_{\hat{\Lambda}}Y$ and $\Pi_{\hat{\alpha}}X$, respectively.
Second, the quantity $(n-r_{1})$ acts as bias correction for ${\rm trace\,}(X'\Pi_{\hat{\alpha}}X)$ in ((ref)). Although the random matrix theory can gives us ${\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}nT+O_{P}(n+T)$, this bound is not enough as it only yields \[ (nT)^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}+O_{P}(\min\{n^{-1},T^{-1}\}). \] To achieve the rate $(nT)^{-1/2}$, we would need an estimate for $\sigma_{u}^{2}$ at the rate $O_{P}((nT)^{-1/2})$, which is faster than $O_{P}(\min\{n^{-1},T^{-1}\})$ unless $n\asymp T$. We derive a more accurate characterization by showing \[ [(n-r_{1})T]^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}+O_{P}((nT)^{-1/2}). \] Therefore, $[T(n-r_{1})]^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)$ has a strictly smaller remainder term unless $T\asymp n$. Similarly, the quantity $n-\hat{r}$ acts as bias correction for ${\rm trace\,}(Y'\Pi_{\hat{\Lambda}}\Pi_{\hat{\alpha}}X)$ since one can show ${\rm trace\,}(Y'\Pi_{\hat{\Lambda}}\Pi_{\hat{\alpha}}X)=(n-\hat{r})T\sigma_{u}^{2}\beta+O_{P}(\sqrt{nT})$.
Third, a key step in analyzing the numerator in ((ref)) is to show $\|\Pi_{\hat{\alpha}}D\|_{F}^{2}=O_{P}((nT)^{1/2})$ regardless of the factor strength. To see why this is crucial, we note that the best guaranteed rate for $\|D-\hat{D}\|_{F}^{2}$ is $\max\{n,T\}$ for any estimator $\hat{D}$, rohde2011estimation,candes2011tight. This rate is worse than $(nT)^{1/2}$ unless $T\asymp n$. One novelty of our analysis is to show that although $\|\Pi_{\hat{\alpha}}X\|_{F}^{2}=\|\Pi_{\hat{\alpha}}(D+u)\|_{F}^{2}=O_{P}(\max\{n,T\})$, we can separate it into $\|\Pi_{\hat{\alpha}}D\|_{F}^{2}=O_{P}((nT)^{1/2})$ and $\|\Pi_{\hat{\alpha}}u\|_{F}^{2}=O_{P}(\max\{n,T\})$.
The condition $T\geq n$ is motivated by the following insight. Under our factor structure, it suffices to estimate either the factors or the factor loadings, not both. Hence, perhaps we should estimate the one with lower dimensionality. If $T\gg n$, then the factor loading whose dimensionality is proportional to $n$ would be easier to estimate, compared to factors whose dimensionality scales with $T$.
We note that the requirement of $T\geq n$ is completely innocuous and can be removed if we consider the following estimator \[ \tilde{\beta}=
\]
Theorem (ref) formally establishes the uniform rate of $(nT)^{-1/2}$ over $\Theta$. Hence, whether factors are strong or weak and whether the number of factors is exactly known would not prevent us from achieving the rate $(nT)^{-1/2}$.
Since the minimax rate of learning $\beta$ does not depend on the strong factor assumption, the natural question is whether the strong factor assumption is not important at all for estimation and inference. Unfortunately, the answer is no. We shall show that (1) the efficiency of inference on $\beta$ crucially depends on the strong factor condition and (2) there is lack of adaptivity in the factor strength, resulting in a tradeoff between efficiency and robustness.
From the perspective of semiparametric estimation, there is a good reason to suspect that the strong factor condition might affect the inference efficiency. We shall view the panel regression problem in ((ref)) as a semiparametric problem, which would yield a natural semiparametric lower bound for the asymptotic variance in estimating $\beta$. However, we then realize that the typical asymptotic variance in the literature bai2009panel,moon2014linear assuming strong factors can be much smaller than this lower bound. This leads us to suspect that the strong factor condition might play a role similar to strong parametric restrictions on nonparametric components of semiparametric models.
To provide an analogy, consider the partial linear model with observations $(Y_{i},Z_{i},W_{i})$ generated by $Y_{i}=f(W_{i})+Z_{i}\beta+\varepsilon_{i}$ and $Z_{i}=g(W_{i})+u_{i}$; under regularity conditions, the asymptotic variance of estimating $\beta$ is bounded below by $\mathbb{E}\varepsilon_{i}^{2}/\mathbb{E} u_{i}^{2}$ when $f$ and $g$ are nonparametric functions or functions with high-dimensional parameters, see e.g., chernozhukov1802.08667,newey2018cross,jankova2018semiparametric. Now we recall the model in ((ref)): $Y=M+X\beta+\varepsilon$ and $X=D+u$, where $M,D\in\mathbb{R}^{n\times T}$ are high-dimensional nuisance parameters. We can view $M$ and $D$ as $f(W_{i})$ and $g(W_{i})$ in the partial linear model, respectively. From this perspective, we would expect the asymptotic variance for estimating $\beta$ to be at least $\sigma_{\varepsilon}^{2}\sigma_{u}^{-2}$ in general. However, the asymptotic variance derived in bai2009panel and moon2014linear is \[ \frac{\sigma_{\varepsilon}^{2}}{\sigma_{u}^{2}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)}, \] where $\Pi_{M}=I_{n}-M(M'M)^{\dagger}M'$ and $\Pi_{M'}=I_{T}-M'(MM')^{\dagger}M$.
To formally address the efficiency problem, we again adopt the framework of adaptivity. For simplicity, we assume that $\sigma_{u}=\sigma_{\varepsilon}=1$ in ((ref)). We consider the parameter space \[ \Theta^{(2)}=\left\{ \theta=(M,D,1,1,\beta):\ {\rm rank}\, M\leq2,\ {\rm rank}\, D=1,\ |\beta|\leq1\right\} . \] Then we focus on adaptivity on a smaller space in which the strong factor condition holds:
where $\kappa_{1},\kappa_{2}>0$ are constants. The difference between $\Theta^{(1)}$ and $\Theta^{(2)}$ is that $\Theta^{(2)}$ allows for potentially weak factors and uncertainty in the number of factors (${\rm rank}\, M$ can be either 1 or 2), while $\Theta^{(1)}$ only considers known number of factors and assumes all the factors are strong. Our analysis will focus on the question of whether robust confidence intervals (uniform validity on $\Theta^{(2)}$) has worse efficiency on $\Theta^{(1)}$ than non-robust confidence intervals (uniform validity only on $\Theta^{(1)}$).
By bai2009panel and moon2014linear (among others), the least-square estimator $\hat{\beta}_{{\rm LS}}$ (i.e., $(\hat{\beta}_{{\rm LS}},\hat{A})=\arg\min_{\beta\in\mathbb{R},\ A\in\mathbb{R}^{n\times T},\ {\rm rank}\, A\leq2}\|Y-A-X\beta\|_{F}^{2}$) satisfies \[ \frac{\sqrt{nT}(\hat{\beta}_{{\rm LS}}(X,Y)-\beta)}{\sigma(\theta)}\overset{d}{\rightarrow}N(0,1), \] over $\theta\in\Theta^{(1)}$, where for $\theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta)$, \[ \sigma(\theta):=\frac{\sigma_{\varepsilon}}{\sqrt{\sigma_{u}^{2}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)}}. \]
For $\theta\in\Theta^{(1)}$, we have that \[ \sigma(\theta)=\frac{\sigma_{\varepsilon}}{\sqrt{\sigma_{u}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)}}=\frac{1}{\sqrt{1+\|D\|_{F}^{2}/(nT)}}\leq\frac{1}{\sqrt{1+\kappa_{2}^{2}}}. \]
Therefore, a natural 95%-confidence interval for parameters in $\Theta^{(1)}$ is
Existing results imply that \[ \liminf_{n,T\rightarrow\infty}\inf_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}\left(\beta\in CI_{*}(X,Y)\right)\geq95\%. \]
In other words, we have that
where the quantity $\mathcal{L}(\Theta^{(1)})$ is defined as before: $\mathcal{L}(\Theta^{(1)})=\inf_{CI\in\Phi_{0.95}(\Theta^{(1)})}\sup_{\theta\in\Theta^{(1)}}\mathbb{E}_{\theta}|CI(X,Y)|$ is the minimax expected length of confidence intervals on $\Theta^{(1)}$ and $\Phi_{0.95}(\Theta^{(1)})=\{CI(\cdot)=[l(\cdot),u(\cdot)]:\ \inf_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}(\beta\in CI(X,Y))\geq0.95\}$ is the set of 95%-confidence intervals. We also consider robust confidence intervals, which have uniform validity over $\Theta^{(2)}$ and form the set $\Phi_{0.95}(\Theta^{(2)})$. To study the impact of the robustness requirement on efficiency, we revisit the concept of adaptivity by studying \[ \mathcal{L}(\Theta^{(1)},\Theta^{(2)})=\inf_{CI\in\Phi_{0.95}(\Theta^{(2)})}\sup_{\theta\in\Theta^{(1)}}\mathbb{E}_{\theta}|CI(X,Y)|. \]
Both $\mathcal{L}(\Theta^{(1)})$ and $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$ measure performance of confidence intervals on $\Theta^{(1)}$. The former considers non-robust confidence intervals (ones with validity over $\Theta^{(1)})$, whereas the latter considers robust confidence intervals (with validity over the larger set $\Theta^{(2)})$. If $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})/\mathcal{L}(\Theta^{(1)})$ is asymptotically larger than one, then the extra robustness on $\Theta^{(2)}\backslash\Theta^{(1)}$ decreases the efficiency even on $\Theta^{(1)}$; if $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})/\mathcal{L}(\Theta^{(1)})$ converges to one, then one can gain extra robustness without sacrificing efficiency. To characterize $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$, we first derive the following result.
Notice that the above lower bound does not depend on $\kappa_{2}$. On the other hand, $\sqrt{nT}|CI_{*}(X,Y)|\asymp(1+\kappa_{2}^{2})^{-1/2}$ decreases with $\kappa_{2}$. Thus, for large enough $\kappa_{2}$, $CI_{*}$ in ((ref)) violates the lower bound in Theorem (ref). Since the lower bound is satisfied by any confidence interval with uniform validity over $\Theta^{(2)}$, it follows that any robust confidence interval will be wider than $CI_{*}$ on $\Theta^{(1)}$. Equivalently, we can state the result in terms of robustness (coverage guarantee for $CI_{*}$).
Notice that $CI_{*}$ satisfies ((ref)) since $|CI_{*}(X,Y)|=3.92(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2}$. By Corollary (ref), any confidence interval that has a width similar to (or shorter than) that of $CI_{*}$ on $\Theta^{(1)}$ will have coverage probability close to 1/2 on $\Theta^{(2)}$. Finally, we compare $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$ and $\mathcal{L}(\Theta^{(1)})$. Applying Theorem (ref) with $\alpha=0.05$, we obtain obtain \[ \mathcal{L}(\Theta^{(1)},\Theta^{(2)})\geq(nT)^{-1/2}\left(1-2\alpha\right)^{2}/2=0.405(nT)^{-1/2}. \]
In light of ((ref)), this means \[ \liminf_{n,T\rightarrow\infty}\frac{\mathcal{L}(\Theta^{(1)},\Theta^{(2)})}{\mathcal{L}(\Theta^{(1)})}\geq\frac{0.405}{3.92}\sqrt{1+\kappa_{2}^{2}}>\frac{\sqrt{1+\kappa_{2}^{2}}}{9.7}. \]
Therefore, any robust confidence interval is asymptotically wider than any non-robust confidence interval whenever $\kappa_{2}>9.65$. In other words, requiring validity on $\Theta^{(2)}\backslash\Theta^{(1)}$ (allowing for weak factors and unknown number of factors) would lead to efficiency loss on $\Theta^{(1)}$ if $\kappa_{2}>9.65$.
Although the constant of 9.65 is not the optimal constant, the analysis highlights the lack of adaptivity in inference. Without strong factor in the fixed effects, strong factor components in $X$ would result in a stark loss of efficiency. This can be explained. When the fixed effects $M$ have strong factors, the projections $\Pi_{M}$ and $\Pi_{M'}$ can be estimated well and thus we can safely identify components in $X$ that cannot be absorbed by the fixed effects; as a result, the variations in $\Pi_{M}D\Pi_{M'}+u$ can be used to learn $\beta$. However, when $M$ does not have strong factors, it is quite difficult to learn projections $\Pi_{M}$ and $\Pi_{M'}$ and hence we cannot clearly tell which part of $X$ is left after removing components correlated with the fixed effects; consequently, we will not be sure that any part of $D$ can be used to learn $\beta$ and instead will only consider variations in $u$ simply to be on the safe side (ensure coverage probability in all cases), resulting in a confidence interval with length unrelated to $\kappa_{2}$ (representing factor strength in $D$).
Another implication is that when there are potential weak factors, uncertainty in the number of factors leads to efficiency loss. This is in contrast with results in moon2014linear, who established that when all the factors are strong, uncertainty in the number of factors does not cause efficiency loss.