The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
69,079 characters
How well can we learn large factor models without assuming strong factors?
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\title{How well can we learn large factor models without assuming strong
factors?\thanks{First version October 23, 2019.}}
\author{Yinchu Zhu\thanks{Email: [email removed]}\textit{}\\
\textit{University of Oregon}}
\maketitle
\begin{abstract}
In this paper, we consider the problem of learning models with a latent
factor structure. The focus is to find what is possible and what is
impossible if the usual strong factor condition is not imposed. We
study the minimax rate and adaptivity issues in two problems: pure
factor models and panel regression with interactive fixed effects.
For pure factor models, if the number of factors is known, we develop
adaptive estimation and inference procedures that attain the minimax
rate. However, when the number of factors is not specified a priori,
we show that there is a tradeoff between validity and efficiency:
any confidence interval that has uniform validity for arbitrary factor
strength has to be conservative; in particular its width is bounded
away from zero even when the factors are strong. Conversely, any data-driven
confidence interval that does not require as an input the exact number
of factors (including weak ones) and has shrinking width under strong
factors does not have uniform coverage and the worst-case coverage
probability is at most 1/2. For panel regressions with interactive
fixed effects, the tradeoff is much better. We find that the minimax
rate for learning the regression coefficient does not depend on the
factor strength and propose a simple estimator that achieves this
rate. However, when weak factors are allowed, uncertainty in the number
of factors can cause a great loss of efficiency although the rate
is not affected. In most cases, we find that the strong factor condition
(and/or exact knowledge of number of factors) improves efficiency,
but this condition needs to be imposed by faith and cannot be verified
in data for inference purposes.
\end{abstract}
\section{Introduction}
Models with hidden factors have been a popular tool for analyzing
economic data. These models provide a convenient framework in describing
datasets that are big in both cross-sectional and time series dimensions.
The basic formulation is
\begin{equation}
X_{i,t}=M_{i,t}+u_{i,t},\label{eq: factor model}
\end{equation}
where $M\in\mathbb{R}^{n\times T}$ is a low-rank matrix and $u_{i,t}\sim N(0,1)$
is i.i.d across $(i,t)$. The low rank property of $M$ is another
way of describing the factor structure: if $M=LF'$ with $L\in\mathbb{R}^{n\times k}$
and $F\in\mathbb{R}^{T\times k}$, then ${\rm rank}\, M\leq k$. For this reason,
we view factor models as a low-rank matrix plus an error matrix. Throughout
the paper, all the idiosyncratic terms are assumed to be i.i.d and
have a normal distribution.
Most of the results in the literature on factor models assume the
strong factor condition or spiked eigenvalue condition, see \citep{Bai2003a,fan2011high,wang2017asymptotics}
among many others. This condition states that the largest $k$ singular
values of $M$ grow at the rate $\sqrt{nT}$ or at least faster than
$\sqrt{n+T}$. Besides the obvious question of whether such an assumption
can be checked in the data, perhaps a more relevant question is whether
we can assess the factor strength well enough for the purpose of estimation
and inference. It is important to know how well we can possibly learn
the factor models when the factor strength is also learned from the
data. In this paper, we can answer this question in pure factor models
and panel regression with interactive fixed effects. We study the
problem of learning scalar (low-dimensional) components, e.g., entries
of large matrices or regression coefficients.
The conceptual tools we use are minimax rates and adaptivity. We mainly
examine two issues.
\begin{itemize}
\item Minimax rates. What is minimax expected length of confidence intervals
for different levels of factor strengths? In particular, whether robust
confidence intervals (i.e., those with validity for arbitrary factor
strengths) and non-robust confidence intervals (i.e., those with validity
only under strong factors) have different rates overall. It turns
out that the answer depends on the problem (pure factor models or
panel regressions with interactive fixed effects).
\item Adaptivity. In terms of efficiency, can robust confidence intervals
match non-robust confidence intervals when the factors are actually
strong? To study this problem, we consider the minimax optimal performance
of robust confidence intervals over a parameter space in which factors
are strong. In general, we find that robust confidence intervals have
much worse efficiency than non-robust confidence intervals when all
the factors are strong. Therefore, requiring robustness (i.e., validity
under arbitrary factor strengths) leads to efficiency loss even in
situations with only strong factors. This robustness-efficiency tradeoff
implies that in order to improve efficiency, the strong factor condition
(and/or exact knowledge of number of factors) needs to be imposed
by faith and is impossible to verify in data for inference purpose;
if we have to learn the factor strengths from the data, there is necessarily
efficiency loss.
\end{itemize}
In a pure factor model, the task is to learn entries in $M$, say
$M_{1,1}$, where $X_{1,1}$ is potentially missing. We characterize
the impact of factor strength on estimation and inference. The main
findings are summarized as follows.
When the number of factors is known, the minimax rate for estimation
of $M_{1,1}$ depends on the factor strength (singular values of $M$).
If the singular values of $M$ does not grow faster than $\sqrt{n+T}$,
then it is impossible to achieve consistency in estimating $M_{1,1}$.
Moreover, the rate of convergence for $M_{1,1}$ can be learned from
the data: one can construct confidence intervals that are valid for
arbitrary factor strength and automatically achieve the optimal rate
in their widths.
When the number of factor is not given a priori, there is a tradeoff
between validity and efficiency. Any confidence interval that is valid
for arbitrary factor strength cannot have widths that shrink to zero
even when applied to strong factor settings. Conversely, if a confidence
interval learns the number of factors from the data and has widths
shrinking to zero when the all the factors are strong, then this confidence
interval cannot have uniform coverage for all factor strengths and
the worst-case coverage probability is below $1/2$. This tradeoff
only applies to inference. Achieving the minimax rate via a data-driven
estimator is entirely possible. The key insight is that although ignoring
weak factors does not reduce the rate in estimation, it is quite damaging
for inference valdity.
In a panel regression with interactive fixed effects, the situation
is much better. The fixed effects in these models have a factor structure
and the task is to learn the regression coefficient. We show that
even if the exact number of factors is unknown, it is possible to
learn the regression coefficient at the rate $(nT)^{-1/2}$, regardless
of the factor strengths in the fixed effects. However, when factors
are allowed to be weak, uncertainty in the number of factors can cause
a great loss of efficiency. This is in drastic contrast with existing
results that under the strong factor condition, it is not important
to know the exact number of factors, see \citet{moon2014linear}.
Our work is related to the literature on weak factors. Problems of
weak factors have been documented in simulations \citep[e.g., ][]{boivin2006more,bai2008large}
that study the performance of standard asymptotic theories. Theoretically,
the inconsistency of PCA (principal component analysis) has been pointed
out by \citet{johnstone2009consistency} and \citet{onatski2012asymptotics}
among others. \citet{onatski2012asymptotics} also derived detailed
asymptotic distribution of PCA when the factors are weak. Inconsistency
under weak factors is one of the main motivations driving developments
in the sparse PCA literature. There sparsity is imposed on factor
loadings to achieve consistent estimation of the factor structure,
see \citep{amini2008high,berthet2013optimal,cai2013sparse,birnbaum2013minimax}
among many others. It turns out that even when the factors are observed,
estimation and inference could encounter non-standard difficulties
if the factors are only weakly influential, \citep[e.g., ][]{kleibergen2009tests,gospodinov2017spurious,anatolyev1807.04094}.
The situation is more complicated if there are potentially missing
factors. We will now proceed to analysis of different models. Additional
discussion on related literature will be presented in Sections \ref{sec: SC}
and \ref{sec: panel reg} regarding pure factor models (with missing
values) and panel regressions, respectively.
\textbf{Notations}. For a matrix $A$, $\sigma_{1}(A)\geq\sigma_{2}(A)\geq\cdots$
denote the singular values of $A$ in decreasing order. For matrix
$A$, $\|A\|_{\infty}=\max_{i,j}|A_{i,j}|$, $\|A\|=\sigma_{1}(A)$
denotes the spectral norm, $\|A\|_{F}$ denotes the Frobenius norm
and $\|A\|_{*}$ denotes the nuclear norm (sum of all singular values
of $A$). For a vector $a$, $\|a\|_{2}$ denotes the Euclidean norm.
In model (\ref{eq: factor model}), the distribution of the data is
indexed by $M$; the probability measure and expectation under $M$
are denoted by $\mathbb{P}_{M}$ and $\mathbb{E}_{M}$, respectively. For two positive
sequences $X_{n},Y_{n}$, we use $X_{n}\lesssim Y_{n}$ to denote
$X_{n}\leq CY_{n}$ for a constant. We use $X_{n}\gtrsim Y_{n}$ to
denote $Y_{n}\lesssim X_{n}$ and $X_{n}\asymp Y_{n}$ means that
$X_{n}\lesssim Y_{n}$ and $X_{n}\gtrsim Y_{n}$. We use $\mathbf{1}\{\cdot\}$
for the indicator function. For a matrix $A\in\mathbb{R}^{n\times T}$, $A_{-1,-1}$
denotes the matrix $A$ with its (1,1) entry replaced by zero. For
a vector $a=(a_{1},...,a_{k})'$, we denote $a_{-1}=(a_{2},...,a_{k})'$.
\section{\label{sec: SC}Entry-wise learning: missing data and synthetic control}
In this section, we work with the following the parameter space:
\[
\mathcal{M}=\left\{ A\in\mathbb{R}^{n\times T}:\ \|A\|_{\infty}\leq\kappa\quad\text{and}\quad{\rm rank}\, A\leq2\right\} ,
\]
where $\kappa>0$ is a constant. The set $\mathcal{M}$ corresponds to factor
models with at most two factors. We assume that entries of the factor
structure are bounded so the factor strength would be a meaningful
quantity. This is related to the incoherence condition; see \citet{Bai1910.06677}
for an excellent discussion on the relation between incoherence and
factor strength. Here, we impose bounded $\|A\|_{\infty}$ simply
to rule out the situations in which the entire factor structure concentrates
on a few entries. For $M\in\mathbb{R}^{n\times T}$ with $M_{i,t}=nT\mathbf{1}\{(i,t)=(1,1)\}$,
the strong factor condition holds $\sigma_{1}(M)\asymp\sqrt{nT}$
but clearly the standard asymptotic theory for principal component
analysis (PCA) in \citet{Bai2003a} does not hold even if all entries
are observed; entries in $X_{-1,-1}$ obviously have no information
on $M_{1,1}$.
We define the set of one-factor models with factor strength $\tau$:
\[
\mathcal{M}(\tau)=\left\{ A\in\mathcal{M}:\ \sigma_{1}(A)\geq\tau\quad\text{and}\quad\sigma_{2}(A)=0\right\} .
\]
Our analysis of the case with known number of factors will deal with
$\mathcal{M}(\tau)$ and later discussion on unknown number of factors mainly
focuses on whether the number of factors is one or two. Restricting
the analysis to at most two factors is for simplicity and the results
can be extended to any fixed number of factors with extra (but perhaps
unnecessary) complications in notations and proofs. In this section,
we assume that the idiosyncratic terms are i.i.d standard normal variables
for simplicity.\footnote{The assumption of unit variance is not too restrictive since the quantity
$(nT)^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}u_{i,t}^{2}$ can always be
consistently estimated at the rate $\max\{n^{-1},T^{-1}\}$, regardless
of the factor strength. To see this, consider the PCA estimator $\hat{M}$
with rank $\bar{k}$, assuming $\bar{k}\geq k$. Since $\|X-\hat{M}\|\leq\|u\|$,
we have $\|\hat{M}-M\|\leq2\|u\|$. Since ${\rm rank}\,(\hat{M}-M)\leq2\bar{k}$,
we have $\|\hat{M}-M\|_{F}\leq2\sqrt{2\bar{k}}\|u\|=O_{P}(\sqrt{n+T})$
due to results from random matrix theory.}
One example that involves missing data is related to the synthetic
control problem. This is a fast growing literature since the seminal
papers by \citep{abadie2003economic,abadie2015comparative} and has
attracted much attention in applied research. One of the leading models
for synthetic control is to use a factor model for the counterfactuals,
see \citep{abadie2015comparative,gobillon2016regional,li2017statistical,athey2017matrix,xu2017generalized,chernozhukov1712.09089,li2018inference}
among many others. We consider a simple formulation.
\begin{example}[Synthetic control]
\label{exa: SC}There are $n$ units observed over $T$ time periods.
The potential outcome without treatment is denoted by $Y_{i,t}^{N}$
for unit $i$ in time period $t$. We assume that $Y_{i,t}^{N}=L_{i}F_{t}+u_{i,t}$
and the treatment is allocated to the first unit in the last time
period. In other words, the observed data is $Y_{i,t}=Y_{i,t}^{N}+D_{i,t}\gamma$
for $1\leq i\leq n$ and $1\leq t\leq T$, where $D_{i,t}=\mathbf{1}\{(i,t)=(1,T)\}$.
The object of interest is the treatment effect $\gamma$.
\end{example}
Like other causal inference problems, learning $\gamma$ is primarily
the task of learning the counterfactual $Y_{1,T}^{N}$, which is unobserved.
If $Y_{1,T}^{N}$ were known, then the obvious 95\% confidence interval
for $\gamma$ is $[Y_{1,T}-Y_{1,T}^{N}-1.96,Y_{1,T}-Y_{1,T}^{N}+1.96]$.
Since $Y_{1,T}^{N}$ is unobserved, we can replace it with a consistent
estimate. However, there is no guarantee that such an estimate exists
when the factors are not strong. If consistent estimation for $Y_{1,T}^{N}$
is impossible, then a 95\% confidence interval for $\gamma$ needs
to have a width larger than $1.96\times2$. To formally investigate
the estimation and inference for $Y_{1,T}^{N}$, we consider the following
problem.
\begin{example}[Missing one entry]
\label{exa: missing one entry}Suppose that the model in (\ref{eq: factor model})
holds. We observe all the entries of $X\in\mathbb{R}^{n\times T}$ except
one entry. Without loss of generality, we assume that $X_{1,1}$ is
missing and thus the observed data is $X_{-1,-1}$. The goal is to
estimate $M_{1,1}$ and conduct inference.
\end{example}
The main feature of Example \ref{exa: missing one entry} is that
the missing pattern is non-random. Since there is only one entry missing,
we essentially observe the entire data. For this reason, the problem
is very closely related to the following one.
\begin{example}[Full observation]
\label{exa: entrywise estimation} We observe $X\in\mathbb{R}^{n\times T}$
from the model in (\ref{eq: factor model}). The goal is to estimate
$M_{1,1}$ and conduct inference.
\end{example}
Example \ref{exa: entrywise estimation} corresponds to the problem
of estimation and inference in standard large factor models. Many
classical results are developed by \citet{Bai2003a} and the references
therein. In fact, the problem missing data in factor models also has
a long history dating back to at least \citet{stock2002macroeconomic}.
Very recently, works by \citet{su2019factor,Bai1910.06677,xiong1910.08273}
provided extensive asymptotic theories for various missing data problems
under the strong factor condition. However, when the factors are not
assumed to be strong, it is still unclear what can be done and what
is impossible.
\begin{example}[Missing at random]
\label{exa: random missing} We observe $\tilde{X},\Xi\in\mathbb{R}^{n\times T}$
with $\tilde{X}_{i,t}=X_{i,t}\Xi_{i,t}$, where $X_{i,t}$ is from the model
in (\ref{eq: factor model}) and $\Xi_{i,t}\in\{0,1\}$ is i.i.d Bernoulli
with $\mathbb{P}(\Xi_{i,t}=1)=\pi$. We assume that $X$ and $\Xi$ are independent.
The goal is to estimate $M_{1,1}$ and conduct inference.
\end{example}
Example \ref{exa: random missing} is also called the matrix completion
problem. This literature aims to estimate $M$ from $(\tilde{X},\Xi)$ and
most of the results are stated in terms of the overall risk in Frobenius
norm: construct an estimate $\hat{M}$ and provide bound on $\|\hat{M}-M\|_{F}$,
see \citep{candes2009exact,candes2010power,recht2010guaranteed,keshavan2010matrix,candes2011tight,rohde2011estimation,koltchinskii2011nuclear}
among others.
These results might seem quite encouraging in that they derive the
bounds without the strong factor condition. In particular, one can
construct an estimate $\hat{M}$ and guarantee
\[
(nT)^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}(\hat{M}_{i,t}-M_{i,t})^{2}=o_{P}(1)
\]
without any assumption on the factor strength. Since the average
(across entries) squared errors tends to zero, one might expect the
estimation error for a typical entry to be small without the strong
factor condition, e.g., $\hat{M}_{1,1}-M_{1,1}=o_{P}(1)$. We now show
that this is not true.
\subsection{\label{subsec: rate SC}Optimal rates: known number of factors}
We focus on the analysis of Example \ref{exa: missing one entry}.
When the number of factors is known to be one, we consider the problem
of how the strength of this factor affects the efficiency of estimation
and inference for $M_{1,1}$. The basic problem of our analysis is
to distinguish between
\[
\mathcal{M}^{{\rm null}}(\tau):=\left\{ M\in\mathcal{M}(\tau):\ M_{1,1}=0\right\}
\]
and
\[
\mathcal{M}_{\rho}(\tau):=\left\{ M\in\mathcal{M}(\tau):\ |M_{1,1}|\geq\rho\right\}
\]
for $\rho>0$. When the factors are strong, i.e., $\tau\gtrsim\sqrt{nT}$,
results in \citet{Bai1910.06677} suggest that non-trivial power can
be achieved when $\rho\gtrsim\min\{n^{-1/2},T^{-1/2}\}$. Our goal
in this subsection is to answer the following questions:
\begin{itemize}
\item for a given $\tau$, how large does $\rho$ need to be to guarantee
non-trivial power in testing $\mathcal{M}^{{\rm null}}(\tau)$ versus $\mathcal{M}_{\rho}(\tau)$?
\item if $\tau$ is not given, do we need a larger $\rho$ to distinguish
between $\mathcal{M}^{{\rm null}}(\tau)$ and $\mathcal{M}_{\rho}(\tau)$?
\end{itemize}
In this section, we fix $\alpha\in(0,1)$. We first give an impossibility
result concerning the first question.
\begin{thm}
\label{thm: lower bound 1 factor SC}Assume $\tau\leq\kappa\sqrt{nT}/12$.
Let $\psi$ be a measurable function taking values in $[0,1]$ such
that
\begin{equation}
\sup_{M\in\mathcal{M}^{{\rm null}}(\tau)}\mathbb{E}_{M}\psi(X_{-1,-1})\leq\alpha.\label{eq: size SC lower bnd}
\end{equation}
Then
\begin{equation}
\inf_{M\in\mathcal{M}_{\rho}(\tau)}\mathbb{E}_{M}\psi(X_{-1,-1})\leq2\alpha,\label{eq: power SC lower bnd}
\end{equation}
where $\rho=C\min\{1,\sqrt{n+T}/\tau\}$and $C>0$ is a constant that
depends only on $\kappa$ and $\alpha$.
\end{thm}
In Theorem \ref{thm: lower bound 1 factor SC}, $\psi$ is a test
for $H_{0}:\ M\in\mathcal{M}^{{\rm null}}(\tau)$ versus $H_{1}:\ M\in\mathcal{M}_{\rho}(\tau)$.
A test is a measurable function that maps the data $X_{-1,-1}$ to
$\{0,1\}$, where 1 denotes rejection of $H_{0}$. We allow for random
tests by extending $\{0,1\}$ to the interval $[0,1]$. The requirement
in (\ref{eq: size SC lower bnd}) ensures size control uniformly in
the null space $\mathcal{M}^{{\rm null}}(\tau)$. The conclusion of Theorem \ref{thm: lower bound 1 factor SC}
states that when $\rho\lesssim\min\{1,\sqrt{n+T}/\tau\}$, no test
can guarantee non-trivial power in distinguishing $\mathcal{M}^{{\rm null}}(\tau)$
and $\mathcal{M}_{\rho}(\tau)$.
As a consequence, it is impossible to achieve consistency in estimating
$M_{1,1}$ if $\tau\lesssim\sqrt{n+T}$. To see this, notice that
in this case, there exists a constant $\rho_{0}>0$ such that no test
can have high power in testing $H_{0}:\ M_{1,1}=0$ against $H_{1}:\ |M_{1,1}|\geq\rho_{0}$.
Hence, if a consistent estimator $\hat{M}_{1,1}$ exists, then the simple
test of $\mathbf{1}\{|\hat{M}_{1,1}|\leq\rho_{0}/2\}$ would have both Type
I error and Type II error tending to zero for $H_{0}:\ M_{1,1}=0$
against $H_{1}:\ |M_{1,1}|\geq\rho_{0}$. However, since such a test
is impossible (due to Theorem \ref{thm: lower bound 1 factor SC}),
a consistent estimator for $M_{1,1}$ does not exist.
Another consequence of Theorem \ref{thm: lower bound 1 factor SC}
is that the factor strength has a direct impact on the rate of learning
$M_{1,1}$. In order to obtain the usual rate of $\min\{n^{-1/2},T^{-1/2}\}$,
the strong factor assumption of $\tau\asymp\sqrt{nT}$ is necessary.
If $\tau\ll\sqrt{nT}$, the rate for $M_{1,1}$ is strictly worse
than $\min\{n^{-1/2},T^{-1/2}\}$.
We now present an implication of Theorem \ref{thm: lower bound 1 factor SC}
on confidence intervals. Although this implication is immediate, the
notations we introduce will be very useful in discussing later issues,
especially with unknown number of factors. For a given parameter space
$\mathcal{M}^{(1)}\subseteq\mathcal{M}$, we define the set of valid $1-\alpha$
confidence intervals as measurable functions that map the data $X_{-1,-1}$
to an interval such that this interval covers $M_{1,1}$ with probability
at least $1-\alpha$ for all parameters in $\mathcal{M}^{(1)}$. In other
words, we define
\[
\Phi(\mathcal{M}^{(1)})=\left\{ CI(\cdot)=[l(\cdot),u(\cdot)]:\ \inf_{M\in\mathcal{M}^{(1)}}\mathbb{P}_{M}\left(M_{1,1}\in CI(X_{-1,-1})\right)\geq1-\alpha\right\} .
\]
The minimax expected length of confidence intervals over $\mathcal{M}^{(1)}$
is
\[
\mathcal{L}(\mathcal{M}^{(1)})=\inf_{CI\in\Phi(\mathcal{M}^{(1)})}\sup_{M\in\mathcal{M}^{(1)}}\mathbb{E}_{M}|CI(X_{-1,-1})|,
\]
where $|CI|$ denotes the length of the confidence interval $CI$,
i.e., $|CI(\cdot)|=u(\cdot)-l(\cdot)$ for $CI(\cdot)=[l(\cdot),u(\cdot)]$.
Minimax rate for the length of confidence intervals is a common way
of examining the efficiency for inference in high-dimensional models\footnote{Compared to stating results in terms of size and power, discussing
length of confidence intervals brings simplicity mainly because we
do not need a notation on the hypothesized value in the null.}, \citep[see e.g., ][]{cai2017confidence,bradic1802.09117}. Theorem
\ref{thm: lower bound 1 factor SC} has the following simple implication;
we omit the proof for brevity.
\begin{cor}
\label{cor: lower bound CI 1-factor SC}Assume $\tau\leq\kappa\sqrt{nT}/12$.
Then
\[
\mathcal{L}(\mathcal{M}(\tau))\gtrsim\min\{1,\sqrt{n+T}/\tau\}.
\]
\end{cor}
Corollary \ref{cor: lower bound CI 1-factor SC} states a lower bound
for $\mathcal{L}(\mathcal{M}(\tau))$. We now show that this lower bound is also
optimal. Since consistency is impossible when $\tau_{n}\lesssim\sqrt{n+T}$,
we focus on $\tau_{n}\gg\sqrt{n+T}$. (If $\tau\lesssim\sqrt{n+T}$,
we can simply use the simple confidence interval $[-\kappa,\kappa]$
assuming $\kappa$ is known.) We first construct an estimator that
has the rate of convergence $\sqrt{n+T}/\tau$ when $\tau_{n}\gtrsim\sqrt{n+T}$.
The idea is based on the following the observation:
\[
M_{1,1}=\frac{L_{1}L_{-1}'M_{-1,1}}{\|L_{-1}\|_{2}^{2}},
\]
where $M_{-1,1}=(M_{2,1},...,M_{n,1})'\in\mathbb{R}^{n-1}$. Hence, a plug-in
estimator would be
\begin{equation}
\hat{M}_{1,1}=\frac{\hat{L}_{1}\hat{L}_{-1}'X_{-1,1}}{\|\hat{L}_{-1}\|_{2}^{2}},\label{eq: PCA estimate SC}
\end{equation}
where $X_{-1,1}=(X_{2,1},...,X_{n,1})'\in\mathbb{R}^{n-1}$. Here, $\hat{L}=(\hat{L}_{1},\hat{L}_{-1}')'\in\mathbb{R}^{n}$
is the PCA estimator for $L$ using data $X_{,-1}=\{X_{i,t}\}_{1\leq i\leq n,\ 2\leq t\leq T}$.
This estimator can be viewed as a version of the tall-wide estimator
by \citet{Bai1910.06677}. The proof in our case is significantly
complicated by the fact that the factor strengths might not be strong.
This leads to difficulties in bounding certain terms; we use a novel
construction that allows us to apply a decoupling argument, see Lemma
\ref{lem: bound key 1} in the appendix. The following result establishes
its rate of convergence under any given factor strength.
\begin{thm}
\label{thm: upper bound 1 factor SC}Consider $\hat{M}_{1,1}$ in (\ref{eq: PCA estimate SC}).
Assume that $\tau\geq4\max\{\sqrt{10}\kappa,2\}\sqrt{3(n+T)}$. Then
for any $\alpha\in(0,1)$, there exists a constant $C>0$ such that
\[
\sup_{M\in\mathcal{M}(\tau)}\mathbb{P}_{M}\left(|\hat{M}_{1,1}-M_{1,1}|>C\sqrt{n+T}/\tau\right)\leq\alpha.
\]
\end{thm}
Fortunately, distinguishing between $\tau\gg\sqrt{n+T}$ and $\tau\lesssim\sqrt{n+T}$
is not difficult, thanks to bounds in random matrix theory. Since
$\|M_{-1,-1}\|_{F}^{2}=\|M\|_{F}^{2}-M_{1,1}^{2}\geq\tau^{2}-\kappa^{2}$,
$\|M_{-1,-1}\|_{F}$ and $\tau$ have the same rate. Since $M_{-1,-1}$
has rank at most 2, $\|M_{-1,-1}\|_{F}\geq\|M_{-1,-1}\|\geq\|M_{-1,-1}\|_{F}/\sqrt{2}$.
Therefore, $\|X_{-1,-1}\|\gtrsim\|M_{-1,-1}\|-\|u_{-1,-1}\|=O_{P}(|\tau-\sqrt{n+T}|)$
by Corollary 5.35 of \citet{vershynin2010introduction}. Hence, the
following estimator always has the optimal rate of convergence:
\begin{equation}
\hat{M}_{1,1}\mathbf{1}\left\{ \|X_{-1,-1}\|>4\max\{\sqrt{10}\bar{\kappa},2\}\sqrt{3(n+T)}\right\} ,\label{eq: estimator adaptivity}
\end{equation}
where $\bar{\kappa}>0$ is a constant that upper bounds $\kappa$.
The idea is that if $\|X_{-1,-1}\|$ is smaller than the above threshold,
we can safely conclude $\tau\lesssim\sqrt{n+T}$ and thus no estimator
is consistent for $M_{1,1}$, which means that the estimator zero
would have the optimal ``rate'' too (since $M_{1,1}$ is bounded).
Therefore, we have the following result.
\begin{cor}
$\mathcal{L}(\mathcal{M}(\tau))\asymp\min\{1,\sqrt{n+T}/\tau\}.$
\end{cor}
Notice that the estimator in (\ref{eq: estimator adaptivity}) is
adaptive in that it automatically achieves the optimal rate $\min\{1,\sqrt{n+T}/\tau\}$
without prior knowledge of $\tau$. Since we can learn the rate for
$\tau$ from $\|X_{-1,-1}\|$, we can estimate the rate $\min\{1,\sqrt{n+T}/\tau\}$
from the data and thereby construct a confidence interval centered
around the estimator in (\ref{eq: estimator adaptivity}): simply
choose a width of $C_{0}\min\{\sqrt{n+T}/\|X_{-1,-1}\|,1\}$ for a
universal constant $C_{0}$ (which can be explicitly determined).
Therefore, the length of the confidence interval is also adaptive
since it automatically adjusts to have the optimal rate. Unfortunately,
this turns out to be true only when the number of factors is known.
We will now see that when we do not know the exact number of factors
(including weak ones), such adaptivity is impossible.
\subsection{Adaptivity: unknown number of factors}
The previous analysis already establishes the optimal rate under knowledge
of the exact number of factors. Now we aim to answer the following
questions:
\begin{itemize}
\item Is it possible to achieve the same rate as established before if the
number of factors is not known?
\item Do we pay a price if we cannot determine the number of weak factors?
\end{itemize}
These questions arise naturally in practice as the researcher is typically
not given the exact number of factors. In these cases, one often needs
to estimate the number of factors from the data and then proceeds
using this estimate.
\begin{comment}
\footnote{Some recent methods allow estimation of factor models via nuclear-norm
penalization and thus avoid explicit knowledge of the number of factors.
However, as we shall see, this does not mean that lack of knowledge
on the number of factors has no cost.}
\end{comment}
Although such a strategy makes intuitive sense, it is still far from
clear whether lack of knowledge on the number of factors has any cost.
This problem is particularly tricky for inference since uniform size
control (or coverage) is a concern; we do not have a notion of uniform
validity for estimation. Suppose that we know there are either one
or two factors. Ideally, we would like to construct a confidence interval
that has $1-\alpha$ coverage probability uniformly over all one-factor
models and all two-factor models, regardless of factor strengths.
Procedures that involve a pre-estimation of the number of factors
are intended to have such robustness. If there is one factor, then
hopefully the procedure will find out that there is one factor and
constructs the confidence interval accordingly; if there are two factors,
then the procedure is supposed to detect that and constructs a valid
confidence interval based on that finding. This leads to a confidence
interval summarized in Figure \ref{fig: PCA estimated num factors}.
Even for methods, such as nuclear-norm-regularized approaches, which
do not need the number of factors as an input, we still hope that
there is a data-driven procedure that would provide valid inference
no matter what the true number of factors is. In light of this robustness
requirement, we are interested in finding out whether these uniformly
valid procedures lose any efficiency, compared to the situation with
known number of factors.
\begin{figure}
\caption{\label{fig: PCA estimated num factors}Pre-tests + PCA}
{\small{}\medskip{}
}{\small\par}
{\small{}\begin{tikzpicture}[sibling distance=25em, every node/.style = {shape=rectangle, rounded corners, draw, align=center, top color=white!10, bottom color=blue!20}]] \node {A priori knowledge: the number of factors $k\in\{1,2\} $.} child { node {Pre-test: produce $\hat{k}\in\{1,2\} $ from the data. } child { node {$\hat{k}=1$} child { node {PCA estimate: keep the first PC} child { node {Compute standard error: \\asymptotic normality with rate $n^{-1/2}+T^{-1/2} $ } } } } child { node {$\hat{k}=2$} child { node {PCA estimate: keep the first two PC's} child { node {Compute standard error: \\asymptotic normality with rate $n^{-1/2}+T^{-1/2} $ } } } }}; \end{tikzpicture}}{\small\par}
{\small{}The rate $n^{-1/2}+T^{-1/2}$ and the asymptotic normality
under strong factors are established in \citet{Bai2003a}. Here, the
term PC refers to principal component. }{\small\par}
\end{figure}
We would not expect any loss of efficiency if the number of factors
can be consistently estimated. Unfortunately, determining the number
of weak factors is quite difficult (if not impossible) since typical
methods \citep[e.g.,][]{bai2002determining,onatski2009testing,ahn2013eigenvalue}
only promise success in detecting strong factors. If we run these
pre-tests and ignore potential weak factors, do we pay a price in
terms of validity or efficiency?
To formally answer this question, we introduce the notion of inference
adaptivity. Consider the class of one-factor models with factor strength
$\tau_{0}$:
\[
\mathcal{M}^{(1)}=\mathcal{M}(\tau_{0}).
\]
From the previous analysis, we know that a confidence interval for
$M_{1,1}$ with validity over $\mathcal{M}^{(1)}$ can have shrinking length
only if $\tau_{0}\gg\sqrt{n+T}$. Since we do not know for sure that
there is only one factor, we would like to consider a confidence interval
that has validity over a larger space. For example, given $\tau_{1}\geq\tau_{2}>0$,
we define
\[
\mathcal{M}^{(2)}=\mathcal{M}(\tau_{1},\tau_{2}):=\left\{ A\in\mathcal{M}:\ \sigma_{1}(A)\geq\tau_{1},\ \sigma_{2}(A)\leq\tau_{2}\right\} ,
\]
where $\tau_{1}\leq\tau_{0}$. Clearly, $\mathcal{M}^{(1)}\subset\mathcal{M}^{(2)}$.
\begin{comment}
However, in practice, the exact number of factors is often unknown.
Hence, we would like to use a data-driven procedure that does not
require the exact number of factors as an input. In essence, we are
considering a procedure that maps the data to an interval (or onto
$\{0,1\}$ as a test) and is valid over a class that contains parameters
corresponding to different number of factors.
\end{comment}
The primary setting is $\tau_{1},\tau_{0}\gg\sqrt{n+T}$ and $\tau_{2}\lesssim\sqrt{n+T}$.
In this setting, $\mathcal{M}^{(2)}$ allows for a second weak factor. The
exercise is to consider all the confidence intervals that have uniform
coverage over $\mathcal{M}^{(2)}$ and then choose the one that has the
best efficiency over $\mathcal{M}^{(1)}$ (in terms of expected length).
The questions listed in the two bullet points above now become an
issue of adaptivity. If there is a confidence interval that is valid
over $\mathcal{M}^{(2)}$ and has length automatically achieves the rate
of $\sqrt{n+T}/\tau_{0}$ over $\mathcal{M}^{(1)}$, then (at least for
the rate) lack of knowledge on the number of factors does not have
a cost and in particular we do not pay a price in terms of size and
power for failing to detect weak factors. If any confidence interval
that has validity over $\mathcal{M}^{(2)}$ necessarily has a length greater
than $O(\sqrt{n+T}/\tau_{0})$ over $\mathcal{M}^{(1)}$, then allowing
for the extra validity on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$ causes
an efficiency loss. In particular, we are interested in the following
quantity
\[
\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})=\inf_{CI\in\Phi(\mathcal{M}^{(2)})}\sup_{M\in\mathcal{M}^{(1)}}\mathbb{E}_{M}|CI(Y_{-1,-1})|.
\]
The above quantity is the best guaranteed expected length over $\mathcal{M}^{(1)}$
among all the confidence intervals with uniform validity on $\mathcal{M}^{(2)}$.
The minimax rate $\mathcal{L}(\mathcal{M}^{(1)})$ defined before is a special
case since $\mathcal{L}(\mathcal{M}^{(1)})=\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(1)})$. Notice
that by construction, we have $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\geq\mathcal{L}(\mathcal{M}^{(1)})$.
If $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\gg\mathcal{L}(\mathcal{M}^{(1)})$, then the
additional validity requirement on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$
has a negative impact on the efficiency on $\mathcal{M}^{(1)}$. If $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\asymp\mathcal{L}(\mathcal{M}^{(1)})$,
then we say that it is possible to achieve adaptive inference; the
procedure can automatically recognize that the parameter is in the
``fast-rate-region'' $\mathcal{M}^{(1)}$ and constructs the confidence
interval accordingly.
\begin{rem}
\label{rem: linear IV}Examples of adaptive inference can be found
in other econometric models. Consider the linear IV model with one
endogenous regressor and one instrumental variable (IV). We know that
when the IV is strong, the confidence interval for the regression
coefficient shrinks at the parametric rate $n^{-1/2}$; when the IV
is weak, the confidence interval might not shrink to zero. One can
invert an identification-robust test\footnote{For example, the Anderson-Rubin test, Lagrange multiplier test by
\citet{kleibergen2005testing} or conditional likelihood ratio test
by \citet{moreira2003conditional}.} to obtain a confidence interval. Clearly, such an interval is valid
for any identification strength. On the other hand, this interval
will shrink at the rate $n^{-1/2}$ when the IV is strong. In this
example, $\mathcal{M}^{(1)}$ corresponds to the parameter space with strong
IV and $\mathcal{M}^{(2)}$ represents all parameter values (i.e., the identification
may or may not be strong).
\end{rem}
We now show that such adaptivity is unfortunately impossible for factor
models when the number of factors is unknown.
\begin{comment}
We interpret this quantity as follows. Since $\mathcal{M}^{(1)}\subseteq\mathcal{M}^{(2)}$,
$\Phi(\mathcal{M}^{(2)})$ contains more ``robust'' confidence intervals
than $\Phi(\mathcal{M}^{(1)})$ because elements of $\Phi(\mathcal{M}^{(2)})$
have validity over the larger space $\mathcal{M}^{(2)}$. Then we consider
the performance of these robust confidence intervals in a smaller
parameter space. Since the minimax rate over a smaller parameter space
is faster, we would like to know whether procedures that are valid
over a larger space can achieve efficiency on a smaller space.
We now use this framework to study the adaptivity with respect to
factor strength.
Of course, we certainly do not wish to underestimate the number of
strong factors
Another way of viewing this proble
we are not given the number of factors. It is important to check whether
lack of such knowledge can greatly reduce the efficiency of inference.
In particular, it is quite difficult to distinguish the following
two settings from the data:
\[
\text{one strong factor versus two factors (one strong and one weak)}.
\]
The key question is what is the rate $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})$
when $\tau_{1},\tau_{0}\gg\sqrt{n+T}$ and $\tau_{2}\lesssim\sqrt{n+T}$.
In this case, $\mathcal{M}^{(1)}$ corresponds to one factor models with
strong factors, whereas $\mathcal{M}^{(2)}$ includes $\mathcal{M}^{(1)}$ and
two-factor models with a weak second factor.
\end{comment}
\begin{thm}
\label{thm: adaptivity SC 1}Assume that $\tau_{1},\tau_{0}\leq\kappa\sqrt{nT}/12$
and $\tau_{2}\geq\kappa_{1}$. There exists a constant $C>0$ such
that
\[
\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\geq C.
\]
\end{thm}
In this paper, confidence intervals with validity over $\mathcal{M}^{(2)}$
will be referred to as robust confidence intervals (or weak-factor-robust
confidence intervals). Theorem \ref{thm: adaptivity SC 1} has a striking
implication when $\tau_{1},\tau_{0}\asymp\sqrt{nT}$ and $\tau_{2}$
is bounded away from zero. In this case, no robust confidence intervals
can guarantee to have shrinking width when there is actually only
one factor and this factor is strong; recall that in this case, $\mathcal{L}(\mathcal{M}^{(1)})\asymp\sqrt{n+T}/\tau_{0}\asymp\min\{n^{-1/2},T^{-1/2}\}$.
Hence, $\mathcal{L}(\mathcal{M}^{(1)},\mathcal{M}^{(2)})\gg\mathcal{L}(\mathcal{M}^{(1)})$. In
other words, robustness necessarily causes efficiency loss. Unlike
the adaptivity in Remark \ref{rem: linear IV} for linear IV models,
no confidence intervals valid over $\mathcal{M}^{(2)}$ can adapt its efficiency
on $\mathcal{M}^{(1)}$. When $\tau_{0},\tau_{1},\tau_{2}\asymp\sqrt{nT}$,
robust confidence intervals will have lengths tending to zero only
if there are two factors and both are strong, but in this case we
are back to the situation with known number of factors.
The other side of this coin is the observation that any confidence
interval that has shrinking width and ignores weak factor necessarily
lacks uniform coverage. We state an explicit result below.
\begin{thm}
\label{thm: bad coverage}Let $\mathcal{I}_{*}(X_{-1,-1})=[l_{*}(X_{-1,-1}),u_{*}(X_{-1,-1})]$
be a random interval. Suppose that for some sequences $\tau_{0}\gtrsim1$,
we have $\sup_{M\in\mathcal{M}^{(1)}}\mathbb{E}_{M}\left|\mathcal{I}_{*}(X_{-1,-1})\right|=o(1)$.
If $\tau_{2}\gtrsim1$, then
\[
\limsup_{n,T\rightarrow\infty}\inf_{M\in\mathcal{M}^{(2)}}\mathbb{P}_{M}\left(M_{1,1}\in\mathcal{I}_{*}(X_{-1,-1})\right)\leq\frac{1}{2}.
\]
\end{thm}
Theorem \ref{thm: bad coverage} reveals the danger of some popular
methods in practice. Consider the case with $\tau_{0},\tau_{1},\tau_{2}\asymp\sqrt{nT}$
(so $\tau_{0},\tau_{2}\gtrsim1$ is satisfied). Recall the procedure
in Figure \ref{fig: PCA estimated num factors} under the simplified
assumption that the number of factors is either one or two. By classical
results, the PCA estimate is asymptotically normal over $\mathcal{M}^{(1)}$
with a standard error shrinking to zero. Hence, the overall procedure
in Figure \ref{fig: PCA estimated num factors}, which uses these
standard errors, produces a random interval with shrinking length
on $\mathcal{M}^{(1)}$. However, Theorem \ref{thm: bad coverage} indicates
that precisely due to its shrinking width on $\mathcal{M}^{(1)}$, the procedure
in Figure \ref{fig: PCA estimated num factors} cannot have uniform
coverage over $\mathcal{M}^{(2)}$ unless the number of factors (including
weak ones) can be consistently determined. Since this result holds
regardless of which pre-test is used, developing better tests for
the number of factors might not significantly improve inference quality.
Moreover, using methods that do not require the number of factors
as an input (e.g., nuclear-norm penalized methods) does not solve
the problem either since Theorem \ref{thm: bad coverage} does not
assume a specific form for $\mathcal{I}_{*}$. In this regard, Theorem \ref{thm: bad coverage}
also serves as a simple check on the uniform validity of a given procedure:
if there exists two different numbers $k_{1}$ and $k_{2}$ such that
the procedure under consideration gives shrinking confidence intervals
in both cases ((1) when there are $k_{1}$ strong factors and no weak
factors and (2) when there are $k_{2}$ strong factors and no weak
factors), then this procedure does not have uniform validity.
\begin{comment}
One such method is to use a pre-test to determine the number of factors
and then use the estimated number of factors for PCA estimation and
construct confidence intervals using the classical asymptotic normality
results. We describe this method in
\end{comment}
Theorems \ref{thm: adaptivity SC 1} and \ref{thm: bad coverage}
indicate an inherent tradeoff between efficiency and validity. On
one hand, Theorem \ref{thm: adaptivity SC 1} states that procedures
with validity over $\mathcal{M}^{(2)}$ cannot provide accurate inference
on $\mathcal{M}^{(1)}$. On the other hand, Theorem \ref{thm: bad coverage}
makes the equivalent claim that procedures that provide accurate inference
on $\mathcal{M}^{(1)}$ cannot guarantee uniform validity on $\mathcal{M}^{(2)}$.
Therefore, requiring validity on $\mathcal{M}^{(2)}\backslash\mathcal{M}^{(1)}$
necessarily reduces the inference efficiency.
\begin{comment}
If we require uniform validity over $\mathcal{M}^{(2)}$,
. Hence, our result shows that no matter which pre-test is used, this
procedure has no uniform validity as long as such ``corresponding
asymptotic theory'' gives a shrinking confidence interval. As a result,
For example, if a pre-test detects only one strong factor and we build
a confidence interval with shrinking length, then this overall procedure
(with the pre-test included) cannot provide uniform coverage over
$\mathcal{M}^{(2)}$. If it could provide uniform coverage, then we would
have found a robust confidence interval whose width shrinks to zero
over $\mathcal{M}^{(1)}$; however, this is impossible due to Theorem \ref{thm: adaptivity SC 1}.
This outlines the danger of some common practice in applied research.
Very often the researcher simply tries to detect the number of factors
via a pre-test, applies PCA and invoke the corresponding asymptotic
theory, which typically suggests a confidence interval of length $O(\min\{n^{-1/2},T^{-1/2}\})$.
\end{comment}
Since Theorem \ref{thm: adaptivity SC 1} only involves the worst-case
width (or power) on $\mathcal{M}^{(1)}$. One might wonder whether robust
procedures could still have decent efficiency on average over $\mathcal{M}^{(1)}$
despite the bad worst-case performance. We now show that this is not
the case: lack of efficiency occurs at many points in $\mathcal{M}^{(1)}$.
\begin{thm}
\label{thm: adaptivitity SC main}For any $\eta\in(0,1)$, define
\[
\mathcal{M}_{*}^{(1)}=\left\{ A\in\mathbb{R}^{n\times T}:\ \|A\|_{\infty}\leq\kappa(1-\eta),\ \sigma_{1}(A)\geq\tau_{0}(1+\eta),\ \sigma_{2}(A)=0\right\} .
\]
Then
\[
\inf_{CI\in\Phi(\mathcal{M}^{(2)})}\inf_{M\in\mathcal{M}_{*}^{(1)}}\mathbb{E}_{M}|CI(X_{-1,-1})|\geq(1-2\alpha)\min\{\kappa\eta,\tau_{0}\eta,\tau_{2}\}.
\]
\end{thm}
Since $\eta$ can be chosen arbitrarily, $\mathcal{M}_{*}^{(1)}$ represents
quite many (if not most) points in $\mathcal{M}^{(1)}$. Theorem \ref{thm: adaptivitity SC main}
states that the efficiency is bad at every point in $\mathcal{M}_{*}^{(1)}$
for any robust confidence interval. Therefore, the impossibility result
in Theorem \ref{thm: adaptivity SC 1} is not driven by only a few
unlucky points in $\mathcal{M}^{(1)}$. This means that there is fundamental
difficulty in entry-wise learning when the number of factors is unknown.
Hence, in order to achieve efficient inference, the number of factors
needs to be given a priori.
\begin{rem}[Estimation adaptivity vs inference adaptivity]
We note that the lack of adaptivity implied by Theorems \ref{thm: adaptivity SC 1},
\ref{thm: bad coverage} and \ref{thm: adaptivitity SC main} is only
about inference. It is entirely possible to have adaptive estimation
when the number of factors is unknown. For example, consider the procedure
in Figure \ref{fig: PCA estimated num factors}. This procedure (after
proper truncation if needed) still provides an estimate with the optimal
rate\footnote{When all the factors are strong (e.g., $\mathcal{M}^{(1)}$ with $\tau_{0}\asymp\sqrt{nT}$),
ignoring the weak factors leads to the best rate. When there are weak
factors, no consistency is possible anyway due to results in Section
\ref{subsec: rate SC}; hence, one can simply use zero as the estimate
via truncation (similar to (\ref{eq: estimator adaptivity})) in this
case. } and hence is an adaptive estimator. However, lack of adaptivity for
inference reflects the fact that we cannot learn from the data what
this optimal rate is.
\begin{comment}
We now summarize all the situations in the following table.\tikzset{every picture/.style={line width=0.75pt}}
\begin{tikzpicture}[x=0.75pt,y=0.75pt,yscale=-1,xscale=1]
\end{tikzpicture}
\begin{tabular}{ccc}
& & \tabularnewline
& {\footnotesize{}Allow weak factors} & {\footnotesize{}All factors are strong}\tabularnewline
{\footnotesize{}Known number of factors} & {\footnotesize{}Rate depends on factor strength with adaptivity.} & {\footnotesize{}Easy.}\tabularnewline
{\footnotesize{}Unknown number of factors} & {\footnotesize{}Lack of adaptivity. Conservative inference.} & {\footnotesize{}Easy to estimate the number of factors.}\tabularnewline
\end{tabular}
\begin{figure}
\caption{\label{fig:plot}Minimax rate and adaptivity}
\centering{}{\footnotesize{}}
\noindent\begin{minipage}[t]{1\columnwidth}
{\footnotesize{}\begin{center}
\tikzset{every picture/.style={line width=0.75pt}}
\begin{tikzpicture}[x=0.75pt,y=0.75pt,yscale=-1,xscale=1]
\draw (10,140) -- (600,140) ;
\draw (304,30) -- (305,250) ;
\draw (155,86) node [align=left] {\textbf{Condition I: }\\
$\bullet$ validity for
allow weak factors\\ $\bullet$ unknown number of factors: either one or two\\\\ Then\\ $\bullet$ No consistency even };
\draw (450,86) node [align=left] {\textbf{Condition II: }\\allow weak factors + known number of factors\\\\ $\bullet$ Rate $\min\{1,\sqrt{n+T}/\tau\}$ achieved adaptively};
\draw (450,200) node [align=left] {\textbf{Condition IV: }\\no weak factors + known number of factors\\\\ $\bullet$ Rate $\min\{1,\sqrt{n+T}/\tau\}$ achieved adaptively};
\draw (155,200) node [align=left] {\textbf{Condition III: }\\no weak factors + unknown number of factors\\\\ $\bullet$ Reduced to Condition IV \\(number of factors can be consistently estimated)};
\end{tikzpicture}
\end{center}}{\footnotesize\par}
{\footnotesize{}The number of factors is $k$. Weak factors refer
to $\tau\lesssim\sqrt{n+T}$.}{\footnotesize\par}
\end{minipage}{\footnotesize\par}
\end{figure}
\end{comment}
\end{rem}
\section{\label{sec: panel reg}Panel regression with interactive fixed effects}
\subsection{Can we achieve the optimal rate without assuming strong factors?}
The panel data model with interactive fixed effects \citep[e.g., ][]{pesaran2006cross,bai2009panel}
assumes
\[
Y_{i,t}=L_{i}'F_{t}+X_{i,t}\beta+\varepsilon_{i,t},
\]
where the observed data is $\{(Y_{i,t},X_{i,t})\}_{1\leq i\leq n,\ 1\leq t\leq T}$.
Here, $\varepsilon_{i,t}$ is i.i.d across $(i,t)$ following $N(0,\sigma_{\varepsilon}^{2})$
and the factor structure $L_{i}'F_{t}$ is the fixed effect with $L_{i},F_{t}\in\mathbb{R}^{r_{0}}$.
For simplicity, here $X_{i,t}$ and $\beta$ are scalars. The main
requirement of $X_{i,t}$ is that it cannot be absorbed by the fixed
effects. Since the fixed effects have a factor structure, we assume
that $X_{i,t}$ is a factor structure plus non-negligible noise:
\[
X_{i,t}=\alpha_{i}'g_{t}+u_{i,t},
\]
where $\alpha_{i},g_{t}\in\mathbb{R}^{r_{1}}$ and $u_{i,t}$ is i.i.d across
$(i,t)$ following $N(0,\sigma_{u}^{2})$. We assume that $u$ and
$\varepsilon$ are mutually independent. Factor structures in the
regressor $X_{i,t}$ have been a common assumption \citep[see e.g.,][]{pesaran2006cross,moon2014linear,zhu2017high,chernozhukvo2018panel}.
Under these assumptions, we can write the model as
\begin{equation}
Y=M+X\beta+\varepsilon\text{ and }X=D+u,\label{eq: panel regression}
\end{equation}
where $M,D,\varepsilon,u\in\mathbb{R}^{n\times T}$, ${\rm rank}\, M\leq r_{0}$
and ${\rm rank}\, D\leq r_{1}$. The distribution of the data $(Y,X)$ is
thus indexed by $\theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta)$.
We consider the following parameter space:
\[
\Theta=\left\{ \theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta):\ \ {\rm rank}\, M\leq r_{0},\ {\rm rank}\, D\leq r_{1},\ \{\sigma_{\varepsilon},\sigma_{u}\}\subset[\kappa^{-1},\kappa],\ |\beta|\leq\kappa\right\} ,
\]
where $\kappa>0$ is a constant and $r_{1},r_{2}>0$ are fixed. The
requirement of $\{\sigma_{\varepsilon},\sigma_{u}\}\subset[\kappa^{-1},\kappa]$
is merely saying that $\sigma_{\varepsilon}$ and $\sigma_{u}$ are
bounded away from zero and infinity. For simplicity, we shall also
require that $\beta$ be bounded. Notice that we do not require boundedness
of $\|M\|_{\infty}$ and $\|D\|_{\infty}$; it turns out that such
a requirement will not be needed. We note that $r_{0}$ and $r_{1}$
are only upper bounds on the number of factors, instead of the exact
number of factors. Moreover, there is no requirement on the relative
magnitude between $n$ and $T$.
The most important feature of $\Theta$ is that there is no assumption
at all regarding the strength of the factors. For estimating $\beta$,
\citet{bai2009panel} showed that when all the factors are strong
and the number of factors is known (or consistently estimable), one
can achieve the rate $(nT)^{-1/2}$. \citet{moon2014linear} showed
that when all the factors are strong, overstating the number of factors
does not have any impact on the rate for estimating $\beta$. These
results still leave two open questions that are quite relevant in
practice:
\begin{itemize}
\item If the number of factors is known and some factors might be weak,
would this create a problem for learning $\beta$?
\item Furthermore, if there are uncertainties regarding both the number
of factors and factor strengths, would there be additional difficulties
learning $\beta$?
\end{itemize}
We now provide an estimator that achieves the rate $(nT)^{-1/2}$
uniformly over $\Theta$. As a result, lack of knowledge on the number
of factors and potentially weak factors do not create any problem
for the rate on learning $\beta$. Notice that the rate $(nT)^{-1/2}$
is minimax optimal by a simple two-point argument.
The basic idea of our estimator is as follows. The assumption in (\ref{eq: panel regression})
implies a factor structure in $Y$:
\[
Y=(M+D\beta)+V,
\]
where ${\rm rank}\,(M+D\beta)\leq r_{0}+r_{1}$ and $V=u\beta+\varepsilon$.
Thus, we can identify $\beta$ as $\beta=\mathbb{E}{\rm trace\,}(V'u)/(nT\sigma_{u}^{2})$.
Since both $V$ and $u$ are idiosyncratic parts in $Y$ and $X$,
respectively, we can estimate them by the typical low-rank estimation
strategies.
Without loss of generality, we assume that $T\geq n$; if $T<n$,
then we flip our data from $(Y,X)$ to $(Y',X')$. For any matrix
$A\in\mathbb{R}^{n\times r}$, we define $\Pi_{A}=I_{n}-P_{A}$ and $P_{A}=A(A'A)^{\dagger}A'$,
where $^{\dagger}$ denoting the Moore-Penrose pseudo-inverse. Let
$\hat{\alpha}\in\arg\min_{A\in\mathbb{R}^{n\times r_{1}}}{\rm trace\,}(X'\Pi_{A}X)$
and $\hat{\Lambda}\in\arg\min_{A\in\mathbb{R}^{n\times k}}{\rm trace\,}(Y'\Pi_{A}Y)$
with $k=r_{0}+r_{1}$, Notice that $\hat{\alpha}$ and $\hat{\Lambda}$ are
simply the eigenvectors corresponding to the largest $r_{1}$ and
$k$ eigenvalues of $XX'$ and $YY'$, respectively. Then the estimator
is defined as
\begin{equation}
\hat{\beta}:=\hat{\beta}(X,Y)=\frac{n-r_{1}}{n-\hat{r}}\cdot\frac{{\rm trace\,}(Y'\Pi_{\hat{\Lambda}}\Pi_{\hat{\alpha}}X)}{{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)},\label{eq: panel estimator}
\end{equation}
where $\hat{r}=k+r_{1}-{\rm trace\,}(P_{\hat{\Lambda}}P_{\hat{\alpha}})$.
We make three comments on the estimator. First, due to the identification
condition of $\beta=\mathbb{E}{\rm trace\,}(V'u)/(nT\sigma_{u}^{2})$, we can view
the estimation problem as learning the expected conditional covariance
\citep[e.g.,][]{newey_fast_rate1801.09138,chernozhukov2018learning,chernozhukov1802.08667}
and thus construct a solution in a similar spirit. The means of $X$
and $Y$ are high-dimensional structures: $A$ and $M+D\beta$. Since
the trace operation defines an inner product, ${\rm trace\,}(\mathbb{E} V'u)$ is
a linear functional of the covariance $\mathbb{E}[Y-(M+D\beta)]'[X-A]$.
The simple plug-in approach is to construct an estimate for $M+D\beta$
and $A$ and replace $\mathbb{E}(\cdot)$ with $(nT)^{-1}{\rm trace\,}(\cdot)$.
Under PCA, the estimate for $Y-(M+D\beta)$ and $X-D$ are $\Pi_{\hat{\Lambda}}Y$
and $\Pi_{\hat{\alpha}}X$, respectively.
Second, the quantity $(n-r_{1})$ acts as bias correction for ${\rm trace\,}(X'\Pi_{\hat{\alpha}}X)$
in (\ref{eq: panel estimator}). Although the random matrix theory
can gives us ${\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}nT+O_{P}(n+T)$,
this bound is not enough as it only yields
\[
(nT)^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}+O_{P}(\min\{n^{-1},T^{-1}\}).
\]
To achieve the rate $(nT)^{-1/2}$, we would need an estimate for
$\sigma_{u}^{2}$ at the rate $O_{P}((nT)^{-1/2})$, which is faster
than $O_{P}(\min\{n^{-1},T^{-1}\})$ unless $n\asymp T$. We derive
a more accurate characterization by showing
\[
[(n-r_{1})T]^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)=\sigma_{u}^{2}+O_{P}((nT)^{-1/2}).
\]
Therefore, $[T(n-r_{1})]^{-1}{\rm trace\,}(X'\Pi_{\hat{\alpha}}X)$ has a strictly
smaller remainder term unless $T\asymp n$. Similarly, the quantity
$n-\hat{r}$ acts as bias correction for ${\rm trace\,}(Y'\Pi_{\hat{\Lambda}}\Pi_{\hat{\alpha}}X)$
since one can show ${\rm trace\,}(Y'\Pi_{\hat{\Lambda}}\Pi_{\hat{\alpha}}X)=(n-\hat{r})T\sigma_{u}^{2}\beta+O_{P}(\sqrt{nT})$.
Third, a key step in analyzing the numerator in (\ref{eq: panel estimator})
is to show $\|\Pi_{\hat{\alpha}}D\|_{F}^{2}=O_{P}((nT)^{1/2})$ regardless
of the factor strength. To see why this is crucial, we note that the
best guaranteed rate for $\|D-\hat{D}\|_{F}^{2}$ is $\max\{n,T\}$ for
any estimator $\hat{D}$, \citep[see e.g.,][]{rohde2011estimation,candes2011tight}.
This rate is worse than $(nT)^{1/2}$ unless $T\asymp n$. One novelty
of our analysis is to show that although $\|\Pi_{\hat{\alpha}}X\|_{F}^{2}=\|\Pi_{\hat{\alpha}}(D+u)\|_{F}^{2}=O_{P}(\max\{n,T\})$,
we can separate it into $\|\Pi_{\hat{\alpha}}D\|_{F}^{2}=O_{P}((nT)^{1/2})$
and $\|\Pi_{\hat{\alpha}}u\|_{F}^{2}=O_{P}(\max\{n,T\})$.
The condition $T\geq n$ is motivated by the following insight. Under
our factor structure, it suffices to estimate either the factors or
the factor loadings, not both. Hence, perhaps we should estimate the
one with lower dimensionality. If $T\gg n$, then the factor loading
whose dimensionality is proportional to $n$ would be easier to estimate,
compared to factors whose dimensionality scales with $T$.
\begin{thm}
\label{thm: upper bnd panel}Consider the estimator $\hat{\beta}$ in (\ref{eq: panel estimator}).
Assume that $T\geq n$. Then for any $\eta\in(0,1)$, there exists
a constant $C_{\eta}>0$ such that
\[
\sup_{\theta\in\Theta}\mathbb{P}_{\theta}\left(\sqrt{nT}\left|\hat{\beta}-\beta\right|>C_{\eta}\right)\leq\eta.
\]
\end{thm}
We note that the requirement of $T\geq n$ is completely innocuous
and can be removed if we consider the following estimator
\[
\tilde{\beta}=\begin{cases}
\hat{\beta}(X,Y) & \text{if }T\geq n\\
\hat{\beta}(X',Y') & \text{otherwise}.
\end{cases}
\]
Theorem \ref{thm: upper bnd panel} formally establishes the uniform
rate of $(nT)^{-1/2}$ over $\Theta$. Hence, whether factors are
strong or weak and whether the number of factors is exactly known
would not prevent us from achieving the rate $(nT)^{-1/2}$.
\subsection{Can the currently known asymptotic variance hold without strong factors?}
Since the minimax rate of learning $\beta$ does not depend on the
strong factor assumption, the natural question is whether the strong
factor assumption is not important at all for estimation and inference.
Unfortunately, the answer is no. We shall show that (1) the efficiency
of inference on $\beta$ crucially depends on the strong factor condition
and (2) there is lack of adaptivity in the factor strength, resulting
in a tradeoff between efficiency and robustness.
From the perspective of semiparametric estimation, there is a good
reason to suspect that the strong factor condition might affect the
inference efficiency. We shall view the panel regression problem in
(\ref{eq: panel regression}) as a semiparametric problem, which would
yield a natural semiparametric lower bound for the asymptotic variance
in estimating $\beta$. However, we then realize that the typical
asymptotic variance in the literature \citep[e.g.,][]{bai2009panel,moon2014linear}
assuming strong factors can be much smaller than this lower bound.
This leads us to suspect that the strong factor condition might play
a role similar to strong parametric restrictions on nonparametric
components of semiparametric models.
To provide an analogy, consider the partial linear model with observations
$(Y_{i},Z_{i},W_{i})$ generated by $Y_{i}=f(W_{i})+Z_{i}\beta+\varepsilon_{i}$
and $Z_{i}=g(W_{i})+u_{i}$; under regularity conditions, the asymptotic
variance of estimating $\beta$ is bounded below by $\mathbb{E}\varepsilon_{i}^{2}/\mathbb{E} u_{i}^{2}$
when $f$ and $g$ are nonparametric functions or functions with high-dimensional
parameters, see e.g., \citep{chernozhukov1802.08667,newey2018cross,jankova2018semiparametric}.
Now we recall the model in (\ref{eq: panel regression}): $Y=M+X\beta+\varepsilon$
and $X=D+u$, where $M,D\in\mathbb{R}^{n\times T}$ are high-dimensional
nuisance parameters. We can view $M$ and $D$ as $f(W_{i})$ and
$g(W_{i})$ in the partial linear model, respectively. From this perspective,
we would expect the asymptotic variance for estimating $\beta$ to
be at least $\sigma_{\varepsilon}^{2}\sigma_{u}^{-2}$ in general.
However, the asymptotic variance derived in \citet{bai2009panel}
and \citet{moon2014linear} is
\[
\frac{\sigma_{\varepsilon}^{2}}{\sigma_{u}^{2}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)},
\]
where $\Pi_{M}=I_{n}-M(M'M)^{\dagger}M'$ and $\Pi_{M'}=I_{T}-M'(MM')^{\dagger}M$.
To formally address the efficiency problem, we again adopt the framework
of adaptivity. For simplicity, we assume that $\sigma_{u}=\sigma_{\varepsilon}=1$
in (\ref{eq: panel regression}). We consider the parameter space
\[
\Theta^{(2)}=\left\{ \theta=(M,D,1,1,\beta):\ {\rm rank}\, M\leq2,\ {\rm rank}\, D=1,\ |\beta|\leq1\right\} .
\]
Then we focus on adaptivity on a smaller space in which the strong
factor condition holds:
\begin{multline*}
\Theta^{(1)}=\Bigl\{\theta=(M,D,1,1,\beta)\in\Theta^{(2)}:\ {\rm rank}\, M=1,\ M'D=0,\ MD'=0,\\
\|M\|_{F}\geq\kappa_{1}\sqrt{nT},\ \|D\|_{F}\geq\kappa_{2}\sqrt{nT}\Bigr\},
\end{multline*}
where $\kappa_{1},\kappa_{2}>0$ are constants. The difference between
$\Theta^{(1)}$ and $\Theta^{(2)}$ is that $\Theta^{(2)}$ allows
for potentially weak factors and uncertainty in the number of factors
(${\rm rank}\, M$ can be either 1 or 2), while $\Theta^{(1)}$ only considers
known number of factors and assumes all the factors are strong. Our
analysis will focus on the question of whether robust confidence intervals
(uniform validity on $\Theta^{(2)}$) has worse efficiency on $\Theta^{(1)}$
than non-robust confidence intervals (uniform validity only on $\Theta^{(1)}$).
By \citet{bai2009panel} and \citet{moon2014linear} (among others),
the least-square estimator $\hat{\beta}_{{\rm LS}}$ (i.e., $(\hat{\beta}_{{\rm LS}},\hat{A})=\arg\min_{\beta\in\mathbb{R},\ A\in\mathbb{R}^{n\times T},\ {\rm rank}\, A\leq2}\|Y-A-X\beta\|_{F}^{2}$)
satisfies
\[
\frac{\sqrt{nT}(\hat{\beta}_{{\rm LS}}(X,Y)-\beta)}{\sigma(\theta)}\overset{d}{\rightarrow}N(0,1),
\]
over $\theta\in\Theta^{(1)}$, where for $\theta=(M,D,\sigma_{\varepsilon},\sigma_{u},\beta)$,
\[
\sigma(\theta):=\frac{\sigma_{\varepsilon}}{\sqrt{\sigma_{u}^{2}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)}}.
\]
For $\theta\in\Theta^{(1)}$, we have that
\[
\sigma(\theta)=\frac{\sigma_{\varepsilon}}{\sqrt{\sigma_{u}+{\rm trace\,}(\Pi_{M'}D'\Pi_{M}D)/(nT)}}=\frac{1}{\sqrt{1+\|D\|_{F}^{2}/(nT)}}\leq\frac{1}{\sqrt{1+\kappa_{2}^{2}}}.
\]
Therefore, a natural 95\%-confidence interval for parameters in $\Theta^{(1)}$
is
\begin{align}
CI_{*}(X,Y) & =\Bigl[\hat{\beta}_{{\rm LS}}(X,Y)-1.96(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2},\nonumber \\
& \qquad\qquad\qquad\qquad\hat{\beta}_{{\rm LS}}(X,Y)+1.96(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2}\Bigr].\label{eq: CI usual}
\end{align}
Existing results imply that
\[
\liminf_{n,T\rightarrow\infty}\inf_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}\left(\beta\in CI_{*}(X,Y)\right)\geq95\%.
\]
In other words, we have that
\begin{equation}
\limsup_{n,T\rightarrow\infty}\frac{\mathcal{L}(\Theta^{(1)})}{3.92(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2}}\leq1,\label{eq: minimax panel strong factor}
\end{equation}
where the quantity $\mathcal{L}(\Theta^{(1)})$ is defined as before: $\mathcal{L}(\Theta^{(1)})=\inf_{CI\in\Phi_{0.95}(\Theta^{(1)})}\sup_{\theta\in\Theta^{(1)}}\mathbb{E}_{\theta}|CI(X,Y)|$
is the minimax expected length of confidence intervals on $\Theta^{(1)}$
and $\Phi_{0.95}(\Theta^{(1)})=\{CI(\cdot)=[l(\cdot),u(\cdot)]:\ \inf_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}(\beta\in CI(X,Y))\geq0.95\}$
is the set of 95\%-confidence intervals. We also consider robust confidence
intervals, which have uniform validity over $\Theta^{(2)}$ and form
the set $\Phi_{0.95}(\Theta^{(2)})$. To study the impact of the robustness
requirement on efficiency, we revisit the concept of adaptivity by
studying
\[
\mathcal{L}(\Theta^{(1)},\Theta^{(2)})=\inf_{CI\in\Phi_{0.95}(\Theta^{(2)})}\sup_{\theta\in\Theta^{(1)}}\mathbb{E}_{\theta}|CI(X,Y)|.
\]
Both $\mathcal{L}(\Theta^{(1)})$ and $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$
measure performance of confidence intervals on $\Theta^{(1)}$. The
former considers non-robust confidence intervals (ones with validity
over $\Theta^{(1)})$, whereas the latter considers robust confidence
intervals (with validity over the larger set $\Theta^{(2)})$. If
$\mathcal{L}(\Theta^{(1)},\Theta^{(2)})/\mathcal{L}(\Theta^{(1)})$ is asymptotically
larger than one, then the extra robustness on $\Theta^{(2)}\backslash\Theta^{(1)}$
decreases the efficiency even on $\Theta^{(1)}$; if $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})/\mathcal{L}(\Theta^{(1)})$
converges to one, then one can gain extra robustness without sacrificing
efficiency. To characterize $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$, we
first derive the following result.
\begin{thm}
\label{thm: panel standard error}Let $CI(\cdot)=[l(\cdot),u(\cdot)]$
be a $(1-\alpha)$ confidence interval that has validity over $\Theta^{(2)}$,
i.e., $\inf_{\theta\in\Theta^{(2)}}\mathbb{P}_{\theta}(\beta\in CI(X,Y))\geq1-\alpha$
with $\alpha\in(0,1/2)$. Then for any $c\in(0,4)$, we have
\[
\sup_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}\left(|CI(X,Y)|\geq c(nT)^{-1/2}\right)\geq1-2\alpha-c.
\]
Moreover,
\[
\sup_{\theta\in\Theta^{(1)}}\mathbb{E}_{\theta}|CI(X,Y)|\geq(nT)^{-1/2}\left(1-2\alpha\right)^{2}/2.
\]
\end{thm}
Notice that the above lower bound does not depend on $\kappa_{2}$.
On the other hand, $\sqrt{nT}|CI_{*}(X,Y)|\asymp(1+\kappa_{2}^{2})^{-1/2}$
decreases with $\kappa_{2}$. Thus, for large enough $\kappa_{2}$,
$CI_{*}$ in (\ref{eq: CI usual}) violates the lower bound in Theorem
\ref{thm: panel standard error}. Since the lower bound is satisfied
by any confidence interval with uniform validity over $\Theta^{(2)}$,
it follows that any robust confidence interval will be wider than
$CI_{*}$ on $\Theta^{(1)}$. Equivalently, we can state the result
in terms of robustness (coverage guarantee for $CI_{*}$).
\begin{cor}
\label{cor: under-coverage panel data}Let $CI(\cdot)=[l(\cdot),u(\cdot)]$
be a random interval. Assume that for any $\eta\in(0,1)$,
\begin{equation}
\limsup_{n,T\rightarrow\infty}\sup_{\theta\in\Theta^{(1)}}\mathbb{P}_{\theta}\left(|CI(X,Y)|>3.92(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2}(1+\eta)\right)=0.\label{eq: panel reg good efficiency}
\end{equation}
Then
\[
\liminf_{n,T\rightarrow\infty}\inf_{\theta\in\Theta^{(2)}}\mathbb{P}_{\theta}\left(\beta\in CI(X,Y)\right)\leq\frac{1}{2}+1.96(1+\kappa_{2}^{2})^{-1/2}.
\]
\end{cor}
Notice that $CI_{*}$ satisfies (\ref{eq: panel reg good efficiency})
since $|CI_{*}(X,Y)|=3.92(nT)^{-1/2}(1+\kappa_{2}^{2})^{-1/2}$. By
Corollary \ref{cor: under-coverage panel data}, any confidence interval
that has a width similar to (or shorter than) that of $CI_{*}$ on
$\Theta^{(1)}$ will have coverage probability close to 1/2 on $\Theta^{(2)}$.
Finally, we compare $\mathcal{L}(\Theta^{(1)},\Theta^{(2)})$ and $\mathcal{L}(\Theta^{(1)})$.
Applying Theorem \ref{thm: panel standard error} with $\alpha=0.05$,
we obtain obtain
\[
\mathcal{L}(\Theta^{(1)},\Theta^{(2)})\geq(nT)^{-1/2}\left(1-2\alpha\right)^{2}/2=0.405(nT)^{-1/2}.
\]
In light of (\ref{eq: minimax panel strong factor}), this means
\[
\liminf_{n,T\rightarrow\infty}\frac{\mathcal{L}(\Theta^{(1)},\Theta^{(2)})}{\mathcal{L}(\Theta^{(1)})}\geq\frac{0.405}{3.92}\sqrt{1+\kappa_{2}^{2}}>\frac{\sqrt{1+\kappa_{2}^{2}}}{9.7}.
\]
Therefore, any robust confidence interval is asymptotically wider
than any non-robust confidence interval whenever $\kappa_{2}>9.65$.
In other words, requiring validity on $\Theta^{(2)}\backslash\Theta^{(1)}$
(allowing for weak factors and unknown number of factors) would lead
to efficiency loss on $\Theta^{(1)}$ if $\kappa_{2}>9.65$.
Although the constant of 9.65 is not the optimal constant, the analysis
highlights the lack of adaptivity in inference. Without strong factor
in the fixed effects, strong factor components in $X$ would result
in a stark loss of efficiency. This can be explained. When the fixed
effects $M$ have strong factors, the projections $\Pi_{M}$ and $\Pi_{M'}$
can be estimated well and thus we can safely identify components in
$X$ that cannot be absorbed by the fixed effects; as a result, the
variations in $\Pi_{M}D\Pi_{M'}+u$ can be used to learn $\beta$.
However, when $M$ does not have strong factors, it is quite difficult
to learn projections $\Pi_{M}$ and $\Pi_{M'}$ and hence we cannot
clearly tell which part of $X$ is left after removing components
correlated with the fixed effects; consequently, we will not be sure
that any part of $D$ can be used to learn $\beta$ and instead will
only consider variations in $u$ simply to be on the safe side (ensure
coverage probability in all cases), resulting in a confidence interval
with length unrelated to $\kappa_{2}$ (representing factor strength
in $D$).
Another implication is that when there are potential weak factors,
uncertainty in the number of factors leads to efficiency loss. This
is in contrast with results in \citet{moon2014linear}, who established
that when all the factors are strong, uncertainty in the number of
factors does not cause efficiency loss.