EconBase
← Back to paper

Robust Estimation and Inference in Panels with Interactive Fixed Effects

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

126,818 characters

Robust Estimation and Inference in Panels with Interactive Fixed Effects



\title{\bf Robust Estimation and Inference\\in Panels with Interactive Fixed
  Effects\thanks{
    We thank the participants of the numerous seminars and conferences for helpful comments
and suggestions.
    We also thank Riccardo D'Adamo and Chen-Wei Hsiang for their excellent research assistance.
    Any remaining errors are our own.
    Armstrong gratefully acknowledges support by the National
    Science Foundation Grant SES-2049765.
    Weidner gratefully acknowledges support through the European Research Council grant ERC-2018-CoG-819086-PANEDA.
    Zeleneev gratefully acknowledges the generous funding from the UK
Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee (Grant Ref:
EP/X02931X/1).
    }}
\author{\setcounter{footnote}{2}Timothy B. Armstrong\thanks{
University of Southern California. Email: \texttt{[email removed]} }
\and
Martin Weidner
\thanks{
University of Oxford. Email: \texttt{[email removed]} }
\and
Andrei Zeleneev
\thanks{
University College London. Email: \texttt{[email removed]} }
}
\date{May 2025}

\maketitle
\thispagestyle{empty}
\setcounter{page}{0}

\bigskip
\begin{abstract}
\noindent
We consider estimation and inference for a regression coefficient in panels with interactive fixed effects (i.e., with a factor structure). We demonstrate that existing estimators and confidence intervals (CIs) can be heavily biased and size-distorted when some of the factors are weak. We propose estimators with improved rates of convergence and bias-aware CIs that remain valid uniformly, regardless of factor strength. Our approach applies the theory of minimax linear estimation to form a debiased estimate, using a nuclear norm bound on the error of an initial estimate of the interactive fixed effects. Our resulting bias-aware CIs take into account the remaining bias caused by weak factors. Monte Carlo experiments show substantial improvements over conventional methods when factors are weak, with minimal costs to estimation accuracy when factors are strong.
\end{abstract}

\vskip 3cm

\newpage

\section{Introduction}

In this paper, we consider a linear panel regression model of the form
 \begin{align}
  Y_{it} = X_{it} \beta + \sum_{k=1}^K Z_{k,it}\delta_k +  \Gamma_{it} + U_{it} ,
  \label{model0}
\end{align}
where $Y_{it}, X_{it}, Z_{1,it}, \ldots, Z_{K,it} \in \mathbb{R}$  are the observed outcome variable and covariates for units $i=1,\ldots,N$ and time periods $t=1,\ldots,T$.
The error components  $\Gamma_{it}  \in \mathbb{R}$ and $U_{it}  \in \mathbb{R}$ are unobserved, and the regression coefficients   $\beta, \delta_1, \ldots, \delta_K \in \mathbb{R}$ are unknown.
The  parameter of interest is $\beta \in \mathbb{R}$, the  coefficient on $X_{it}$. We are interested in ``large panels'', where both $N$ and $T$ are relatively large.

The error component $U_{it}$ is modelled as a mean-zero random shock that is uncorrelated with the regressors $X_{it}$ and $Z_{k,it}$
and that is at most weakly autocorrelated across $i$ and over $t$. By contrast, the error component $\Gamma_{it}$ can be correlated with $X_{it}$ and $Z_{k,it}$
and can also be strongly autocorrelated across $i$ and over $t$. Of course, further restrictions on $\Gamma_{it}$ are required to allow estimation and inference on $\beta$.
For example, the additive fixed effect model imposes that $\Gamma_{it} = \alpha_i + \gamma_t$, where
$\alpha_i$ accounts for any omitted variable that is constant over time, and $\gamma_t$ for any omitted variable that is constant across units.
Instead of this additive fixed effect model we consider the so-called interactive fixed effect model, where
 \begin{align}
     \Gamma_{it} =  \sum_{r=1}^R \,  \lambda_{ir}\, f_{tr} \; .
     \label{FactorModel}
\end{align}
Here, the $\lambda_{ir}$ and $f_{tr} $ can either be interpreted as unknown parameters or as unobserved shocks. This model for $\Gamma_{it} $ is also known as a factor model, with factor loadings $\lambda_{ir}$ and factors $f_{tr}$. We will use the terms factor and interactive fixed effect interchangeably. The number of factors $R$
is unknown, but is assumed to be small relative to $N$ and $T$. The interactive fixed effect model is attractive because it introduces enough restrictions to allow
estimation and inference on $\beta$ while still incorporating or approximating a large class of data generating processes (DGPs) for $ \Gamma_{it} $.

The existing econometrics literature on panel regressions with interactive fixed
effects is  quite large. Since the seminal work of \cite{Pesaran2006estimation} and
\cite{bai2009panel}, developing tools for estimation and inference on $\beta$ in
model~\eqref{model0}-\eqref{FactorModel} under large $N$ and large $T$
asymptotics has been a primary focus of this literature. Specifically, \citet{Pesaran2006estimation} introduces the common correlated effects (CCE) estimator, which uses cross-sectional averages of the observed variables as proxies for the unobserved factors. \citet{bai2009panel} derives the large $N$, $T$ properties of the least-squares (LS) estimator that jointly minimizes the sum of squared residuals over the regression coefficients, factors, and factor loadings.\footnote{This estimator was first introduced by \cite{Kiefer1980}.}



\cite{bai2009panel} shows that, under appropriate assumptions, the LS estimator for the regression coefficients is $\sqrt{NT}$-consistent and asymptotically normally distributed
as both $N$ and $T$ grow to infinity. One of the key assumptions imposed for this result is the so-called ``strong factor assumption'', which requires all the factor loadings
$ \lambda_{ir}$ and factors $f_{tr}$ to have sufficient variation across $i$ and over $t$, respectively. If the strong factor assumption is violated, then the LS estimator
for $ \lambda_{ir}$ and  $f_{tr}$  may be unable to pick up the true loadings and factors correctly, because the ``weak factors''\footnote{
See, for example, \cite{Onatski2010,Onatski2012} for a discussion and formalization of the notion of weak factors.
}  in $ \Gamma_{it}$ cannot be
distinguished from the noise in $U_{it}$.
This can lead to substantial bias and misleading inference, due to omitted
variables bias from $\Gamma_{it}$ that is not picked up by the estimator.



\begin{figure}[tb]
    \begin{center}
      \begin{subfigure}{0.49\textwidth}
        \includegraphics[width=1.0\linewidth]{figures/N50_T100_R1_kappa000.pdf}
        \caption{No ID ($\kappa = 0.00$)}
      \end{subfigure}
      \begin{subfigure}{0.49\textwidth}
        \includegraphics[width=1.0\linewidth]{figures/N50_T100_R1_kappa010.pdf}
        \caption{Weak ID ($\kappa = 0.10$)}
      \end{subfigure}
      \begin{subfigure}{0.49\textwidth}
        \includegraphics[width=1.0\linewidth]{figures/N50_T100_R1_kappa020.pdf}
        \caption{Weak ID ($\kappa = 0.20$)}
      \end{subfigure}
      \begin{subfigure}{0.49\textwidth}
        \includegraphics[width=1.0\linewidth]{figures/N50_T100_R1_kappa100.pdf}
        \caption{Strong ID ($\kappa = 1.00$)}
      \end{subfigure}
    \end{center}
  \caption{Finite sample distributions of the LS and the debiased estimators, $N=100$, $T=50$, $R=1$}
  \label{fig: intro}
\end{figure}

To illustrate how this can lead to problems with conventional estimates and CIs
for $\beta$, Figure \ref{fig: intro} presents a subset of the results of our Monte Carlo
study.\footnote{A detailed description of the numerical experiment is provided
  in Section \ref{ssec: MC}.}
When the factors are nonexistent (panel a) or strongly identified
(panel d), the distribution of the LS estimator (in blue) is centered at the
true parameter value $\beta$ (equal to $0$ in this case).  However, when the
factors are present but weak enough that they are difficult to estimate (panels b
and c), the LS estimator is heavily biased and non-normally distributed.  In our
Monte Carlo study in Section \ref{sec: numerical}, we show that this indeed leads to
severe coverage distortion, with conventional CIs based on the LS estimator
having almost zero coverage.





In this paper, we address this issue by developing new tools for estimation and
inference on $\beta$ in the model \eqref{model0}.
We develop a debiased estimator along with a bound on the remaining bias, which
we use to construct a bias-aware confidence interval.
As illustrated in Figure \ref{fig: intro}, our debiased estimator (shown in red)
substantially decreases the bias of the LS estimator when factors are weak,
leading to a large improvement in overall estimation error.
In addition, this improved performance under weak factors does not come at a
substantial cost to performance when factors are strong or nonexistent: our
debiased estimator performs similarly to the LS estimator in these cases.
Importantly, our CI
requires only an upper bound on the number of factors: we show that it is valid
uniformly over a large class of DGPs that allows for weak, strong or nonexistent
factors up to a specified upper bound on the number of factors.  We derive rates
of convergence that hold uniformly over this class of DGPs, and we show that our
estimator achieves a faster uniform rate of convergence than existing approaches
when weak factors are allowed.  In the case where $N$ and $T$ grow at the same
rate, our estimator achieves the parametric $\sqrt{NT}$ rate.



Our debiasing approach uses a preliminary estimate $\hat\Gamma_{\rm pre}$ of the
individual effect matrix $\Gamma$ along with a bound $\hat C$ on the nuclear norm $\Vert \Gamma - \hat
\Gamma_{\rm pre}\Vert_*$ of its estimation error.
Letting $\tilde \Gamma:= \Gamma - \hat \Gamma_{\rm pre}$, we then consider
the augmented outcomes
\begin{align*}
  \tilde Y_{it} := Y_{it} - \hat \Gamma_{{\rm pre}, it} = X_{it} \beta + \sum_{k=1}^K Z_{k,it}\delta_k + \tilde \Gamma_{it} + U_{it}.
\end{align*}
Treating $\tilde \Gamma_{it}$ as nuisance parameters satisfying a convex
constraint $\Vert \tilde \Gamma \Vert_{*} \le \hat C$, we derive linear weights
$A_{it}$ such that the estimator $\sum_{i=1}^N\sum_{t=1}^TA_{it}\tilde Y_{it}$   for $\beta$
optimally uses this constraint, using the theory of minimax linear estimators \citep[see][]{ibragimov_nonparametric_1985,donoho1994statistical,armstrong2018optimal}.
In particular, the resulting weights $A_{it}$ control the remaining omitted
variables bias $\sum_{i=1}^N\sum_{t=1}^TA_{it}\tilde \Gamma_{it}$ due to
possible weak factors in $\tilde \Gamma=\Gamma-\hat\Gamma_{\rm pre}$ not picked
up by the initial estimate $\hat\Gamma_{\rm pre}$.

A key step in deriving our CI is the construction of the preliminary estimator
$\hat\Gamma_{\rm pre}$ and bound $\hat C$ on the nuclear norm of its estimation
error.  Our CI is bias-aware: it uses the bound $\hat C$ to explicitly take into
account any remaining bias in the debiased estimator.  Our bound is feasible
once an upper bound on the number of factors is specified.  In our Monte Carlo
study, we find that, while our CIs are often conservative, they are about as
wide as an ``oracle'' CI that uses an infeasible critical value to correct the
coverage of a CI based on the standard LS estimator.

While our results allow for arbitrary sequences of weak factors, our conditions on other aspects of the model are similar to \citet{bai2009panel} and \citet{MoonWeidner2015}.  An important condition is that the covariate of interest $X_{it}$ must not itself be entirely explained by a low dimensional factor model.
For example, in a panel where $X_{it}$ is the minimum hourly wage in state $i$ and year $t$, we would require that states change their minimum wage laws in different years, and that this is done sufficiently often to generate variation in $X_{it}$ that cannot be explained by a small number of factors $f_t$.
This rules out settings where $X_{it}$ is an indicator variable for a policy that affects a subset of the units and occurs only during a single time period: in this case, $X_{it}=\lambda_i\cdot f_t$ where $\lambda_i$ is an indicator variable for unit $i$ undergoing the policy change and $f_t$ is an indicator variable for periods after the policy change.
See Section \ref{asymptotic_section} for formal conditions and further discussion.


A special case of the factor model is the grouped unobserved heterogeneity model
considered by \citet{bonhomme2015grouped}.  In this model,
$\Gamma_{it}=\alpha_{g(i),t}$, where $g(\cdot)$ is an unknown function mapping
individuals $i$ to a group index $g(i)\in \{1,\ldots, R\}$.  This takes the form
of the factor model (\ref{FactorModel}) with $\lambda_{ir}=1$ if $g(i)=r$ and
$0$ otherwise, and with $f_{tr}=\alpha_{r,t}$.  The strong factor assumption
corresponds to the strong group separation assumption imposed in this literature
(e.g., Assumption 2(b) in
\citealp{bonhomme2015grouped}) which imposes that the group means $\alpha_{r,\cdot}=(\alpha_{r,1},\ldots,\alpha_{r,T})'$
are sufficiently far away for different groups $r$.  Our results apply in this
setting and allow for this assumption to be relaxed.  An interesting question
for future research is whether it is possible to modify our approach to take
advantage of the additional structure in the grouped unobserved heterogeneity
model.





















\bigskip

{\noindent \bf Related literature}


\noindent
The papers by \citet{Pesaran2006estimation} and \citet{bai2009panel} mentioned previously have motivated a large follow up literature on large $N$ and $T$ analysis of panel models with interactive effects.  \cite{bai2016econometric} provides a review with further references.  Another literature has proposed alternative estimation methods along with asymptotic analysis in the regime with $T$ fixed and $N$ increasing.  This includes the quasi-difference approach of \citet{HoltzEakin-Newey-Rosen1988} and generalized method of moments approaches of \cite{AhnLeeSchmidt2001,AhnLeeSchmidt2013}.
More recent papers analyzing the fixed $T$ large $N$ regime include \cite{robertson2015iv}, \cite{juodis2018fixed}, \cite{westerlund2019cce}, \cite{higgins2021fixed}, \cite{juodis2022linear}.
 None of these papers provide inference methods that remain valid when factors are weak or rank-deficient (e.g.\ $f=0$).
\cite{Chamberlain2009} derive estimators that satisfy a Bayes-minimax property over a certain class of priors in a finite sample setting that includes a version of the model (\ref{FactorModel}).
This Bayes-minimax property does not, however, translate to a guarantee on coverage or estimation error under weak factors.

A special case of the violation of the  strong factor assumption is when some factor are equal to zero, while all other factors are strong;
the inference results of \cite{bai2009panel} are usually robust towards this specific violation of the strong factor assumption
\citep{MoonWeidner2015}. This robustness, however, does not carry over to more general weak factors in the DGP of $\Gamma_{it}$,
 as illustrated by Figure~\ref{fig: intro}.






The problem of weak factors is related to the problem of omitted variable bias of LASSO estimators in high dimensional regression that is the focus of debiased LASSO estimators \citep[see][]{belloni_inference_2014,javanmard_confidence_2014,van_de_geer_asymptotically_2014,zhang_confidence_2014}.
Just as LASSO estimators omit variables with coefficients that are large enough to cause omitted variables bias but too small to distinguish from zero, weak factors in $\Gamma$ can be difficult to estimate, leading to omitted variables bias in conventional estimates of $\beta$.
Our approach to using minimax linear estimation to debias an initial estimate
mirrors the approach of \citet{javanmard_confidence_2014} to debiasing the LASSO.
We discuss this connection further in Section \ref{sec:comparison_to_other_results}.
\citet{hirshberg_augmented_2020} provide a general discussion and further
references for minimax linear debiasing; we refer to this general approach as augmented
linear estimation following their terminology.  Minimax linear estimation itself
goes back at least to \citet{ibragimov_nonparametric_1985}, with further results
on this approach and its optimality properties in \citet{donoho1994statistical},
\citet{armstrong2018optimal} and \citet{yata_optimal_2021}, among others.  The
particular form of the minimax estimator used for debiasing in our setup follows
from a formula given in \citet{armstrong_bias-aware_2020}.



Requiring $\Gamma_{it}$ to have the factor structure~\eqref{FactorModel} is
equivalent to requiring the matrix of unobserved effects $\Gamma$ to have rank
at most $R$, i.e., having ${\rm rank}(\Gamma) \leq R$. Bounding the nuclear norm of $\tilde \Gamma$ or $\Gamma$ instead can also be seen as a convex relaxation of this requirement. Similar convexifications of the rank constraint have been widely used in the matrix completion literature
(e.g., \citealt{RechtFazelParrilo2010} and \citealt{Hastieetal2015} for recent surveys),
and for reduced rank regression estimation
(e.g., \citealt{RohdeTsybakov2011}).
In the econometrics literature, the numerous applications of this idea include, for example, estimation of pure factor models \citep{BaiNg2017}, estimation of panel regression models with homogeneous \citep{moon2018nuclear,beyhum2019square} and heterogeneous coefficients \citep{chernozhukov2019inference}, estimation of treatment effects \citep{athey_matrix_2021,fernandez2021low}, and many others.\footnote{For example, recent economic applications of nuclear norm and related penalization methods also include latent community detection \citep{alidaee2020recovering,ma2022detecting}, quantile regression \citep{belloni2019high,wang2022low,feng_2023}, and estimation of panel threshold models and high-dimensional VARs (\citealp{miao2020panel} and \citealp{miao2023high}).}
However, none of these papers obtain asymptotically valid CIs or improved rates of convergence under weak factors.

In recent work, \citet{chetverikov2022spectral} propose an estimator that, like ours, achieves a faster rate of convergence than conventional approaches under weak factors.\footnote{The main focus of \citet{chetverikov2022spectral} is the grouped effects model of \citet{bonhomme2015grouped}, which is a special case of the interactive fixed effects setting we consider here.  However, the authors extend their results to the general interactive fixed effects setting.}
While \citet{chetverikov2022spectral} allow for weak factors in some of their estimation results, they assume strong factors when constructing CIs.
The estimation approach in \citet{chetverikov2022spectral} also differs from our approach by using modelling assumptions that place a factor structure on the covariate matrix $X$.

Our focus is on allowing for weak factors without imposing additional assumptions on the error term $U$, such as homoskedasticity or full independence from the individual effects $\Gamma$ and regressor $X$.  Such additional structure allows for further identifying information by making it easier to distinguish between the error term $U$ and the individual effects $\Gamma$, leading to a fundamentally different analysis.  \citet{zhu2019well} derives asymptotic upper and lower bounds for estimators and CIs in a setting with possible weak factors under homoskedastic and fully independent errors.  The estimators and CIs constructed by \citet{zhu2019well} take advantage of the additional structure of \citeauthor{zhu2019well}'s setting, making them inapplicable in ours.  However, the \emph{lower} bounds derived by \citet{zhu2019well} are immediately relevant: they show that no CI can be asymptotically valid under weak factors while mimicking the performance of the CI of \citet{bai2009panel} when factors are strong.

As discussed above, our assumptions rule out the case where $X_{it}$ is an indicator variable for a policy that affects a subset of units starting in the same time period.
Recent papers that analyze such settings include \citet{ferman_synthetic_2021} and \citet{arkhangelsky_synthetic_2021}.
The fact that $X_{it}$ is collinear with the confounding factor model in this setting presents a fundamental identification issue that requires placing additional conditions on the model.
In contrast to this literature, our goal is to leverage variation in $X_{it}$ that cannot be explained by a low dimensional factor model in settings where such variation exists.



\citet{beyhum2022factor}, \citet{fan2022learning}, and \citet{bai2023approximate} consider estimation and inference in various settings under a regime in which a lower bound on the strength of the factors can decrease with $N$ and $T$, but is large enough that factors can be consistently estimated.  This is analogous to the ``semi-strong'' regime in weak instrument and related settings; see \citet{andrews2012estimation}.  While the semi-strong regime requires careful theoretical analysis, the fact that factors can be consistently estimated leads to asymptotically unbiased and normal estimators for the main effect $\beta$.  Our results apply to semi-strong and strong regimes as well, while also allowing for weak factor regimes in which factors cannot be consistently estimated.




Finally, \citet{cox2024weak} develops tools for inference in low-dimensional factor models with weak identification. In \citet{cox2024weak}, the primary objects of interests are the covariance of the factors and the loadings. The baseline model in \citet{cox2024weak} does not include observed covariates, whereas we focus on estimation and inference on $\beta$, the coefficient on $X_{it}$, exclusively.\footnote{\citet{cox2024weak} mentions that observed covariates could, in principle, be incorporated in his framework as long as they are uncorrelated with the unobserved effects, which is a primary worry in the panel literature.}

\bigskip

The rest of this paper is organized as follows. Section \ref{sec:MainIdea} introduces the framework and describes construction of the debiased estimator and bias-aware CI. Section \ref{sec:linear_factor_implementation} provides implementation details. Section \ref{asymptotic_section} provides formal statistical guarantees. Section \ref{sec: numerical} considers numerical and empirical illustrations.
A supplementary appendix contains all proofs
and additional results for the numerical and empirical illustrations.







\section{Construction of robust estimates and confidence intervals}
\label{sec:MainIdea}


\subsection{Setup}

We consider a panel setting in which we observe a scalar outcome
$Y_{it}$, a scalar covariate $X_{it}$ of interest
and additional control covariates $\{Z_{k,it}\}_{k=1}^K$
for $i=1,\ldots,N$, $t=1,\ldots, T$,
which follow
the regression model (\ref{model0}).
The error term $U_{it}$ is assumed to be mean zero conditional on $X$,
$\{Z_{k,it}\}_{k=1}^K$ and $\Gamma$,\footnote{We note that this requires strict
  exogeneity and in particular rules out
  using lagged outcomes as covariates.  We leave extensions to models with lagged outcomes as a topic for future research.}
but we allow for
heteroskedasticity, which may depend on $X_{it}$ and $\Gamma_{it}$, as well as
some weak dependence.
We write the model in matrix notation as
\begin{align}\label{model0_matrix}
  Y = X \beta + Z\cdot \delta +  \Gamma + U \, ,
  \qquad\qquad
  \mathbb{E}[U|X,Z,\Gamma] = 0 \, ,
\end{align}
where
$Z$ denotes the three dimensional array
$\{Z_{k,it}\}$
and we define $Z\cdot \delta = \sum_{k=1}^K Z_{k}\delta_k$ where $Z_k$ denotes
the matrix with $i,t$-th element $Z_{k,it}$.
We use $\lambda$ to denote the $N\times R$ matrix of loadings $\lambda_{ir}$ and $f$ to denote the $T\times R$ matrix of factors $f_{tr}$, so that (\ref{FactorModel}) can be written in matrix form as $\Gamma=\lambda f'$.

We are interested in the coefficient $\beta$ of $X_{it}$, which can be
interpreted as the effect of a treatment variable $X_{it}$ in a constant
treatment effects model (we discuss extensions to heterogeneous treatment
effects in Remark \ref{het_te_remark}).
For concreteness, we use panel notation, and we refer to $i$ and $t$ as
individuals and time periods respectively.  However, we allow for other settings
such as network data in which $i$ and $t$ both index individuals in a network.
While we will assume a low rank structure on $\Gamma$, we allow for arbitrary
dependence between the covariate $X_{it}$ and the individual effect $\Gamma_{it}$.


A key ingredient in our approach is an initial estimate $\hat\Gamma$ of $\Gamma$ and a bound on its estimator in the nuclear norm, which holds with probability approaching one:
\begin{align}
  \big\| \widetilde \Gamma  \big\|_*
  \leq \hat C \, ,
  \qquad\text{where}\qquad
  \widetilde\Gamma := \Gamma - \hat\Gamma \, .
  \label{NNcondition}
\end{align}
Here, $\|\cdot\|_*$ denotes the nuclear norm of the argument matrix,
and  $\hat C \geq 0$ is a known or estimated constant.
We describe our estimate $\hat\Gamma$ and bound $\hat C$ in Section \ref{sec:linear_factor_implementation}, and we state a formal result giving conditions under which the bound holds with probability approaching one in Section \ref{asymptotic_section}.
This bound depends on an upper bound for the number of factors $R$, which must be specified a priori.  Importantly, our approach does not require specifying the exact number of factors: some of the factors may be zero, in addition to the possibility of being ``weak'' in the sense of being close to zero.
We emphasize that obtaining an computable upper bound $\hat C$ that enables construction of our CI is itself one of the main technical contributions of this paper.\footnote{As we discuss further in Section \ref{asymptotic_section}, a tighter CI can be constructed by bounding the difference between the estimate $\hat\Gamma$ and $\Gamma+P_{\lambda} U$, where $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$, with $M^+$ denoting the Moore–Penrose inverse of a matrix $M$.
The implementation in Section \ref{sec:linear_factor_implementation} is for the tighter CI that uses these arguments.
For ease of exposition, however, we focus on using the bound \eqref{NNcondition} directly in the remainder of this section.}


\begin{remark}
    Although the main focus of this paper is on models with the linear factor
    structure~\eqref{FactorModel}, the methodology presented in this section applies to general
    interactive fixed effects models as long as it is possible to construct a preliminary estimator $\hat \Gamma$ and a bound $\hat C$ satisfying \eqref{NNcondition}. For example, we conjecture that our method can also be extended to nonlinear factor models with $\Gamma_{it} = g(\lambda_i,f_t)$, where $g(\cdot,\cdot)$ is some unknown function (e.g., \citealp{zeleneev2019identification,freeman2023linear}). As noted, for example, in \citet{fernandez2021low}, such $\Gamma$ can be approximated by a low-rank matrix with a (slowly) growing rank $R$. Hence, we expect that our method can be applied in this setting with $\hat \Gamma$ constructed using a growing $R$ and $\hat C$ adjusted for the low-rank approximation error (if needed), in the same spirit as sieve approximations are used in nonparametric estimation.
\end{remark}




\subsection{Augmented linear estimators and CIs}

We first define a class of estimators and CIs, indexed by an $N\times T$ matrix
$A$.  We then provide a choice of the matrix $A$, based on finite
sample optimality in an idealized setting.
Our class of estimators is given in the following definition.

\begin{definition}
  \label{DefAugLinEst}
  Let $A=A(X,Z)$ be an $N \times T$ matrix of weights $A_{it} \in \mathbb{R}$ that can depend on the matrix $X$
  and array $Z$.  Let $\hat\Gamma$ be an initial estimate of $\Gamma$, and let
  $\widetilde Y=Y-\hat\Gamma$.  The \emph{augmented linear estimator} with
  weight matrix $A$ and initial estimate $\hat\Gamma$ is given by
  \begin{align}
    \hat\beta_A
    &:= \sum_{i=1}^N \sum_{t=1}^T A_{it}\widetilde Y_{it}
      =\langle A, \widetilde Y \rangle_F.
      \label{DefHatBeta}
  \end{align}

\end{definition}
 Here, $\langle \cdot, \cdot \rangle_F$ denotes the   entry-wise inner product between the argument matrices.


The estimator $\hat\beta_A=\langle
A,\widetilde Y \rangle_F$ applies a linear estimator after an initial estimation
step in which the initial estimate $\hat\Gamma$ is subtracted from the outcome
$Y$.
This mirrors applications of this idea in other settings going back to
\citet{javanmard_confidence_2014}; see \citet{hirshberg_augmented_2020} for
references (the term ``augmented linear estimation'' is used in the latter
paper).


To analyze this class of estimators, note that subtracting the initial estimate
from both sides of the equation (\ref{model0_matrix}) gives
\begin{align}
  \widetilde Y = X \beta + Z\cdot \delta + \widetilde \Gamma + U
  \label{model}
\end{align}
(recall that $\widetilde Y=Y-\hat\Gamma$ and $\widetilde
\Gamma=\Gamma-\hat\Gamma$).
Our choice of the matrix $A$ will be motivated by a heuristic in which we consider the model (\ref{model}) with $\widetilde Y$ as the observed outcome and $\widetilde \Gamma$ a nuisance parameter such that $U$ is mean zero conditional on $X,Z$ and $\tilde \Gamma$, with the bound (\ref{NNcondition}) interpreted as a deterministic bound that holds with $\hat C$ nonrandom.
This heuristic is not literally true, since $\widetilde\Gamma$ depends on $U$ through the estimation error in the initial estimate $\hat\Gamma$.
Nonetheless, the CIs and estimators we obtain will be asymptotically valid and consistent respectively, under conditions that we give in Section \ref{asymptotic_section}.

Following this heuristic, we consider the decomposition
\begin{align}\label{bias_variance_decomposition}
  &\hat\beta_A-\beta
    =
    \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A)
    + \langle A, U \rangle_F
\end{align}
where
\begin{align}\label{bias_def_eq}
  \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A):=\left( \langle A,X
  \rangle_F - 1 \right)\beta + \langle A,Z\cdot \delta \rangle_F + \langle A,
  \widetilde\Gamma \rangle_F.
\end{align}
Under the heuristic where $\tilde\Gamma$ is nuisance parameter in the model (\ref{model}), $\operatorname{bias}_{\beta,\delta,\widetilde \Gamma}$ gives the bias of the estimator $\hat\beta_{A}$ conditional on $X,Z$ and $\tilde \Gamma$.
In reality, $\operatorname{bias}_{\beta,\delta,\widetilde \Gamma}$ does not literally
give the bias or conditional bias of $\hat\beta_A$, since conditioning on
$\widetilde\Gamma=\Gamma-\hat\Gamma$ means conditioning on an information set
that depends on $Y$ through the preliminary estimate $\hat\Gamma$.
We nonetheless refer to $\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ as a bias term, following our heuristic.

Let $\widehat{\operatorname{se}}$ be an estimate
of the standard deviation of $\langle A,U \rangle_F=\sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it}$.
For example, to allow for arbitrary
heteroskedasticity in $U_{it}$ while imposing independence across $i$ and $t$,
we can use $\widehat{\operatorname{se}}=\sqrt{\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{it}^2}$
where $\hat U_{it}$ denotes residuals from an initial regression.
If $\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ were zero,
then we could form a CI by adding and subtracting a normal critical value times
$\widehat{\operatorname{se}}$.
To take into account the possibility that
$\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ will in general be nonnegligible in
our setting, we use the bound (\ref{NNcondition}) to obtain an upper bound on
the bias term.
In particular,
when (\ref{NNcondition}) holds, we have
$\left| \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A) \right|
\le \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$, where
for general $C \geq 0$ we define
\begin{align}
  \overline{\operatorname{bias}}_C(\hat\beta_A)
  &:= \sup_{ \beta,\delta,\widetilde\Gamma: \|\widetilde\Gamma\|_*\le C} \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A)
 \nonumber \\
 &   =
   \begin{cases}
     \displaystyle \sup_{ \widetilde\Gamma: \|\widetilde\Gamma\|_*\le C}
  \langle A, \widetilde\Gamma \rangle_F & \text{if} \, \langle A, X \rangle_F=1,\, \text{and } \langle A, Z_k \rangle_F=0,\,  \text{for } k=1,\ldots K,  \\
     \infty & \text{otherwise}
   \end{cases}
 \nonumber \\
 &   =
   \begin{cases}
     C s_1(A) & \text{if} \, \langle A, X \rangle_F=1,\, \text{and } \langle A, Z_k \rangle_F=0,\,  \text{for } k=1,\ldots K,  \\
     \infty & \text{otherwise.}
   \end{cases}
   \label{WorstCaseBias}
\end{align}
Here, for the second equality we used that the supremum over $\beta$ and $\delta$ is unbounded unless
$\langle A, X \rangle_F=1$ and $\langle A, Z_k \rangle_F=0$, and for the final step we used that
the nuclear norm $ \| \cdot \|_*$ is dual to the spectral norm,  which we  denote by $s_1(\cdot)$ since it is equal to the
largest singular value of the argument matrix.
We refer to $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ as the worst-case bias of the
estimator $\hat\beta_A$ (again, this terminology reflects the heuristic in which $\tilde\Gamma$ is treated as a nuisance parameter in (\ref{model}) rather than estimation error from the initial estimate $\hat\Gamma$).

Note that, whereas
$\operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A)$ depends on the unknown
matrix of individual effects $\Gamma$ through the matrix
$\widetilde\Gamma=\Gamma-\hat\Gamma$,
$\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ is feasible to compute once a bound $\hat C$ is given.
Taking into account the possible bias leads to a \emph{bias-aware} CI:
\begin{align}\label{general_A_CI_eq}
  \left\{ \hat\beta_A \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2} \widehat{\operatorname{se}} \right]   \right\}.
\end{align}
To motivate this CI, note that the probability that the lower endpoint is greater
than $\beta$ is
\begin{align*}
  &P\left( \hat\beta_A - \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) - z_{1-\alpha/2} \widehat{\operatorname{se}} > \beta
    \right)
    = P\left( \sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it} + \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A) > \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2} \widehat{\operatorname{se}}
    \right)  \\
  &\le P\left( \sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it} > z_{1-\alpha/2} \widehat{\operatorname{se}}
    \right)
    \approx \alpha/2,
\end{align*}
where the last step assumes that $\sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it}$ is approximately normally distributed with
zero mean and standard deviation close to $ \widehat{\operatorname{se}}$. We provide formal justifications for this later. By a similar argument, the probability that the upper endpoint is less than $\beta$ can be bounded by $\alpha/2$, and these calculations together imply that the coverage of our CI is approximately at least $1-\alpha$.


\begin{remark}\label{het_te_remark}
  In principle, our approach can be extended to a heterogeneous treatment effect model where the constant coefficient $\beta$ is replaced by an individual specific coefficient $\beta_{it}$ that is allowed to vary with $i$ and $t$.  In particular, if a bound on the nuclear norm of the matrix of coefficients $\beta_{it}$ or on the error of preliminary estimates of these coefficients is available in addition to such a bound for $\Gamma$, we can use minimax linear debiasing to estimate a linear functional of the individual specific effects $\beta_{it}$.  For example, the linear functional $\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T \beta_{it}$ gives the average treatment effect of a one-unit change in $X_{it}$ over the $NT$ units in a setting where $\beta_{it}$ is interpreted as the causal effect of a change in the variable $X_{it}$.
  Deriving a computable bound on the nuclear norm error of an initial estimate of the coefficients $\beta_{it}$ in this case is nontrivial, however, and we leave this question for future research.
\end{remark}


\subsection{Choice of weights $A=(A_{it})$}\label{choice_of_weight_section}

As described in the last subsection,
one can construct valid confidence intervals for $\beta$
of the form \eqref{general_A_CI_eq} for any choice of weight matrix $A$, subject to weak regularity conditions.
To get a simple baseline procedure, we compute weights that are optimal
in an idealized setting where $U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$ independently of $X,Z$ and $\tilde \Gamma$ (again, this involves invoking the heuristic of treating $\tilde\Gamma$ as a nuisance parameter in (\ref{model}) rather than estimation error from a preliminary estimate).
In this idealized setting, $\hat\beta_A$ is then normally distributed with
variance $\sigma^2\sum_{i=1}^N\sum_{t=1}^TA_{it}^2=\sigma^2\|A\|_F^2$ (where
$\|\cdot\|_F$ denotes the Frobenius norm), and with bias ranging from
$-\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ to $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$.
Thus, if we choose worst-case MSE under i.i.d. normal errors as our criterion function for the weights, then the optimal weights
are obtained by minimizing  $\left( \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) \right)^2 + \sigma^2
\|A\|_F^2 $.
By substituting the formula for $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ from
(\ref{WorstCaseBias}), we obtain the following baseline choice of weights, indexed by
a tuning parameter ${b}$ that corresponds to~$\hat C/\sigma$.


\begin{definition}
\label{DefOptimalWeights}
For ${b}>0$, define the ``optimal'' $N\times T$ weight matrix by
\begin{align*}
  A^*_{{b}} &:= \operatorname*{argmin}_{A \in \mathbb{R}^{N \times T} } \, {b}^2 s_1(A)^2 + \|A\|_F^2
      \;\; \qquad\text{s.t.}\quad  \text{$\langle A,X \rangle_F = 1$ and $\langle A,Z_k\cdot \delta \rangle_F = 0$,}
\end{align*}
Here, the constraint $\langle A,Z_k\cdot \delta \rangle_F = 0$
is imposed for all $k \in \{1,\ldots,K\}$.
\end{definition}

Heuristically, we expect that a
good choice of ${b}$ will correspond to $\hat C/\sigma$ such that the
bound $\hat C$ on the
nuclear norm holds with high probability.
Conveniently, our nuclear norm bound in the exact factor model in Section
\ref{sec:linear_factor_implementation} scales with the standard deviation $\sigma$ in the
homoskedastic case, which gives us a simple and feasible choice of the tuning
parameter ${b}$.


We emphasize again that while the definition of $A^*_{b}$ is
motivated by the idealized setting $U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$, we do {\it not} assume that the error terms $U_{it}$ satisfy this strong
assumption.
Choosing $A=A^*_{b}$
to construct the estimator  $\hat\beta_A $ and the confidence intervals \eqref{general_A_CI_eq} under more general error distributions just means that
the resulting estimates and confidence intervals will not be optimal (in finite samples), but we will nevertheless show them to be consistent and valid, respectively.


\begin{remark}\label{other_criteria_remark}
  While we have used MSE to motivate our baseline choice of weights $A^*_{b}$, one
  could use other criteria corresponding to different weights on bias and
  variance.  For example, optimizing CI length when $\hat C/\sigma=b$
  would give the criterion ${b} s_1(A)+ z_{1-\alpha}\|A\|_F$.  If $\beta$ gives
  the net welfare gain of an all-or-nothing policy change, then one can target
  minimax welfare regret as in \citet{ishihara_evidence_2021} and
  \citet{yata_optimal_2021}.
  In our Monte Carlo simulations however, we find that the exact choice of
  criterion has little effect on performance.
\end{remark}


\subsection{Practical implementation}\label{general_practical_implementation_section}


The definition of $A^*_{b}$ is a convex optimization problem
that can easily be solved numerically for any given input $X$, $Z$,
${b}$.
Using results from
\citet{armstrong_bias-aware_2020}, it follows that $A^*_{b}$
can also be computed using the residuals of a nuclear norm regularized
regression of $X$ on $Z_1,\ldots,Z_K$ and a matrix of individual effects.
When there are no additional covariates~$Z$, this nuclear
norm regularized regression simplifies further: it can be solved by computing
the singular value decomposition of $X$, and then performing soft thresholding
on the singular values.  The resulting weights $A^*_{b}$ obtained from
the residuals of this regression replace the largest singular values of $X$
with a constant.
We provide details in Appendix~\ref{computational_details_appendix}.

In addition to giving alternative methods for computing the weights $A^*_{b}$,
these results provide some intuition for these weights.
The
residuals from this nuclear norm regularized regression of $X$ on $Z_1,\ldots,
Z_K$ and the individual effects ``partial out'' potential correlation of $X$ with the
estimation error $\widetilde\Gamma$, similar to the estimator of
\citet{robinson_root-n-consistent_1988} in the partially linear model.
When there are no additional covariates $Z$, this amounts to removing the largest singular values of $X$ and replacing them with a constant.















To summarize, we can compute an estimator $\hat\beta_{A}$ using Definition
\ref{DefAugLinEst} using any matrix of weights $A$.  We can also compute a CI $\left\{ \hat\beta_A \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2}
    \widehat{\operatorname{se}} \right]   \right\}$
as in (\ref{general_A_CI_eq}),
once we have a standard error $\widehat{\operatorname{se}}$ and an upper bound $\hat C$ for the
nuclear norm of the error in the initial estimate of $\Gamma$.  Definition \ref{DefOptimalWeights}
gives us a heuristic for computing a reasonable choice of the matrix $A$, once
we have an initial choice of ${b}$ for the ratio $\hat C/\sigma$ of the nuclear norm
bound to variance of $U_{it}$.

Thus, to apply our approach, we need an initial choice
${b}$ to compute the weights $A^*_{b}$
using Definition \ref{DefOptimalWeights}.
We also need a robust upper bound $\hat C$ such that the bound
(\ref{NNcondition}) holds with high probability.
Finally, we need a robust standard error $\widehat{\operatorname{se}}$.
Our CI then takes the form in (\ref{general_A_CI_eq}) with $A=A^*_{b}$
and the given bound $\hat C$ and standard error $\widehat{\operatorname{se}}$.
In Section \ref{sec:linear_factor_implementation}, we give details of these choices, as well as how to compute
the initial estimate of $\Gamma$.












\section{Implementation}\label{sec:linear_factor_implementation}


In this section, we describe the implementation of our approach.
Our approach relies on bounds for the nuclear norm of the initial estimate of $\Gamma$, derived formally in Section \ref{asymptotic_section}.
As explained in Section \ref{asymptotic_section}, a tighter CI can be derived using a more nuanced argument that bounds the difference between $\hat\Gamma$ and $\Gamma+P_{\lambda} U$, where $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$, and $M^+$ denotes the Moore–Penrose inverse of a matrix $M$.
In particular, we show that
the bound
$\hat C\approx 2Rs_1(U)$
can be used.
The bound $R$ on the number of factors must be specified by the researcher, similar to other methods in this literature \citep[e.g.][]{bai2009panel}.
Furthermore, the weights $A^*_{b}$ are designed to be optimal when
$U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$, which leads to the approximation
$s_1(U)/\sigma\approx \sqrt{N}+\sqrt{T}$ \citep{geman1980limit}.  We therefore use
${b}={b^*}:=2R(\sqrt{N}+\sqrt{T})$ as our default choice to calibrate
$\hat C/\sigma$ when computing the weights in Definition \ref{DefOptimalWeights}.
We then use an upper bound $\hat C$ that is valid under heteroskedasticity when
computing $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{b^*}})$ in the construction of the CI.

Our initial estimate is formed in two steps.  First, we form the least squares estimate $\hat\Gamma_{\rm LS}$.  We then apply our debiasing approach to get an estimate of the coefficients $\beta$ and $\delta$ and form an $\hat\Gamma_{\rm pre}$ by applying least squares to estimate $\Gamma$, with $\beta$ and $\delta$ fixed at this initial debiased estimate.  This estimate $\hat\Gamma_{\rm pre}$ is then used as the initial estimate in our procedure.
Thus, our procedure involves applying our debiasing approach twice.  This appears to be necessary to get an initial estimate $\hat\Gamma_{\rm pre}$ the best possible nuclear norm bounds on the estimation error.



Below we provide the
details of our implementation algorithm.\footnote{Implementation of this algorithm in R is also available at \url{https://github.com/chenweihsiang/PanelIFE/tree/main}. We thank Chen-Wei Hsiang for his excellent assistance in preparing this R package.}\textbf{}

\begin{algorithm}[Implementation for the factor model]\mbox{}
  \label{alg: factor model}
  \begin{description}
  \item[Input] Data $Y,X,Z$ and $R$ pre-specified by the user, along with tuning
    parameter $\varepsilon$.

  \item[Output] Estimator and CI for $\beta$.

  \end{description}

  \begin{enumerate}
  \item Compute the least squares (LS) estimator
    \begin{align}
      \left(\hat \beta_{\rm LS}, \hat \delta_{\rm LS}, \hat \Gamma_{\rm LS}\right) = \operatorname*{argmin}_{\left\{\beta \in \mathbb R, \delta \in \mathbb R^K, G \in \mathbb{R}^{N \times T} \, : \, {\rm rank}(G) \leq R \right\}}
      \sum_{i=1}^N \sum_{t=1}^T \left( Y_{it} - X_{it} \beta - Z_{it}' \delta - G_{it} \right)^2 .
      \label{eq: LS def}
    \end{align}

  \item Compute $\widetilde Y_{\rm pre} = Y - \hat \Gamma_{\rm LS}$
    and let ${b^*}=2R(\sqrt{N}+\sqrt{T})$.
    Let
    \begin{align*}
      \hat\beta_{\rm pre} = \langle A^*_{{b^*}}, \widetilde Y_{\rm pre} \rangle_F.
    \end{align*}
    Construct $\hat\delta_{\rm pre}$ with the $j$-th element $\hat\delta _{{\rm
  pre},j}$ computed in the same way as $\hat\beta_{\rm pre}$, but with $X$ and $Z_j$
switched.



  \item Compute $\hat \Gamma_{\rm pre}$ as
  \begin{align*}
    \hat \Gamma_{\rm pre} = \operatorname*{argmin}_{\left\{ G \in \mathbb{R}^{N \times T} \, : \, {\rm rank}(G) \leq R \right\}}
    \sum_{i=1}^N \sum_{t=1}^T \left( Y_{it} - X_{it} \hat \beta_{\rm pre} - Z_{it}' \hat \delta_{\rm pre} - G_{it} \right)^2 .
  \end{align*}
  The solution $\hat \Gamma_{\rm pre}$ to this least squares problem is simply
  given by the leading $R$ principal components of the residuals $Y_{it} -
  X_{it} \hat \beta_{\rm pre} - Z_{it}' \hat \delta_{\rm pre}$.
  Compute $\widetilde Y = Y - \hat \Gamma_{\rm pre}$.

  \item Compute the final estimate
    \begin{align*}
      \hat\beta = \hat\beta_{A^*_{{b^*}}} = \langle A^*_{{b^*}}, \widetilde Y \rangle_F.
    \end{align*}
    To compute the CI, let $\hat C = (2 +
    \varepsilon) R s_1(\hat U_{\rm pre})$ and $\widehat{\operatorname{se}}^2=\sum_{i=1}^N\sum_{t=1}^T A_{b^*,it}^{*2}\hat U_{{\rm pre},it}^2$, where
    \begin{align*}
      \hat U_{\rm pre} = Y - X \hat \beta_{\rm pre} -Z \cdot \hat \delta_{\rm pre} - \hat \Gamma_{\rm pre}.
    \end{align*}
    Compute the CI
    \begin{align}
      \label{eq: bias aware CI}
      \hat\beta_{A^*_{{b^*}}} \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{{b^*}}}) + z_{1-\alpha/2}\widehat{\operatorname{se}}\right]
    \end{align}
    where $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{{b^*}}})=\hat C s_1(A^*_{{b^*}})$.

  \end{enumerate}
\end{algorithm}








\begin{remark}[Behavior of estimator under strong factors]
  \label{remark: strong factors CS}
  In Section~\ref{ssec: semi strong}, we show that, in the absence of weak
  factors, $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{{b^*}}})$ becomes negligible
  relative to $\widehat{\operatorname{se}}$ in the construction of the CI in \eqref{eq: bias
    aware CI}.  Thus, the CI
  \begin{align}\label{eq:non_bias_aware_ci}
    \hat\beta_{A^*_{{b^*}}}
    \pm z_{1-\alpha/2}\widehat{\operatorname{se}},
  \end{align}
  which uses our bias-corrected estimator but
  ignores bias when computing the critical value, will have correct asymptotic
  coverage in the strong factor case.
  In our Monte Carlos provided in Section~\ref{ssec: MC non-robust coverage},
  we find that this CI is (i) comparable to alternative non-robust CIs in terms
  of the length and (ii) despite its non-robustness, substantially less
  size-distorted if there is a weak factor(s) since it is based on the debiased
  estimator.
  In settings where the bias-aware CI (\ref{eq: bias aware CI}) is too wide to
  yield precise inference, we recommend reporting the CI
  (\ref{eq:non_bias_aware_ci}) alongside the bias-aware CI (\ref{eq: bias aware
    CI}) as a compromise between ignoring weak factors and a fully robust
  approach.

\end{remark}

\begin{remark}[Choice of $R$ and $\varepsilon$]
  The  quantity $\varepsilon$ is used in the bound $\hat C = (2 +
  \varepsilon) R s_1(\hat U_{\rm pre})$ on $\|\widetilde \Gamma\|_*$ needed
  to compute the CI in the final step. While $\varepsilon>0$ is necessary for theoretical guarantees,
   in our Monte Carlos, we find that  we get good coverage when choosing
   $\varepsilon=0$.

  In contrast, the choice of $R$ has a substantive effect on the CI, both
  through the bound $\hat C = (2 + \varepsilon) R s_1(\hat U_{\rm pre})$ and
  through the point estimate.
  Since the number of weak factors cannot be determined from the data, the
  researcher must specify an a priori bound on the total number of factors $R$.
  Nonetheless, the data can be informative about the number of strong or semi-strong
  factors, which provides a lower bound for $R$.  We recommend forming an
  estimate $\hat R_s$ of the number of strong factors using one of the standard
  methods (e.g., \citealp{BaiNg2002,Onatski2010,AhnHorenstein2013}) and using
  this as a starting point for examining the sensitivity of the results to the
  choice of $R$.  For example, by taking $R = \hat R_s + 1$, the researcher
  allows for the potential presence of an additional weak factor (or $R - \hat
  R_s$ weak factors in general for a bigger $R$).
\end{remark}


\begin{remark}[Lindeberg condition]\label{lindeberg_remark}
  The asymptotic validity of the CI depends on asymptotic normality of the
  stochastic term $\langle A, U \rangle_F$
  where $A=A^*_{{b^*}}$ is a non-random matrix of weights.  This, in
  turn, depends on a Lindeberg condition on the weights $A$.
  To ensure that this holds, we can modify our optimization procedure for
  computing the weights $A=A^*_{{b^*}}$ by imposing a bound on the
  Lindeberg weights
  \begin{align}\label{lindeberg_weights_def_eq}
    \operatorname{Lind}(A)=\frac{\max_{1\le i\le N, 1\le t\le T} A_{it}^2}{\sum_{i=1}^N\sum_{t=1}^T A_{it}^2}.
  \end{align}
  A similar approach to showing asymptotic validity is taken in
  \citet{javanmard_confidence_2014} in a different setting.

  To make this approach practical, we need guidance on what makes
  $\operatorname{Lind}(A)$ ``small enough to use the central limit theorem'' in
  a given sample size.  A formal answer to this question is elusive, due to the
  difficulty of obtaining finite sample bounds on approximation error in the
  central limit theorem that are practically useful.  As a heuristic, we can use
  comparisons to other settings where the central limit theorem is used.  For
  example, the sample mean $\bar W=\frac{1}{n}\sum_{i=1}^n W_i$
  with $n$ observations corresponds to an estimator with Lindeberg constant
  $(1/n)^2/[n\cdot (1/n)^2]=1/n$.  If we are comfortable using the normal
  approximation in such a setting with, say, $n=50$, then we can impose a bound
  $\operatorname{Lind}(A)\le 1/50$.
  \citet{noack_bias-aware_2024} provide some discussion of these issues in a
  related setting involving inference in fuzzy regression discontinuity.

  In our Monte Carlos, we find that $\operatorname{Lind}(A)$ is very small for
  the weights used in Algorithm \ref{alg: factor model} once $N$ and $T$ are
  larger than, say, 20.  Thus, imposing a bound on these weights does not appear
  to be necessary in practice in the data generating processes we have examined.

\end{remark}


\begin{remark}[Standard error]
  The standard error $\widehat{\operatorname{se}}^2=\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{{\rm pre},it}^2$ assumes that $U_{it}$ is
  uncorrelated across $i$ and $t$, but allows for heteroskedasticity.  Such an
  assumption will be reasonable if $\Gamma_{it}$ captures all of the dependence
  in errors for the outcome.  However, incorporating all dependence in
  $\Gamma_{it}$ may lead to an unnecessarily conservative choice of the upper bound $R$ on the number of factors, leading to a wider CI.  To avoid such
  conservative bounds on $\Gamma$, one can incorporate any dependence that is
  not directly correlated with $X_{it}$ into the error term $U_{it}$, and allow
  for such dependence when constructing the standard error. For example, to allow for (arbitrary) time dependence of $U_{it}$ (while maintaining uncorrelatedness across $i$), one could simply use clustered standard errors (see, e.g., \citealp{arellano1987computing,hansen2007asymptotic})
  \begin{align*}
    \widehat{\operatorname{se}}^2=\sum_{i=1}^N \left(\sum_{t=1}^T A_{it}\hat U_{{\rm pre},it}\right)^2.
  \end{align*}
\end{remark}





\section{Asymptotic results}\label{asymptotic_section}








This section gives formal asymptotic results for the estimators and CIs given in
Sections \ref{sec:MainIdea} and \ref{sec:linear_factor_implementation}.
We consider the following decomposition of our regression model:
\begin{align}
  \label{eq: augmented model}
  \widetilde Y: = Y - \hat \Gamma_{\rm pre} = X \beta + Z\cdot \delta + \Gamma-\hat\Gamma_{\rm pre} + U = X \beta + Z\cdot \delta + \widetilde \Gamma + \widetilde U.
\end{align}
Here, $\widetilde\Gamma$ and $\widetilde U$ an be any $N\times T$ matrices
chosen compatibly so that $\widetilde \Gamma + \widetilde U =
\Gamma-\hat\Gamma_{\rm pre} + U$.
While our discussion so far has focused on the case where $\widetilde \Gamma =
\Gamma - \hat \Gamma_{\rm pre}$ and $\widetilde U = U$, it turns out that
allowing for other choices of $\widetilde\Gamma$ and $\widetilde U$ allows for
an improvement in the width of our CI.


To formally state asymptotic results that allow for weak factors and an unknown
error distribution, we introduce some additional notation.
We consider uniform-in-the-underlying distribution asymptotics over a set
$\mathcal{P}$ of distributions $P$ for $\Gamma$ and $X,Z_1,\ldots,Z_K,U$ and a
set $\Theta$ of parameters
$\theta=(\beta,\delta')'$.
While we treat $\Gamma,X,Z_1,\ldots,Z_k$ as random variables determined by the
unknown probability distribution $P$ for notational purposes, we note that a
fixed design setting in which $\Gamma,X,Z_1,\ldots,Z_k$ are non-random
(sequences of) matrices can be incorporated by considering a set $\mathcal{P}$
that places a probability one mass on a given value of
$\Gamma,X,Z_1,\ldots,Z_k$.
We
use $\mathbb{P}_{P,\theta}$ to denote probability under the given distribution $P$ and
parameters $\theta$.  Formally, we consider large $N$, large $T$ asymptotics in
which $N=N_n\to\infty$ and $T=T_n\to\infty$, and we consider sequences of
distributions $\mathcal{P}=\mathcal{P}_n$ and parameter spaces
$\Theta=\Theta_n$.
Asymptotic statements are then taken in the
sequence $n$.  However, we suppress the dependence on an index sequence $n$ in
order to save on notation.
For a sequence of vectors or matrices $A_{N,T}=A_{N,T}(\theta,P)$ of fixed
dimension (which may depend on $\theta,P$),
we use the notation $A_{N,T}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T})$ when, for every
$\varepsilon>0$, there exists $C_\varepsilon$ such that
\begin{align*}
  \limsup \sup_{P\in\mathcal{P},\theta\in\Theta} \mathbb{P}_{P,\theta}\left( r_{N,T}^{-1} \| A_{N,T} \|\ge C_\varepsilon \right)\le \varepsilon,
\end{align*}
and we use the notation $A_{N,T}=o_{\Theta,\mathcal{P}}(r_{N,T})$ when, for
every $\varepsilon>0$, we have
\begin{align*}
  \limsup \sup_{P\in\mathcal{P},\theta\in\Theta} \mathbb{P}_{P,\theta}\left( r_{N,T}^{-1} \| A_{N,T} \|\ge \varepsilon \right) \to 0.
\end{align*}
We use the notation $A_{N,T}\asymp_{\Theta,\mathcal{P}}r_{N,T}$ when
$A_{N,T}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T})$ and
$A_{N,T}^{-1}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T}^{-1})$.
We use the notation
$A_{N,T} \underset{\Theta,\mathcal{P}}{\overset{d}{\to}} \mathcal{L}$
to denote the statement
\begin{align*}
  \limsup
  \sup_{\theta\in\Theta,P\in\mathcal{P}}
  \Big|  \mathbb{P}_{\theta,P}\left( A_{N,T} \le t \right) - F_{\mathcal{L}}(t) \Big| \to 1
  \text{ for all }t
\end{align*}
where $F_{\mathcal{L}}$ denotes the cdf of the probability law $\mathcal{L}$.


\subsection{General results}

We first show asymptotic validity of
the CI (\ref{general_A_CI_eq}) under the following high level assumption imposed on the augmented model \eqref{eq: augmented model} and the weights $A_{it}$.

\begin{assumption}\label{ass: augmented high level}
  \mbox{}
  \begin{enumerate}[(i)]
  \item\label{item: augmented high level C hat bound} $\inf_{\theta\in\Theta,P\in\mathcal{P}}\mathbb{P}_{\theta,P}\left( \|\widetilde \Gamma\|_* \le \hat C \right)\to 1$;

  \item\label{item: augmented high level CLT} $\frac{\langle A, \widetilde U
      \rangle_F}{\widehat{\operatorname{se}}}
    \underset{\Theta,\mathcal{P}}{\overset{d}{\to}}
    N(0,1)$.
  \end{enumerate}
\end{assumption}

\begin{theorem}\label{new_high_level_validity_thm}
  Suppose that Assumption \ref{ass: augmented high level} holds.
  Then
  \begin{align*}
    \liminf \inf_{\theta\in\Theta,P\in\mathcal{P}} \mathbb{P}_{\theta,P}\left( \beta \in \left\{ \hat\beta_A \pm \left[  \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2}\widehat{\operatorname{se}}  \right] \right\} \right) \ge 1-\alpha.
  \end{align*}
\end{theorem}



\subsection{Primitive conditions}

We now apply these results to the initial estimate and bound given in Section
\ref{sec:linear_factor_implementation}, under the assumption of a linear factor
model for $\Gamma$.
We allow for a side condition on the
Lindeberg weights
$\operatorname{Lind}(A)$
defined in (\ref{lindeberg_weights_def_eq}), as described in Remark
\ref{lindeberg_remark}.
Let $A^*_{{b}, c}$ be defined in the same way as $A^*_{{b}}$, with the modification that
we impose the constraint $\operatorname{Lind}(A)\le c$:
\begin{align}\label{weight_optimization_lindeberg}
  &\min_{A} \left\| A \right\|_F^2  +  {b}^2 s_1(A)^2,  \nonumber  \\
  &\text{s.t.}
    \quad \operatorname{Lind}(A)\le c,
    \quad \langle A,X \rangle_F = 1,
    \quad \langle A,Z_k \rangle_F = 0 \text{ for } k=1,\ldots,K.
\end{align}
In particular, the
weights used in Algorithm \ref{alg: factor model} are given by
$A^*_{{b^*},\infty}=A^*_{{b^*}}$, and the weights
$A^*_{{b^*},c}$ with $c<\infty$ correspond to the modification described in
Remark \ref{lindeberg_remark}.
In our asymptotic theory we will require
$c=c_{NT}$ to converge to zero as $N,T \rightarrow \infty$.




We impose the following conditions.

\begin{assumption}[Factor Model]
  \label{ass: factor model}
  Suppose that $\text{rank}\left(\Gamma\right) \le R$, i.e., $\Gamma = \lambda f'$ for some $N \times R$ matrix $\lambda$ and some $T \times R$ matrix $f$,
  with probability one for all $P\in\mathcal{P}$
  and the following conditions hold:
  \leavevmode
  \begin{enumerate}[(i)]
    \item \label{item: NC}
     Write $W$ for $X,Z_1,\ldots, Z_K$ and $W\cdot \gamma=X\beta +
     \sum_{k=1}^KZ_k\delta_k$ where $\gamma=(\beta,\delta')'$.  We assume that
     there exists $\underline s^2 > 0$ such that
    \begin{align*}
      \min_{\gamma \in \mathbb R^{K+1}: \left\Vert \gamma\right\Vert = 1}\frac{1}{NT} \sum_{r = 2 R + 1}^{\min\{N,T\}} s_r^2 (W \cdot \gamma) \ge \underline s^2
    \end{align*}
    with probability approaching 1 uniformly over $P \in \mathcal P$;


    \item \label{item: SN} $s_1(X) = \mathcal{O}_{\Theta,\mathcal{P}} \left(\sqrt{NT}\right)$, $s_1(Z_k) = \mathcal{O}_{\Theta,\mathcal{P}} \left(\sqrt{NT}\right)$ for $k \in \{1, \ldots, K\}$, and $s_1(U) \asymp_{\Theta,\mathcal{P}} \max \{\sqrt{N},\sqrt{T}\} $;

          \item \label{item: EX} $\langle X, U \rangle_F  =
            \mathcal{O}_{\Theta,\mathcal{P}} (\sqrt{NT})$ and $\langle Z_k, U
              \rangle_F = \mathcal{O}_{\Theta,\mathcal{P}} (\sqrt{NT})$ for $k \in \{1, \ldots, K\}$;


    \item \label{item: high level singular values of U} $\left(s_1(U) - s_{r}(U)\right)/s_1(U) = o_{\Theta,\mathcal{P}}(1)$ for any fixed positive integer $r$;



    \item \label{A_U_inner_product_bound} For any sequence of matrices $A=A_{N,T}(X,Z)$ that is a function of
    $X,Z_1,\ldots,Z_k$, we have $\langle A, U \rangle_F =
    \mathcal{O}_{\Theta,\mathcal{P}}(\|A\|_F)$.

  \end{enumerate}
\end{assumption}


Assumption \ref{ass: factor model}\eqref{item: NC}
is a generalized non-collinearity condition, which
requires that there is enough variation in the regressors after concentrating out $2R$ arbitrary factors.
It is closely related to Assumption~A of \cite{bai2009panel}, but our version here avoids mentioning the unobserved factor loadings. The same generalized non-collinearity assumption is imposed in \cite{MoonWeidner2015}.
The assumption would be violated if
some linear combination $W \cdot \gamma$
of the covariates were to have rank smaller or equal
to $2R$. In particular,
 ``low-rank regressors'' are ruled out by this condition.  Intuitively, Assumption \ref{ass: factor model}(\ref{item: NC}) holds provided that $X_{it}$ and $Z_{it}$ have non-collinear idiosyncratic components. This intuition is formalized in Appendix~\ref{appendix:VerifyAss2i}, where we also show that, when $X$ is the only regressor in the model (i.e.,\ $K=0$), then Assumption~\ref{ass: X decomposition}\eqref{V_assump} and \eqref{H_assump} below already guarantee that
 Assumption \ref{ass: factor model}(\ref{item: NC})  holds.






Assumption \ref{ass: factor model}(\ref{item: SN}) places mild bounds on $X$ and
$Z_k$. For example,
if  the second moments of $X_{it}$ are (uniformly) bounded, then $
   \mathbb{E}[s_{1}(X)^2] \leq \mathbb{E}[\| X \|_F^2] = \sum_{i=1}^N \sum_{t=1}^T \mathbb{E}[X_{it}^2] = \mathcal O_{\Theta,\mathcal{P}}(NT),
   $
   which, by Markov's inequality, implies
   $s_{1}(X) = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$.
In addition Assumption \ref{ass: factor model}(\ref{item: SN}) also places a rate restriction on $s_1(U)$
that will hold as long as $U_{it}$
does not exhibit too much dependence over $i$ and $t$.
This rate for $s_1(U)$ is closely related to
Assumption \ref{ass: factor model}(\ref{item: high level singular values of U}), which is discussed below.
Assumption \ref{ass: factor model}(\ref{item: EX}) again holds as long as $U_{it}$ does not
exhibit too much dependence over $i$ and $t$, and is uncorrelated with $X_{it}$ and $Z_{it}$.
Finally,
Assumption \ref{ass: factor model}(\ref{A_U_inner_product_bound}) holds as long as
$U$ is mean zero given $X$ and $Z$ and satisfies bounds on dependence and second
moments.

Assumption \ref{ass: factor model}(\ref{item: high level singular values of U})  is a high level assumption on the first few singular values of $U$ (note that $r$
is fixed as $N$ and $T$ converge to infinity).
The singular values of $U$ are
the square roots of the eigenvalues of $UU'$.
The random matrix theory literature shows that,
if $U$ is an appropriate noise matrix,  the
largest few eigenvalues of  $UU'$ converge to the Tracy-Widom law, after appropriate rescaling:
if $N$ and $T$ grow at the same rate, then
each of the largest eigenvalues of $UU'$
grows at rate $N$, while the gaps between them grow
at rate $N^{1/3}$.
\cite{Johnstone2001} establish the  Tracy-Widom law
for the largest  eigenvalues of $UU'$, for
the case of i.i.d.\ normal
error $U_{it}$. The subsequent literature has
shown the universality of this result for
more general error distributions,
see e.g.\ \cite{Soshnikov2002}, \cite{pillai2012edge}
and \cite{yang2019edge}.


We also place conditions on the matrix $X$ requiring that there is sufficient
variation after controlling for individual effects and the additional covariates
$Z$.
The constant $c=c_{N,T}$ in the following
assumption is the one that appears
in our construction of $A^*_{b,c}$.


\begin{assumption}\label{ass: X decomposition}
  For all $P\in\mathcal{P}$, there exists uniformly bounded $\pi=\pi_P$ and random matrices $H$
  and $V$ such that $X=Z\cdot \pi + H + V$ and the following conditions hold:
  \begin{enumerate}[(i)]
  \item \label{V_assump} $\left\Vert V\right\Vert_F \asymp_{\Theta,\mathcal{P}} \sqrt{NT}$, $s_1 (V) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$;

  \item \label{H_assump} $\left\Vert H\right\Vert_F = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$ and $\langle H,V
    \rangle_F = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$;

  \item \label{Z_V_inner_product_assump} $\|Z_k\|_F=\mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$ and
    $\langle Z_k, V \rangle_F = \mathcal O_{\Theta,\mathcal{P}} (\sqrt{N T})$ for $k \in \{1, \ldots, K\}$;

  \item \label{Z_prime_Z_assump} $\left(\bf{ Z{}' Z}\right)^{-1} = \mathcal O_{\Theta,\mathcal{P}}\left(\frac{1}{NT}\right)$ where
    ${\bf Z} = [{\rm vec}(Z_1), \ldots, {\rm vec}(Z_K)]$;

  \item \label{max_V_Z_assump} $\max_{i,t} V_{it}^2 = o_{\Theta,\mathcal{P}} (NT c_{N,T})$ and $\max_{i,t} Z_{k,it}^2 = o_{\Theta,\mathcal{P}} \left( (NT)^2 c_{N,T}\right)$ for $k \in \{1,\ldots,K\}$.
  \end{enumerate}

\end{assumption}


Assumption \ref{ass: X decomposition} uses a decomposition of $X_{it}$ that
depends on an individual effect $H_{it}$ and a random variable $V_{it}$ that is
approximately independent and uncorrelated with $Z_{1,it}\ldots,Z_{k,it}$ as
well as being approximately uncorrelated with the individual effect $H_{it}$.
Importantly, the individual effect $H_{it}$ can be arbitrarily correlated with
$\Gamma_{it}$ and with the variables $Z_{k,it}$.  Note also that we do not place
any assumptions on the rank or nuclear norm of the matrix $H_{it}$.


Part (\ref{max_V_Z_assump}) holds under a tail bound on $V_{it}$ and $Z_{k,it}$.
For example, if $V_{it}$ are (uniformly) sub-Gaussian then $\max_{i,t} V_{it}^2 = \mathcal O_{\Theta,\mathcal{P}} ( \log
(N + T) )$, and the condition $\max_{i,t} V_{it}^2 = o_{\Theta,\mathcal{P}} (NT c_{N,T})$ is
satisfied provided that $N T c_{N,T} / \log (N + T) \rightarrow \infty$.
The only other requirement on $c_{N,T}$ is the requirement that
$c_{N,T} \max\{N,T\}\to 0$ in Theorem
\ref{new_validity_theorem} below.  Thus, our results allow
for a range of choices of $c_{N,T}$.





Define $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$ where $M^+$ denotes the Moore–Penrose inverse of a matrix $M$.


\begin{theorem}\label{new_nuclear_bound_error_thm}
  Let $\hat \Gamma_{\rm pre}$  be defined in Algorithm \ref{alg: factor model},
  with the modification described in Remark \ref{lindeberg_remark}.
  Suppose that Assumption \ref{ass: factor model} holds, and that
  Assumption \ref{ass: X decomposition} holds as stated and with $Z_k$ and $X$
  interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$. Then, for any $\varepsilon > 0$, Assumption \ref{ass: augmented high level}\eqref{item: augmented high level C hat bound} holds with
  \begin{enumerate}[(i)]
    \item $\widetilde \Gamma = \Gamma - \hat \Gamma_{\rm pre}$ and $\hat C = 3 R s_1 (\hat U_{\rm pre}) (1 + \varepsilon) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$; \label{item: 3 R bound}

    \item $\widetilde \Gamma = \Gamma + P_\lambda U - \hat \Gamma_{\rm pre}$ and $\hat C = 2 R s_1 (\hat U_{\rm pre}) (1 + \varepsilon) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$. \label{item: 2 R bound}
  \end{enumerate}
\end{theorem}

Theorem \ref{new_nuclear_bound_error_thm} is the main novel technical result that allows us to construct a feasible CI.  It provides an explicit bound on the nuclear norm error of our initial estimate.  As we show in the proof of Theorem \ref{new_validity_theorem} below, the term $\langle A, P_\lambda U\rangle_F$ is asymptotically negligible under our assumptions.  Thus, redefining the target parameter to be $\Gamma + P_\lambda U$ instead of $\Gamma$ and using the bound in part (ii) of the theorem does not affect the construction of the CI.
This leads to a shorter CI using the bound in part (ii) compared to using the bound in part (i).  For this reason, we use the bound in part (ii) in the implementation described in Section \ref{sec:linear_factor_implementation} and in our formal coverage results below.


We now turn to the rate of convergence of the debiased estimator and the coverage of the CI.  The proofs of these theorems use the nuclear norm bounds in Theorem \ref{new_nuclear_bound_error_thm}.





\begin{theorem}\label{new_beta_rates_thm}
    Let $\hat \beta=\hat\beta_{A^{*}_{{b^*},c}}$ be defined in Algorithm \ref{alg: factor model},
    with the modification described in Remark \ref{lindeberg_remark}.
    Suppose that Assumption \ref{ass: factor model} holds, and that
    Assumption \ref{ass: X decomposition} holds as stated and with $Z_k$ and $X$
    interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$.   Then
    \begin{align*}
      \hat\beta-\beta=\mathcal{O}_{\Theta,\mathcal{P}}(1/\min\{N,T\}).
    \end{align*}
\end{theorem}


To obtain primitive conditions for a central limit theorem and asymptotic
validity of the confidence interval, we impose that the errors are independent,
but not necessarily identically distributed, conditional on $X$, $Z$ and $\Gamma$.

\begin{assumption}\label{U_conditional_moment_bound_assump}
  There exist constants $\underline \sigma>0$ and $\eta>0$ such that, for all
  $P\in\mathcal{P}$, $U_{it}$ is independent over $i,t$ conditional on
  $W,\Gamma$ and, for all $i,t$,
  \begin{align*}
    \mathbb{E}_P [U_{it}|W,\Gamma]=0,
    \quad \mathbb{E}_P [U_{it}^2|W,\Gamma]>\underline\sigma^2,
    \quad \mathbb{E}_P [U_{it}^{4}|W,\Gamma]<1/\eta.
  \end{align*}
\end{assumption}

\begin{theorem}\label{new_validity_theorem}
    Let $\hat \beta=\hat\beta_{A^{*}_{{b^*},c}}$ and $\hat C = 2Rs_1(\hat
  U_{\rm pre})(1+\varepsilon)$ be defined in Algorithm \ref{alg: factor model},
  with the modification described in Remark \ref{lindeberg_remark} for
  $c=c_{N,T}$ with $c_{N,T} \max \{N,T\}\to 0$.
  Suppose that Assumptions \ref{ass: factor model}(\ref{item: NC})-(\ref{item: high level singular values of U}) hold, and that
  Assumption \ref{ass: X decomposition} holds as stated and with $Z_k$ and $X$
  interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$,
  and that Assumption \ref{U_conditional_moment_bound_assump} holds.
  Let
  $\widehat{\operatorname{se}}^2=\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{it}^2$
  where $A=A^{*}_{{b^*},c}$ and $\hat U_{it}$ is the residual from the least
  squares estimator.
  Then
  \begin{align*}
    \hat\beta-\beta=\mathcal{O}_{\Theta,\mathcal{P}}(1/\min\{N,T\})
  \end{align*}
  and
  \begin{align*}
    \liminf \inf_{\theta\in\Theta,P\in\mathcal{P}} \mathbb{P}_{\theta,P}\left( \beta \in \left\{ \hat\beta \pm \left[  \overline{\operatorname{bias}}_{\hat C}(\hat\beta) + z_{1-\alpha/2}\widehat{\operatorname{se}}  \right] \right\} \right) \ge 1-\alpha.
  \end{align*}
\end{theorem}



\subsection{Strong factor case}
\label{ssec: semi strong}

Numerous studies on the estimation of panel regressions with unobserved factors assume that these factors are ``strong'' or ``semi-strong''. This assumption implies that the unobserved error structure, $\Gamma_{it} + U_{it}$ in model \eqref{model0}, viewed as an $N \times T$ matrix, contains an $R = \operatorname{rank}(\Gamma)$ factor component $\Gamma = \lambda f'$ with singular values that asymptotically diverge faster than the largest singular value, $s_1(U)$, of the idiosyncratic error part $U$. Specifically, as $N, T \rightarrow \infty$, the ratio $s_1(U) / s_R(\Gamma)$ approaches zero (at a certain rate) under the semi-strong (strong) factor assumptions. Both \cite{Pesaran2006estimation} and \cite{bai2009panel} impose conditions that imply strong factors in this sense, as do many subsequent papers.

A key motivation for the estimation approach in this paper is to avoid assuming strong factors, instead providing an inference method that remains uniformly valid regardless of factor strength. Nevertheless, it is natural to consider how our approach behaves when factors are, in fact, strong, if only to facilitate comparison with much of the existing literature. The following theorem, therefore, extends Theorem~\ref{new_nuclear_bound_error_thm} to accommodate the case of strong factors.

\begin{theorem}
   \label{th:SemiStrong}
     Let $\hat \Gamma_{\rm pre}$  be defined in Algorithm \ref{alg: factor model},
  with the modification described in Remark \ref{lindeberg_remark}. Suppose that the hypotheses of Theorem~\ref{new_validity_theorem} hold, and furthermore, assume that
   $s_1(U) / s_R(\Gamma)= o_{\Theta,\mathcal{P}} (1)$. Then
   Assumption \ref{ass: augmented high level}\eqref{item: augmented high level C hat bound} holds with
    $\widetilde \Gamma =  \Gamma + M_{\lambda} U P_{f}  + P_{\lambda} U M_{f}  - \hat \Gamma_{\rm pre}$
   and
   $\hat C =  o_{\Theta, \mathcal P} \left(   s_1(U)   \right)    = o_{\Theta, \mathcal P} (\max \{\sqrt{N},\sqrt{T}\}) $,
   where $P_f = f (f' f)^+ f'$, $M_\lambda = \mathbb I_N - P_\lambda$, and $M_f = \mathbb I_T - P_f$.
\end{theorem}


Theorem~\ref{th:SemiStrong} additionally requires $s_1(U) / s_R(\Gamma)= o_P(1)$, i.e., that the factors
are semi-strong or strong, and also that $R = \operatorname{rank}(\Gamma)$. This implies that
$\hat \Gamma_{\rm pre}$ converges to
$\Gamma + M_{\lambda} U P_{f}  + P_{\lambda} U M_{f}$ in second leading order
(see e.g.\ Lemma S.3 in the supplement to \citealp{MoonWeidner2015}).
Theorem~\ref{th:SemiStrong} states that with this change of target matrix, the nuclear norm bound
$\hat C$ is smaller than $ s_1(U) \asymp_{\Theta,\mathcal{P}} {(\max \{\sqrt{N},\sqrt{T}\})}.$ This implies that, when $N$ and $T$ grow at the same rate, the worst-case bias $\overline{\operatorname{bias}}_{\hat C}(\hat\beta)$ used in the construction of the CI in Theorem~\ref{new_validity_theorem} becomes negligible relative to the standard error $\widehat{\operatorname{se}}$. At the same time, the
extra term $\langle A, M_{\lambda} U P_{f}  + P_{\lambda} U M_{f} \rangle_F$ is asymptotically negligible by exactly the same arguments given in the proof of Theorem~\ref{new_validity_theorem} for the negligibility of the term $\langle A, P_\lambda U \rangle_F$. As a result, in the considered regime, the debiased estimator is asymptotically unbiased and the (non-bias) aware CI \eqref{eq:non_bias_aware_ci} is asymptotically valid.

\begin{remark}
  \label{rem: R_w bound}

  Theorem \ref{th:SemiStrong} shows that the bias term is asymptotically negligible when $N$ and $T$ grow at the same rate and all $R$ factors are strong.  This justifies the CI (\ref{eq:non_bias_aware_ci}) discussed in Remark \ref{remark: strong factors CS} in the strong factor setting.  More generally, in the case where $R_w$ factors are weak and $R_s=R-R_w$ factors are strong, we conjecture that Theorem \ref{th:SemiStrong} could be extended to show that the bias-aware CI (\ref{eq: bias aware CI}) is valid with $\hat C = 2 R_w s_1 (\hat U_{\rm pre}) (1 + \varepsilon)$.
 In other words, we conjecture that the worst-case bias of our estimator only depends on the number of weak factors.
\end{remark}






\subsection{Comparison to other results in the literature}\label{sec:comparison_to_other_results}

Our debiasing approach leads to the faster rate $\min \{N,T\}$ compared to the rate $\min \{\sqrt{N},\sqrt{T}\}$ for $\hat \beta_{\rm LS}$ (see, e.g., \citealp{MoonWeidner2015}).
While our results appear to be the first to demonstrate a $\min\{N,T\}$ rate of convergence under the conditions above, recent papers have proposed estimators that use additional structure to construct estimators that achieve the same or better rates.
\citet{chetverikov2022spectral} impose a factor structure on $X$, which corresponds to imposing a low-rank assumption on the matrix $H$ in our Assumption \ref{ass: X decomposition}.  They use this assumption to construct an estimator that, like ours, achieves a $\min\{N, T\}$ rate under weak factors.  \citet{zhu2019well} imposes homoskedastic and independent errors in addition to a factor structure on $X$, and shows that this allows for a faster $\sqrt{NT}$ rate of convergence, even under weak factors.

While robust to weak factors, our CI will be wider than a CI based on the strong factor asymptotics in \citet{bai2009panel}.  Ideally, one would like to form a CI that is \emph{adaptive} to the strength of factors.  Such a CI would be robust to weak factors, while being asymptotically equivalent to the CI in \citet{bai2009panel} when factors are strong.  However, as shown by \citet{zhu2019well}, such an adaptive CI cannot be obtained, even if one imposes homoskedastic errors and additional structure on the covariate matrix $X$.  Thus, while there may be some room for efficiency gains over our CI, one must allow for some increase in CI length relative to the CI in \citet{bai2009panel} in order to allow for weak factors.

As discussed in the introduction, our debiasing approach is analogous to the approach to debiasing the LASSO taken in \citet{javanmard_confidence_2014} and, more broadly, other papers in the debiased LASSO literature such as \citet{belloni_inference_2014}, \citet{van_de_geer_asymptotically_2014} and \citet{zhang_confidence_2014}.
Interestingly, this analogy extends to the rates of convergence in our asymptotic results.  The debiased lasso applies to a high dimensional regression model with $s$ nonzero coefficients and $n$ observations.  The resulting estimator has bias of order $s/n$, up to log terms, and variance $1/n$.  Note that $s$ is the dimension of the constraint set for the unknown parameter, while $n$ is the total number of observations.  In our setting, the debiased estimator has bias of order $\max\{N,T\}/(NT)$ and variance $1/(NT)$.  The set of matrices $\Gamma$ with rank at most $R$ has dimension of order $\max\{N,T\}$ so, just as with the debiased lasso, the bias term is of the same order of magnitude as the ratio of the dimension of the constraint set to the total number of observations.  In the debiased lasso setting, one can justify a CI that ignores bias by assuming that $s$ increases slowly enough relative to $n$ for the order $s/n$ bias term to be asymptotically negligible relative to the order $1/\sqrt{n}$ standard deviation term.  Unfortunately, this cannot occur in our setting even if $R=1$, since the bias term is of order $\max\{N,T\}/(NT)$ which is always of at least the same order of magnitude as the standard deviation $1/\sqrt{NT}$.  This necessitates our bias-aware approach.






\section{Numerical Evidence}
\label{sec: numerical}
\subsection{Simulation Study}
\label{ssec: MC}
We consider the following design:
\begin{align*}
  Y_{it} &= X_{it} \beta + \sum_{r=1}^R \kappa_r \lambda_{ir} f_{tr} + U_{it} ,\\
  X_{it} &= \sum_{r=1}^R \lambda_{ir} f_{tr} + V_{it},
\end{align*}
where $\kappa_r$ controls the strength of factor $f_{tr}$, and $R$ stands for the number of factors. In addition, $\lambda_i$, $f_t$, $U_{it}$ and $V_{it}$ are all mutually independent across both $i$, $t$, and $(i,t)$, and
\begin{align*}
  \lambda_i \sim N(0,I_R) \perp f_t \sim N(0,I_R) \perp  \begin{pmatrix} U_{it} \\ V_{it} \end{pmatrix} \sim N \left(\begin{pmatrix} 0 \\ 0 \end{pmatrix}, \begin{pmatrix} \sigma_U^2 & 0 \\ 0 & \sigma_V^2 \end{pmatrix}\right).
\end{align*}

In the designs considered below, we fix $(\beta, \sigma_U^2, \sigma_V^2) = (0,1,1)$ and vary $N$, $T$, the number of factors $R$, and their strengths controlled by $\kappa_r$. The number of simulations in all of the considered designs is 5000.
As before, we are interested in estimation of and inference on $\beta$.

In Tables \ref{tab: N100 R1}-\ref{tab: N100 T50 R2 inference}, we report the bias, standard deviation, and rmse for the benchmark LS estimator of \citet{bai2009panel} and for the proposed debiased estimator in various designs with 1 and 2 factors.\footnote{Note that the CCE estimator of \cite{Pesaran2006estimation} would not work in these designs, regardless of whether the factors are strong or not, because the cross-sectional averages of $\lambda_{ir}$ equal zero.} We also report the size of the corresponding tests (with $5\%$ nominal size) and the average length of the CIs (with $95\%$ nominal coverage). For simplicity and for brevity of the reported results, we assume that the number of factors is known. We consider a case when the number of factors is overspecified in Appendix \ref{ssec: MC extended}.

The LS estimator is heavily biased and the associated tests and CIs are heavily size distorted unless all the factors are strong. At the same time, the proposed estimator effectively reduces the ``weak factors'' bias without inflating the variance. As a result, the potential efficiency gains from using the debiased estimator can be very large when there is a weak factor, especially for larger sample sizes (see Appendix \ref{ssec: additional simulation results} for additional simulation results). Importantly, even if all the factors are strong, the debiased estimator performs comparably to the LS estimator.

When weak factors are present, the LS CIs can have zero coverage because they are (i) centered around the biased LS estimator and (ii) too short. Hence, the average length of the LS CIs is not a proper benchmark to compare the average length of the bias-aware CIs. To provide a relevant comparison, we also construct identification robust CIs by inverting the (absolute value of the) LS based t-statistic using appropriate identification robust critical values (instead of $z_{1-\alpha/2}$). Specifically, for a given design (here, for fixed $N$, $T$, and $R$), we (numerically) compute the least favorable (over $\kappa$) critical value for the absolute value of the t-statistic based on the LS estimator. We also construct analogous CIs by inverting the (absolute value of the) t-statistic based on the debiased estimator using the corresponding least favorable critical values. We refer to such CIs as the LS and debiased oracle CIs (because they are based on unknown design-specific least favorable critical values) and report their average length denoted by ``length*'' in the tables below.

Notice that the average length of the LS oracle CIs (the ``length*'' column under the LS heading)
is at least comparable to but mostly significantly greater than the actual length of the bias-aware CIs (the ``length'' column under the debiased heading), especially for larger sample sizes (again, see Appendix \ref{ssec: additional simulation results} for additional simulation results).
Thus, the bias-aware CI outperforms the LS CI once one corrects the LS CI to compensate for its severe undercoverage.

Another important comparison is between the actual length and oracle length of the bias-aware CI.  Throughout most of the designs, the oracle length of the bias-aware CI is slightly less than half the length of the actual bias-aware CI that we compute.  This gives a bound on how conservative our CI is: our bias-aware critical value cannot be decreased by more than a factor of about two without sacrificing coverage in these Monte Carlos.  There are two possible sources of this conservativeness: (1) the bound in Theorem \ref{new_nuclear_bound_error_thm} may be conservative or (2) there may be some additional structure in the initial error or its correlation with the data that our nuclear norm debiasing method does not exploit.  While further improving the CI using the proof techniques in this paper appears difficult, we cannot rule out these possibilities.  On the other hand, it is possible that these Monte Carlos overstate the conservativeness of our bias-aware CI: there may be other DGPs for which our bias-aware critical value cannot be decreased without sacrificing coverage.



Despite the simplicity of the design considered in this section, the presented findings seem to be characteristic of more complicated and settings. Specifically, in Appendix \ref{ssec: MC extended}, we consider a design with an additional covariate and non-Gaussian, heteroskedastic, and serially correlated errors and establish qualitatively similar results regardless of whether the correct number of factors is known or overspecified.


\subsection{Empirical Illustration}
\label{ssec: empirical}
In this section, we illustrate the finite sample properties of the proposed estimator and confidence intervals in a numerical experiment calibrated to imitate an actual empirical setting. Specifically, we calibrate our experiment based on the seminal studies of the effects of unilateral divorce law reforms on the US divorce rates by \citet{Friedberg1998} and \citet{Wolfers2006}, subsequently revisited by \citet{KimOka2014} and \citet{MoonWeidner2015} in the context of interactive fixed effects models.

For simplicity of the experiment, as a benchmark, we use the following static specification also considered in \citet{Friedberg1998} and \citet{Wolfers2006}
\begin{align*}
  Y_{it} = X_{it} \beta + \alpha_i + \zeta_i t + \nu_i t^2 + \phi_t + U_{it},
\end{align*}
where $Y_{it}$ denotes the annual divorce rate (per 1,000 persons) in state $i$ in year $t$, and $X_{it}$ is a dummy variable indicating if state $i$ had a unilateral divorce law in year $t$. Following \citet{Friedberg1998} and \citet{Wolfers2006}, we also control for state-specific quadratic time trends and time effects.

We follow \citet{KimOka2014} and use their data to construct a balanced panel with $N = 48$ states and $T = 33$ years. As in \citet{MoonWeidner2015}, first we profile out the individual trends and time effects from $Y_{it}$ and $X_{it}$ to form the projected model
\begin{align*}
  Y_{it}^{\perp} = X_{it}^{\perp} \beta + U_{it}^{\perp}
\end{align*}
and obtain the estimates $\hat \beta$ and $\hat \sigma_{U^\perp}^2$. We also extract the first principal component of the matrix of regressors $X^\perp$ denoted by $\Gamma^{X^\perp} = \lambda_i^{X^\perp} f_t^{X^\perp}$.

In our numerical experiment, we fix $X^\perp$, $\{\lambda_i^{X^\perp}\}_{i=1}^N$, and $\{f_t^{X^\perp}\}_{t=1}^T$, and consider the following DGP
\begin{align*}
  Y_{it}^{\perp} = X_{it}^\perp \hat \beta + \kappa \lambda_i^{X^\perp} f_t^{X^\perp} + U_{it}^\perp,
\end{align*}
where we introduce an additional factor $f_{t}^{X^\perp}$ and a parameter $\kappa$ controlling the strength of $f_{t}^{X^\perp}$. For every repetition, we draw $U_{it}^{\perp}$ as iid $N(0, \hat \sigma_U^2)$ and treat all the other parts of the DGP as fixed.

As before, we compare the LS estimator and inference performance with the proposed approach for various values of $\kappa$. Both approaches use the correctly specified number of factors $R = 1$. The results are based on 5,000 simulations and provided in Table \ref{tab: empriical MC}. We report the same statistics as in Section \ref{ssec: MC}.

The results are qualitatively similar to the results in Section \ref{ssec: MC}. The LS estimator is heavily biased when the factor is weak, and the standard tests and confidence intervals are severely size distorted. Compared to the LS estimator, the debiased estimator has a substantially smaller bias, standard deviation, and rmse when the factor is weak. It also performs competitively if the factor is strong. The LS CIs are much shorter than the bias-aware CIs but have very poor coverage. The oracle CIs based on the LS estimator have the correct coverage and are also considerably wider than the naive CIs and comparable with the bias-aware CIs. Again, the oracle CIs based on the debiased estimator are considerably shorter than the bias-aware CIs and LS oracle CIs, indicating that there is a potential scope for improvement.

Overall, our empirically calibrated simulation study shows that the presence of a weak factor can lead to poor performance of conventional estimators and inference procedures in an actual empirical setting. It also demonstrates that in such settings, the gains from using the debiased estimator could be substantial.

Finally, we also report estimation and inference results for the actual data set. For consistency with the numerical experiment above, we focus on the same single covariate $X_{it}$. In Appendix \ref{sec: empirical app}, we also consider a specification with dynamic treatment effects as in \citet{Wolfers2006}. Similarly to \citet{KimOka2014} and \citet{MoonWeidner2015}, we estimate
\begin{align*}
  Y_{it} = X_{it} \beta + \alpha_i + \zeta_i t + \nu_i t^2 + \phi_t + \sum_{r=1}^R \lambda_{ir} f_{tr} + U_{it}
\end{align*}
for various values of $R$ using the LS and the debiased approaches and construct 95\% CIs for $\beta$. As before, we first profile out the individual trends and time effects, and then use the residual outcomes and regressors as inputs for the LS and debiased estimators.

The results are provided in Table \ref{tab: empirical res}.
We construct three types of CIs based on the debiased estimator using $\hat C$ as in Remark~\ref{rem: R_w bound} with different values of $R_w$. The first one is constructed assuming that there are no weak factors among $R$ factors ($R_w = 0$), i.e., under the same assumption under which the standard LS CI is valid. This is the same CI as introduced in Remark~\ref{remark: strong factors CS}. In this application, these CIs are as short as or even shorter than the LS CIs. Thus, if the researcher wants to obtain shorter CIs at the cost of non-robustness to the potential presence of weak factors, they can still do that using our debiased estimator. As pointed out in Remark~\ref{remark: strong factors CS} and documented in Section~\ref{ssec: MC non-robust coverage}, such CIs are still likely to have much better coverage than the LS ones when there is a weak factors since they are based on the debiased estimator.


The second type of CIs is constructed assuming that among $R$ factors there is up to one weak factor ($R_w = 1$). The corresponding bias-aware CIs are substantially wider than the non-robust ones. However, as the numerical experiment considered earlier in this sections suggests, this is how wide identification robust CIs appear to have to be in this setting. In the considered application, we find that the potential presence of one weak factor is likely to be sufficient to nullify the significance of the previously obtained non-robust estimates.

Finally, we also report our bias-aware CIs as in \eqref{eq: bias aware CI} corresponding to $R_w = R$.  These CIs are uniformly valid regardless of the strength of identification of the factors.

The results for a specification with dynamic treatment effects are qualitatively similar and provided in Appendix~\ref{sec: empirical app}.





\begin{table}[H]
\caption{Simulation results for the experiment in Section \ref{ssec: MC}, $N = 100$, $R = 1$}
\label{tab: N100 R1}
\begin{center}
\resizebox{\textwidth}{!}{
\begin{tabular}{c c c c c c c| c c c c c c}
&\multicolumn{6}{c}{LS}&\multicolumn{6}{c}{Debiased} \\
\cmidrule(lr){2-7}  \cmidrule(lr){8-13}
{$\kappa$}&{bias}&{std}&{rmse}&{size}&{length}&{length*}&{bias}&{std}&{rmse}&{size}&{length}&{length*}\\
\midrule
\multicolumn{13}{c}{$T = 20$}\\
\midrule
0.00&-0.0000&0.0171&0.0171&7.1&0.061&0.299&0.0002&0.0206&0.0206&0.0&0.304&0.137 \\
0.05&0.0242&0.0178&0.0300&37.3&0.062&0.300&0.0095&0.0207&0.0228&0.0&0.304&0.137 \\
0.10&0.0478&0.0200&0.0518&79.3&0.062&0.302&0.0181&0.0215&0.0281&0.0&0.305&0.137 \\
0.15&0.0690&0.0249&0.0734&91.6&0.063&0.308&0.0244&0.0235&0.0339&0.0&0.306&0.138 \\
0.20&0.0792&0.0382&0.0879&85.7&0.067&0.324&0.0250&0.0276&0.0372&0.0&0.309&0.138 \\
0.25&0.0670&0.0531&0.0855&64.8&0.074&0.358&0.0189&0.0306&0.0360&0.0&0.311&0.139 \\
0.50&0.0049&0.0244&0.0248&8.2&0.087&0.425&0.0013&0.0239&0.0240&0.0&0.314&0.139  \\
1.00&0.0004&0.0232&0.0232&5.9&0.088&0.427&0.0001&0.0237&0.0237&0.0&0.315&0.140  \\
\midrule
\multicolumn{13}{c}{$T = 50$}\\
\midrule
0.00&-0.0002&0.0103&0.0103&5.9&0.039&0.228&-0.0001&0.0136&0.0136&0.0&0.173&0.079\\
0.05&0.0244&0.0108&0.0267&67.5&0.039&0.228&0.0064&0.0137&0.0151&0.0&0.173&0.080 \\
0.10&0.0484&0.0124&0.0500&98.2&0.039&0.230&0.0121&0.0143&0.0187&0.0&0.174&0.080 \\
0.15&0.0683&0.0189&0.0709&96.8&0.040&0.237&0.0135&0.0164&0.0213&0.0&0.175&0.080 \\
0.20&0.0580&0.0390&0.0699&72.4&0.046&0.269&0.0084&0.0180&0.0198&0.0&0.177&0.080 \\
0.25&0.0229&0.0306&0.0382&33.5&0.053&0.308&0.0032&0.0164&0.0167&0.0&0.177&0.080 \\
0.50&0.0016&0.0144&0.0145&5.7&0.055&0.324&0.0002&0.0151&0.0151&0.0&0.177&0.080  \\
1.00&0.0001&0.0142&0.0142&5.1&0.055&0.324&-0.0001&0.0151&0.0151&0.0&0.178&0.080 \\
\midrule
\multicolumn{13}{c}{$T = 100$}\\
\midrule
0.00&-0.0001&0.0073&0.0073&6.1&0.028&0.183&-0.0001&0.0108&0.0108&0.0&0.122&0.057 \\
0.05&0.0246&0.0077&0.0258&91.0&0.028&0.183&0.0051&0.0108&0.0120&0.0&0.122&0.057  \\
0.10&0.0486&0.0093&0.0495&99.9&0.028&0.185&0.0089&0.0117&0.0147&0.0&0.123&0.057  \\
0.15&0.0619&0.0224&0.0658&92.9&0.030&0.197&0.0068&0.0132&0.0149&0.0&0.124&0.058  \\
0.20&0.0239&0.0267&0.0358&47.4&0.037&0.243&0.0023&0.0124&0.0127&0.0&0.124&0.058  \\
0.25&0.0077&0.0120&0.0143&17.9&0.039&0.256&0.0009&0.0119&0.0120&0.0&0.124&0.058  \\
0.50&0.0009&0.0103&0.0103&5.9&0.039&0.260&0.0001&0.0118&0.0118&0.0&0.124&0.058   \\
1.00&0.0001&0.0102&0.0102&5.4&0.039&0.260&-0.0000&0.0118&0.0118&0.0&0.124&0.058  \\
\midrule
\multicolumn{13}{c}{$T = 300$}\\
\midrule
0.00&-0.0000&0.0042&0.0042&5.3&0.016&0.121&0.0000&0.0056&0.0056&0.0&0.080&0.033 \\
0.05&0.0247&0.0046&0.0252&100.0&0.016&0.122&0.0047&0.0056&0.0073&0.0&0.080&0.033\\
0.10&0.0482&0.0070&0.0487&99.8&0.016&0.123&0.0057&0.0067&0.0088&0.0&0.080&0.033 \\
0.15&0.0178&0.0173&0.0248&60.5&0.021&0.161&0.0015&0.0063&0.0065&0.0&0.081&0.033 \\
0.20&0.0047&0.0064&0.0080&16.4&0.022&0.170&0.0005&0.0061&0.0061&0.0&0.081&0.033 \\
0.25&0.0023&0.0060&0.0064&7.6&0.023&0.171&0.0003&0.0061&0.0061&0.0&0.081&0.033  \\
0.50&0.0003&0.0057&0.0058&4.9&0.023&0.172&0.0001&0.0060&0.0060&0.0&0.081&0.033  \\
1.00&0.0001&0.0057&0.0057&4.9&0.023&0.172&0.0000&0.0060&0.0060&0.0&0.081&0.033  \\
\bottomrule
\bottomrule
\multicolumn{13}{p{1.2\textwidth}}{${\rm Lind}(A) \in \{0.0063,0.0028,0.0015,0.0006\}$ for $T \in \{20,50,100,300\}$.}
\end{tabular}
}
\end{center}
\end{table}

\newgeometry{left=1.5cm,right=1.5cm,top=3cm,bottom=2cm}
\begin{landscape}

\begin{table}[H]
\caption{Simulation results for the experiment in Section \ref{ssec: MC}, $N = 100$, $T = 50$, $R = 2$}
\label{tab: N100 T50 R2 estimation}
\begin{scriptsize}
\begin{center}
\begin{tabular}{ l| c c c c c c c c c c| c c c c c c c c c c}
     & \multicolumn{10}{c}{LS} & \multicolumn{10}{c}{Debiased}\\
     \diagbox[width=12mm,height=8mm]{$\kappa_1$}{$\kappa_2$}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}\\
     \midrule
     \multicolumn{21}{c}{bias}\\
     \midrule
     0.00 & -0.000&0.016&0.031&0.035&0.018&0.008&0.004&0.002&0.001&0.000&-0.000&0.006&0.010&0.010&0.005&0.002&0.001&0.000&0.000&-0.000\\
     0.05 & 0.016&0.033&0.049&0.060&0.052&0.036&0.030&0.027&0.026&0.025&0.006&0.012&0.017&0.018&0.014&0.010&0.009&0.008&0.007&0.007\\
     0.10 & 0.031&0.049&0.066&0.080&0.083&0.067&0.057&0.051&0.050&0.049&0.010&0.017&0.022&0.025&0.022&0.018&0.015&0.014&0.014&0.013\\
     0.15 & 0.035&0.060&0.080&0.097&0.108&0.099&0.082&0.072&0.070&0.068&0.010&0.018&0.025&0.029&0.028&0.024&0.020&0.018&0.017&0.017\\
     0.20 & 0.018&0.052&0.083&0.108&0.123&0.114&0.089&0.068&0.064&0.060&0.005&0.014&0.022&0.028&0.029&0.024&0.018&0.014&0.013&0.012\\
     0.25 & 0.008&0.036&0.067&0.099&0.114&0.088&0.054&0.032&0.028&0.025&0.002&0.010&0.018&0.024&0.024&0.017&0.010&0.006&0.005&0.005\\
     0.30 & 0.004&0.030&0.057&0.082&0.089&0.054&0.027&0.015&0.012&0.010&0.001&0.009&0.015&0.020&0.018&0.010&0.005&0.003&0.002&0.002\\
     0.40 & 0.002&0.027&0.051&0.072&0.068&0.032&0.015&0.007&0.005&0.004&0.000&0.008&0.014&0.018&0.014&0.006&0.003&0.001&0.001&0.001\\
     0.50 & 0.001&0.026&0.050&0.070&0.064&0.028&0.012&0.005&0.004&0.002&0.000&0.007&0.014&0.017&0.013&0.005&0.002&0.001&0.001&0.000\\
     1.00 & 0.000&0.025&0.049&0.068&0.060&0.025&0.010&0.004&0.002&0.000&-0.000&0.007&0.013&0.017&0.012&0.005&0.002&0.001&0.000&-0.000\\
     \midrule
     \multicolumn{21}{c}{std}\\
     \midrule
     0.00 & 0.009&0.009&0.012&0.019&0.019&0.013&0.011&0.011&0.010&0.010&0.013&0.013&0.014&0.015&0.015&0.014&0.014&0.014&0.014&0.014\\
     0.05 & 0.009&0.009&0.010&0.014&0.021&0.015&0.012&0.011&0.011&0.011&0.013&0.013&0.013&0.015&0.015&0.014&0.014&0.014&0.014&0.014\\
     0.10 & 0.012&0.010&0.010&0.012&0.019&0.020&0.014&0.013&0.013&0.013&0.014&0.013&0.014&0.015&0.016&0.015&0.015&0.014&0.014&0.014\\
     0.15 & 0.019&0.014&0.012&0.012&0.017&0.025&0.021&0.019&0.019&0.020&0.015&0.015&0.015&0.016&0.016&0.017&0.016&0.016&0.016&0.016\\
     0.20 & 0.019&0.021&0.019&0.017&0.025&0.043&0.045&0.040&0.039&0.039&0.015&0.015&0.016&0.016&0.018&0.020&0.020&0.019&0.019&0.019\\
     0.25 & 0.013&0.015&0.020&0.025&0.043&0.064&0.055&0.037&0.034&0.032&0.014&0.014&0.015&0.017&0.020&0.023&0.021&0.018&0.018&0.017\\
     0.30 & 0.011&0.012&0.014&0.021&0.045&0.055&0.036&0.021&0.019&0.019&0.014&0.014&0.015&0.016&0.020&0.021&0.018&0.016&0.016&0.016\\
     0.40 & 0.011&0.011&0.013&0.019&0.040&0.037&0.021&0.016&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.018&0.016&0.016&0.016&0.016\\
     0.50 & 0.010&0.011&0.013&0.019&0.039&0.034&0.019&0.015&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.018&0.016&0.016&0.015&0.015\\
     1.00 & 0.010&0.011&0.013&0.020&0.039&0.032&0.019&0.015&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.017&0.016&0.016&0.015&0.015\\
     \midrule
     \multicolumn{21}{c}{rmse}\\
     \midrule
     0.00 & 0.009&0.019&0.033&0.039&0.026&0.015&0.012&0.011&0.010&0.010&0.013&0.014&0.017&0.018&0.016&0.014&0.014&0.014&0.014&0.014\\
     0.05 & 0.019&0.034&0.050&0.061&0.056&0.039&0.032&0.029&0.028&0.027&0.014&0.017&0.021&0.023&0.021&0.017&0.016&0.016&0.016&0.016\\
     0.10 & 0.033&0.050&0.066&0.081&0.085&0.070&0.058&0.053&0.051&0.050&0.017&0.021&0.026&0.029&0.027&0.023&0.021&0.020&0.020&0.020\\
     0.15 & 0.039&0.061&0.081&0.098&0.109&0.102&0.085&0.075&0.073&0.071&0.018&0.023&0.029&0.033&0.033&0.029&0.026&0.024&0.024&0.023\\
     0.20 & 0.026&0.056&0.085&0.109&0.125&0.122&0.099&0.079&0.075&0.071&0.016&0.021&0.027&0.033&0.034&0.032&0.027&0.024&0.023&0.022\\
     0.25 & 0.015&0.039&0.070&0.102&0.122&0.109&0.077&0.049&0.044&0.041&0.014&0.017&0.023&0.029&0.032&0.028&0.023&0.019&0.018&0.018\\
     0.30 & 0.012&0.032&0.058&0.085&0.099&0.077&0.045&0.026&0.023&0.021&0.014&0.016&0.021&0.026&0.027&0.023&0.018&0.016&0.016&0.016\\
     0.40 & 0.011&0.029&0.053&0.075&0.079&0.049&0.026&0.017&0.016&0.016&0.014&0.016&0.020&0.024&0.024&0.019&0.016&0.016&0.016&0.016\\
     0.50 & 0.010&0.028&0.051&0.073&0.075&0.044&0.023&0.016&0.015&0.015&0.014&0.016&0.020&0.024&0.023&0.018&0.016&0.016&0.016&0.015\\
     1.00 & 0.010&0.027&0.050&0.071&0.071&0.041&0.021&0.016&0.015&0.015&0.014&0.016&0.020&0.023&0.022&0.018&0.016&0.016&0.015&0.015\\
     \bottomrule
     \bottomrule
\end{tabular}
\end{center}
\end{scriptsize}
\end{table}

\begin{table}[H]
\caption{Simulation results for the experiment in Section \ref{ssec: MC}, $N = 100$, $T = 50$, $R = 2$}
\label{tab: N100 T50 R2 inference}
\begin{scriptsize}
\begin{center}
\begin{tabular}{ l| c c c c c c c c c c| c c c c c c c c c c}
     & \multicolumn{10}{c}{LS} & \multicolumn{10}{c}{Debiased}\\
     \diagbox[width=12mm,height=8mm]{$\kappa_1$}{$\kappa_2$}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}\\
     \midrule
     \multicolumn{21}{c}{size}\\
     \midrule
     0.00 & 7.6&52.3&88.3&80.0&42.8&17.9&10.6&6.7&6.3&5.9&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.05 & 52.3&96.2&99.8&99.6&95.9&88.3&81.6&73.7&70.8&68.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.10 & 88.3&99.8&100.0&100.0&100.0&99.7&99.2&98.6&98.4&98.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.15 & 80.0&99.6&100.0&100.0&99.8&99.5&98.7&97.9&97.4&96.6&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.20 & 42.8&95.9&100.0&99.8&98.3&93.1&87.3&80.0&76.9&73.8&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.25 & 17.9&88.3&99.7&99.5&93.1&76.1&60.5&45.2&39.9&36.2&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.30 & 10.6&81.6&99.2&98.7&87.3&60.5&38.9&23.3&18.9&16.2&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.40 & 6.7&73.7&98.6&97.9&80.0&45.2&23.3&11.3&9.2&7.9&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     0.50 & 6.3&70.8&98.4&97.4&76.9&39.9&18.9&9.2&7.4&6.4&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     1.00 & 5.9&68.0&98.0&96.6&73.8&36.2&16.2&7.9&6.4&5.7&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\
     \midrule
     \multicolumn{21}{c}{length}\\
     \midrule
     0.00 & 0.031&0.031&0.032&0.034&0.037&0.038&0.039&0.039&0.039&0.039&0.257&0.257&0.258&0.260&0.261&0.262&0.262&0.262&0.262&0.262\\
     0.05 & 0.031&0.031&0.032&0.033&0.036&0.038&0.038&0.039&0.039&0.039&0.257&0.257&0.258&0.259&0.261&0.262&0.262&0.262&0.262&0.262\\
     0.10 & 0.032&0.032&0.032&0.032&0.034&0.037&0.039&0.039&0.039&0.039&0.258&0.258&0.258&0.260&0.261&0.262&0.263&0.263&0.263&0.263\\
     0.15 & 0.034&0.033&0.032&0.032&0.033&0.036&0.039&0.040&0.040&0.040&0.260&0.259&0.260&0.261&0.263&0.264&0.265&0.265&0.265&0.265\\
     0.20 & 0.037&0.036&0.034&0.033&0.034&0.037&0.042&0.045&0.045&0.046&0.261&0.261&0.261&0.263&0.265&0.266&0.267&0.267&0.268&0.268\\
     0.25 & 0.038&0.038&0.037&0.036&0.037&0.043&0.049&0.052&0.052&0.052&0.262&0.262&0.262&0.264&0.266&0.268&0.268&0.269&0.269&0.269\\
     0.30 & 0.039&0.038&0.039&0.039&0.042&0.049&0.053&0.054&0.054&0.054&0.262&0.262&0.263&0.265&0.267&0.268&0.269&0.269&0.269&0.269\\
     0.40 & 0.039&0.039&0.039&0.040&0.045&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.267&0.269&0.269&0.269&0.269&0.270\\
     0.50 & 0.039&0.039&0.039&0.040&0.045&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.268&0.269&0.269&0.269&0.270&0.270\\
     1.00 & 0.039&0.039&0.039&0.040&0.046&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.268&0.269&0.269&0.270&0.270&0.270\\
     \midrule
     \multicolumn{21}{c}{length*}\\
     \midrule
     0.00 & 0.345&0.346&0.353&0.373&0.406&0.420&0.424&0.427&0.427&0.428&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.115\\
     0.05 & 0.346&0.346&0.349&0.360&0.391&0.416&0.424&0.427&0.428&0.429&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.115\\
     0.10 & 0.353&0.349&0.348&0.354&0.375&0.409&0.424&0.430&0.432&0.432&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.116\\
     0.15 & 0.373&0.360&0.354&0.354&0.365&0.399&0.428&0.441&0.443&0.445&0.115&0.115&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116\\
     0.20 & 0.406&0.391&0.375&0.365&0.371&0.410&0.459&0.491&0.497&0.503&0.115&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116\\
     0.25 & 0.420&0.416&0.409&0.399&0.410&0.475&0.536&0.568&0.573&0.576&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116&0.116\\
     0.30 & 0.424&0.424&0.424&0.428&0.459&0.536&0.579&0.595&0.598&0.599&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116&0.117\\
     0.40 & 0.427&0.427&0.430&0.441&0.491&0.568&0.595&0.604&0.606&0.607&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117\\
     0.50 & 0.427&0.428&0.432&0.443&0.497&0.573&0.598&0.606&0.608&0.609&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117\\
     1.00 & 0.428&0.429&0.432&0.445&0.503&0.576&0.599&0.607&0.609&0.610&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117&0.117\\
     \bottomrule
     \bottomrule
\end{tabular}
\end{center}
\end{scriptsize}
\end{table}


\end{landscape}
\restoregeometry


\begin{table}[H]
\begin{center}
\caption{Simulation results for the empirically calibrated experiment, $N = 48$, $T = 33$, $R = 1$}
\label{tab: empriical MC}
\resizebox{\textwidth}{!}{
\begin{tabular}{c c c c c c c| c c c c c c}
&\multicolumn{6}{c}{LS}&\multicolumn{6}{c}{Debiased}\\
\cmidrule(lr){2-7}  \cmidrule(lr){8-13}
{$\kappa$}&{bias}&{std}&{rmse}&{size}&{length}&{length*}&{bias}&{std}&{rmse}&{size}&{length}&{length*}\\
\midrule
0.00& -0.0007& 0.0647& 0.0647& 6.9& 0.236& 1.062& -0.0010& 0.0797& 0.0797& 0.0& 1.374& 0.683  \\
0.20& 0.0920& 0.0656& 0.1130& 35.0& 0.237& 1.067& 0.0517& 0.0805& 0.0957& 0.0& 1.376& 0.685  \\
0.40& 0.1822& 0.0703& 0.1953& 81.9& 0.240& 1.083& 0.1013& 0.0842& 0.1317& 0.0& 1.381& 0.690  \\
0.60& 0.2620& 0.0890& 0.2767& 93.4& 0.248& 1.116& 0.1392& 0.0966& 0.1694& 0.0& 1.392& 0.698  \\
0.80& 0.2999& 0.1467& 0.3339& 84.6& 0.262& 1.180& 0.1428& 0.1267& 0.1910& 0.0& 1.406& 0.703  \\
1.00& 0.2356& 0.2168& 0.3202& 57.5& 0.286& 1.288& 0.0972& 0.1470& 0.1762& 0.0& 1.418& 0.700  \\
1.20& 0.1134& 0.1955& 0.2260& 26.8& 0.307& 1.381& 0.0420& 0.1258& 0.1326& 0.0& 1.424& 0.693  \\
1.40& 0.0427& 0.1168& 0.1243& 12.6& 0.315& 1.419& 0.0177& 0.1042& 0.1057& 0.0& 1.426& 0.690  \\
1.60& 0.0235& 0.0921& 0.0950& 8.8& 0.317& 1.428& 0.0104& 0.0992& 0.0997& 0.0& 1.427& 0.690  \\
1.80& 0.0157& 0.0879& 0.0893& 7.3& 0.318& 1.431& 0.0070& 0.0981& 0.0984& 0.0& 1.428& 0.689  \\
2.00& 0.0112& 0.0867& 0.0875& 6.8& 0.318& 1.432& 0.0050& 0.0977& 0.0978& 0.0& 1.428& 0.689  \\
\bottomrule
\bottomrule
\end{tabular}
}
\end{center}
\end{table}

\begin{table}[H]
\caption{LS and debiased estimates and 95\% CIs for $\beta$}
\label{tab: empirical res}
\begin{small}
\begin{center}
\begin{tabular}{l c c c c c c c c c}
& $R=1$& $R=2$& $R=3$& $R=4$& $R=5$& $R=6$ \\
\midrule
\multicolumn{7}{c}{LS}\\
\midrule
 & $0.047$& $0.160$& $0.101$& $0.043$& $0.028$& $0.091$ \\
& $[-0.06,0.15]$& $[0.04,0.28]$& $[-0.02,0.22]$& $[-0.07,0.16]$& $[-0.10,0.16]$& $[-0.04,0.22]$ \\
\midrule
\multicolumn{7}{c}{Debiased}\\
\midrule
  & $0.089$& $0.162$& $0.130$& $0.084$& $0.071$& $0.106$ \\
$R_w = 0$& $[-0.01,0.19]$& $[0.07,0.26]$& $[0.05,0.21]$& $[0.01,0.16]$& $[-0.01,0.15]$& $[0.04,0.18]$ \\
$R_w = 1$& $[-0.77,0.95]$& $[-0.56,0.88]$& $[-0.45,0.71]$& $[-0.40,0.57]$& $[-0.34,0.48]$& $[-0.24,0.45]$ \\
$R_w = R$& $[-0.77,0.95]$& $[-1.18,1.50]$& $[-1.43,1.69]$& $[-1.62,1.79]$& $[-1.67,1.81]$& $[-1.61,1.82]$ \\
\bottomrule
\bottomrule
\end{tabular}
\end{center}
\end{small}
\end{table}






















\setlength{\bibsep}{2pt}
\begin{thebibliography}{}

\bibitem[\protect\citeauthoryear{Ahn and Horenstein}{Ahn and
  Horenstein}{2013}]{AhnHorenstein2013}
Ahn, S.~C. and A.~R. Horenstein (2013).
\newblock Eigenvalue ratio test for the number of factors.
\newblock {\em Econometrica\/}~{\em 81\/}(3), 1203--1227.

\bibitem[\protect\citeauthoryear{Ahn, Lee, and Schmidt}{Ahn, Lee and
  Schmidt}{2001}]{AhnLeeSchmidt2001}
Ahn, S.~C., Y.~H. Lee, and P.~Schmidt (2001).
\newblock {GMM} estimation of linear panel data models with time-varying
  individual effects.
\newblock {\em Journal of Econometrics\/}~{\em 101\/}(2), 219--255.

\bibitem[\protect\citeauthoryear{Ahn, Lee, and Schmidt}{Ahn, Lee and
  Schmidt}{2013}]{AhnLeeSchmidt2013}
Ahn, S.~C., Y.~H. Lee, and P.~Schmidt (2013).
\newblock Panel data models with multiple time-varying individual effects.
\newblock {\em Journal of Econometrics\/}~{\em 174\/}(1), 1--14.

\bibitem[\protect\citeauthoryear{Alidaee, Auerbach, and Leung}{Alidaee,
  Auerbach and Leung}{2020}]{alidaee2020recovering}
Alidaee, H., E.~Auerbach, and M.~P. Leung (2020).
\newblock Recovering network structure from aggregated relational data using
  penalized regression.
\newblock {\em arXiv preprint arXiv:2001.06052\/}.

\bibitem[\protect\citeauthoryear{Andrews and Cheng}{Andrews and
  Cheng}{2012}]{andrews2012estimation}
Andrews, D.~W. and X.~Cheng (2012).
\newblock Estimation and inference with weak, semi-strong, and strong
  identification.
\newblock {\em Econometrica\/}~{\em 80\/}(5), 2153--2211.

\bibitem[\protect\citeauthoryear{Arellano}{Arellano}{1987}]{arellano1987computing}
Arellano, M. (1987).
\newblock Computing robust standard errors for within-groups estimators.
\newblock {\em Oxford Bulletin of Economics \& Statistics\/}~{\em 49\/}(4).

\bibitem[\protect\citeauthoryear{Arkhangelsky, Athey, Hirshberg, Imbens, and
  Wager}{Arkhangelsky, Athey, Hirshberg, Imbens and
  Wager}{2021}]{arkhangelsky_synthetic_2021}
Arkhangelsky, D., S.~Athey, D.~A. Hirshberg, G.~W. Imbens, and S.~Wager (2021,
  December).
\newblock Synthetic {Difference}-in-{Differences}.
\newblock {\em American Economic Review\/}~{\em 111\/}(12), 4088--4118.

\bibitem[\protect\citeauthoryear{Armstrong and Koles{\'a}r}{Armstrong and
  Koles{\'a}r}{2018}]{armstrong2018optimal}
Armstrong, T.~B. and M.~Koles{\'a}r (2018).
\newblock Optimal inference in a class of regression models.
\newblock {\em Econometrica\/}~{\em 86\/}(2), 655--683.

\bibitem[\protect\citeauthoryear{Armstrong, Kolesár, and Kwon}{Armstrong,
  Kolesár and Kwon}{2020}]{armstrong_bias-aware_2020}
Armstrong, T.~B., M.~Kolesár, and S.~Kwon (2020).
\newblock Bias-{Aware} {Inference} in {Regularized} {Regression} {Models}.
\newblock {\em arXiv:2012.14823 [econ, stat]\/}.

\bibitem[\protect\citeauthoryear{Athey, Bayati, Doudchenko, Imbens, and
  Khosravi}{Athey, Bayati, Doudchenko, Imbens and
  Khosravi}{2021}]{athey_matrix_2021}
Athey, S., M.~Bayati, N.~Doudchenko, G.~Imbens, and K.~Khosravi (2021).
\newblock Matrix {Completion} {Methods} for {Causal} {Panel} {Data} {Models}.
\newblock {\em Journal of the American Statistical Association\/}~{\em 0\/}(0),
  1--15.

\bibitem[\protect\citeauthoryear{Bai}{Bai}{2009}]{bai2009panel}
Bai, J. (2009).
\newblock Panel data models with interactive fixed effects.
\newblock {\em Econometrica\/}~{\em 77\/}(4), 1229--1279.

\bibitem[\protect\citeauthoryear{Bai and Ng}{Bai and Ng}{2002}]{BaiNg2002}
Bai, J. and S.~Ng (2002, January).
\newblock Determining the number of factors in approximate factor models.
\newblock {\em Econometrica\/}~{\em 70\/}(1), 191--221.

\bibitem[\protect\citeauthoryear{Bai and Ng}{Bai and Ng}{2017}]{BaiNg2017}
Bai, J. and S.~Ng (2017).
\newblock Principal components and regularized estimation of factor models.
\newblock {\em arXiv preprint arXiv:1708.08137\/}.

\bibitem[\protect\citeauthoryear{Bai and Ng}{Bai and
  Ng}{2023}]{bai2023approximate}
Bai, J. and S.~Ng (2023).
\newblock Approximate factor models with weaker loadings.
\newblock {\em Journal of Econometrics\/}~{\em 235\/}(2), 1893--1916.

\bibitem[\protect\citeauthoryear{Bai and Wang}{Bai and
  Wang}{2016}]{bai2016econometric}
Bai, J. and P.~Wang (2016).
\newblock Econometric analysis of large factor models.
\newblock {\em Annual Review of Economics\/}~{\em 8}, 53--80.

\bibitem[\protect\citeauthoryear{Belloni, Chen, Madrid~Padilla, and
  Wang}{Belloni, Chen, Madrid~Padilla and Wang}{2023}]{belloni2019high}
Belloni, A., M.~Chen, O.~H. Madrid~Padilla, and Z.~Wang (2023).
\newblock High-dimensional latent panel quantile regression with an application
  to asset pricing.
\newblock {\em The Annals of Statistics\/}~{\em 51\/}(1), 96--121.

\bibitem[\protect\citeauthoryear{Belloni, Chernozhukov, and Hansen}{Belloni,
  Chernozhukov and Hansen}{2014}]{belloni_inference_2014}
Belloni, A., V.~Chernozhukov, and C.~Hansen (2014).
\newblock Inference on {Treatment} {Effects} after {Selection} among
  {High}-{Dimensional} {Controls}.
\newblock {\em The Review of Economic Studies\/}~{\em 81\/}(2), 608--650.

\bibitem[\protect\citeauthoryear{Beyhum and Gautier}{Beyhum and
  Gautier}{2019}]{beyhum2019square}
Beyhum, J. and E.~Gautier (2019).
\newblock Square-root nuclear norm penalized estimator for panel data models
  with approximately low-rank unobserved heterogeneity.
\newblock {\em arXiv preprint arXiv:1904.09192\/}.

\bibitem[\protect\citeauthoryear{Beyhum and Gautier}{Beyhum and
  Gautier}{2022}]{beyhum2022factor}
Beyhum, J. and E.~Gautier (2022).
\newblock Factor and factor loading augmented estimators for panel regression
  with possibly nonstrong factors.
\newblock {\em Journal of Business \& Economic Statistics\/}, 1--12.

\bibitem[\protect\citeauthoryear{Billingsley}{Billingsley}{1995}]{billingsley1995probability}
Billingsley, P. (1995).
\newblock {\em Probability and measure\/} (3rd ed.).
\newblock A Wiley-Interscience publication. John Wiley and Sons.

\bibitem[\protect\citeauthoryear{Bonhomme and Manresa}{Bonhomme and
  Manresa}{2015}]{bonhomme2015grouped}
Bonhomme, S. and E.~Manresa (2015).
\newblock Grouped patterns of heterogeneity in panel data.
\newblock {\em Econometrica\/}~{\em 83\/}(3), 1147--1184.

\bibitem[\protect\citeauthoryear{Chamberlain and Moreira}{Chamberlain and
  Moreira}{2009}]{Chamberlain2009}
Chamberlain, G. and M.~J. Moreira (2009).
\newblock Decision theory applied to a linear panel data model.
\newblock {\em Econometrica\/}~{\em 77\/}(1), 107--133.

\bibitem[\protect\citeauthoryear{Chernozhukov, Hansen, Liao, and
  Zhu}{Chernozhukov, Hansen, Liao and Zhu}{2019}]{chernozhukov2019inference}
Chernozhukov, V., C.~Hansen, Y.~Liao, and Y.~Zhu (2019).
\newblock Inference for heterogeneous effects using low-rank estimation of
  factor slopes.

\bibitem[\protect\citeauthoryear{Chetverikov and Manresa}{Chetverikov and
  Manresa}{2022}]{chetverikov2022spectral}
Chetverikov, D. and E.~Manresa (2022).
\newblock Spectral and post-spectral estimators for grouped panel data models.
\newblock {\em arXiv preprint arXiv:2212.13324\/}.

\bibitem[\protect\citeauthoryear{Cox}{Cox}{2024}]{cox2024weak}
Cox, G.~F. (2024).
\newblock Weak identification in low-dimensional factor models with one or two
  factors.
\newblock {\em The Review of Economics and Statistics\/}~{\em 03}, 1--45.

\bibitem[\protect\citeauthoryear{Donoho}{Donoho}{1994}]{donoho1994statistical}
Donoho, D.~L. (1994).
\newblock Statistical estimation and optimal recovery.
\newblock {\em The Annals of Statistics\/}, 238--270.

\bibitem[\protect\citeauthoryear{Fan and Liao}{Fan and
  Liao}{2022}]{fan2022learning}
Fan, J. and Y.~Liao (2022).
\newblock Learning latent factors from diversified projections and its
  applications to over-estimated and weak factors.
\newblock {\em Journal of the American Statistical Association\/}~{\em
  117\/}(538), 909--924.

\bibitem[\protect\citeauthoryear{Fan}{Fan}{1951}]{fan1951maximum}
Fan, K. (1951).
\newblock Maximum properties and inequalities for the eigenvalues of completely
  continuous operators.
\newblock {\em Proceedings of the National Academy of Sciences\/}~{\em
  37\/}(11), 760--766.

\bibitem[\protect\citeauthoryear{Feng}{Feng}{2023}]{feng_2023}
Feng, J. (2023).
\newblock Nuclear norm regularized quantile regression with interactive fixed
  effects.
\newblock {\em Econometric Theory\/}, 1–31.

\bibitem[\protect\citeauthoryear{Ferman and Pinto}{Ferman and
  Pinto}{2021}]{ferman_synthetic_2021}
Ferman, B. and C.~Pinto (2021).
\newblock Synthetic controls with imperfect pretreatment fit.
\newblock {\em Quantitative Economics\/}~{\em 12\/}(4), 1197--1221.

\bibitem[\protect\citeauthoryear{Fern{\'a}ndez-Val, Freeman, and
  Weidner}{Fern{\'a}ndez-Val, Freeman and Weidner}{2021}]{fernandez2021low}
Fern{\'a}ndez-Val, I., H.~Freeman, and M.~Weidner (2021).
\newblock Low-rank approximations of nonseparable panel models.
\newblock {\em The Econometrics Journal\/}~{\em 24\/}(2), C40--C77.

\bibitem[\protect\citeauthoryear{Freeman and Weidner}{Freeman and
  Weidner}{2023}]{freeman2023linear}
Freeman, H. and M.~Weidner (2023).
\newblock Linear panel regressions with two-way unobserved heterogeneity.
\newblock {\em Journal of Econometrics\/}~{\em 237\/}(1), 105498.

\bibitem[\protect\citeauthoryear{Friedberg}{Friedberg}{1998}]{Friedberg1998}
Friedberg, L. (1998).
\newblock Did unilateral divorce raise divorce rates? {E}vidence from panel
  data.
\newblock {\em American Economic Review\/}, 608--627.

\bibitem[\protect\citeauthoryear{Geman}{Geman}{1980}]{geman1980limit}
Geman, S. (1980).
\newblock A limit theorem for the norm of random matrices.
\newblock {\em The Annals of Probability\/}~{\em 8\/}(2), 252--261.

\bibitem[\protect\citeauthoryear{Hansen}{Hansen}{2007}]{hansen2007asymptotic}
Hansen, C.~B. (2007).
\newblock Asymptotic properties of a robust variance matrix estimator for panel
  data when t is large.
\newblock {\em Journal of Econometrics\/}~{\em 141\/}(2), 597--620.

\bibitem[\protect\citeauthoryear{Hastie, Tibshirani, and Wainwright}{Hastie,
  Tibshirani and Wainwright}{2015}]{Hastieetal2015}
Hastie, T., R.~Tibshirani, and M.~Wainwright (2015).
\newblock {\em Statistical learning with sparsity: the lasso and
  generalizations}.
\newblock CRC press.

\bibitem[\protect\citeauthoryear{Higgins}{Higgins}{2021}]{higgins2021fixed}
Higgins, A. (2021).
\newblock Fixed {$T$} estimation of linear panel data models with interactive
  fixed effects.
\newblock {\em arXiv preprint arXiv:2110.05579\/}.

\bibitem[\protect\citeauthoryear{Hirshberg and Wager}{Hirshberg and
  Wager}{2020}]{hirshberg_augmented_2020}
Hirshberg, D.~A. and S.~Wager (2020).
\newblock Augmented minimax linear estimation.
\newblock arXiv: 1712.00038.

\bibitem[\protect\citeauthoryear{Holtz-Eakin, Newey, and Rosen}{Holtz-Eakin,
  Newey and Rosen}{1988}]{HoltzEakin-Newey-Rosen1988}
Holtz-Eakin, D., W.~Newey, and H.~S. Rosen (1988).
\newblock Estimating vector autoregressions with panel data.
\newblock {\em Econometrica\/}~{\em 56\/}(6), 1371--95.

\bibitem[\protect\citeauthoryear{Horn and Johnson}{Horn and
  Johnson}{2013}]{horn2013matrix}
Horn, R.~A. and C.~R. Johnson (2013).
\newblock {\em Matrix Analysis\/} (2 ed.).
\newblock Cambridge University Press.

\bibitem[\protect\citeauthoryear{Ibragimov and Khas’minskii}{Ibragimov and
  Khas’minskii}{1985}]{ibragimov_nonparametric_1985}
Ibragimov, I. and R.~Khas’minskii (1985).
\newblock On {Nonparametric} {Estimation} of the {Value} of a {Linear}
  {Functional} in {Gaussian} {White} {Noise}.
\newblock {\em Theory of Probability \& Its Applications\/}~{\em 29\/}(1),
  18--32.

\bibitem[\protect\citeauthoryear{Ishihara and Kitagawa}{Ishihara and
  Kitagawa}{2021}]{ishihara_evidence_2021}
Ishihara, T. and T.~Kitagawa (2021).
\newblock Evidence {Aggregation} for {Treatment} {Choice}.
\newblock {\em arXiv preprint arXiv:2108.06473\/}.

\bibitem[\protect\citeauthoryear{Javanmard and Montanari}{Javanmard and
  Montanari}{2014}]{javanmard_confidence_2014}
Javanmard, A. and A.~Montanari (2014).
\newblock Confidence {Intervals} and {Hypothesis} {Testing} for
  {High}-{Dimensional} {Regression}.
\newblock {\em Journal of Machine Learning Research\/}~{\em 15\/}(82),
  2869--2909.

\bibitem[\protect\citeauthoryear{Johnstone}{Johnstone}{2001}]{Johnstone2001}
Johnstone, I. (2001).
\newblock {On the distribution of the largest eigenvalue in principal
  components analysis}.
\newblock {\em Annals of Statistics\/}~{\em 29\/}(2), 295--327.

\bibitem[\protect\citeauthoryear{Juodis and Sarafidis}{Juodis and
  Sarafidis}{2018}]{juodis2018fixed}
Juodis, A. and V.~Sarafidis (2018).
\newblock Fixed {$T$} dynamic panel data estimators with multifactor errors.
\newblock {\em Econometric Reviews\/}~{\em 37\/}(8), 893--929.

\bibitem[\protect\citeauthoryear{Juodis and Sarafidis}{Juodis and
  Sarafidis}{2022}]{juodis2022linear}
Juodis, A. and V.~Sarafidis (2022).
\newblock A linear estimator for factor-augmented fixed-{$T$} panels with
  endogenous regressors.
\newblock {\em Journal of Business \& Economic Statistics\/}~{\em 40\/}(1),
  1--15.

\bibitem[\protect\citeauthoryear{Kiefer}{Kiefer}{1980}]{Kiefer1980}
Kiefer, N. (1980).
\newblock {A time series-cross section model with fixed effects with an
  intertemporal factor structure}.
\newblock {\em Unpublished manuscript, Department of Economics, Cornell
  University\/}.

\bibitem[\protect\citeauthoryear{Kim and Oka}{Kim and Oka}{2014}]{KimOka2014}
Kim, D. and T.~Oka (2014).
\newblock Divorce law reforms and divorce rates in the usa: An interactive
  fixed-effects approach.
\newblock {\em Journal of Applied Econometrics\/}.

\bibitem[\protect\citeauthoryear{Ma, Su, and Zhang}{Ma, Su and
  Zhang}{2022}]{ma2022detecting}
Ma, S., L.~Su, and Y.~Zhang (2022).
\newblock Detecting latent communities in network formation models.
\newblock {\em The Journal of Machine Learning Research\/}~{\em 23\/}(1),
  13971--14031.

\bibitem[\protect\citeauthoryear{Miao, Li, and Su}{Miao, Li and
  Su}{2020}]{miao2020panel}
Miao, K., K.~Li, and L.~Su (2020).
\newblock Panel threshold models with interactive fixed effects.
\newblock {\em Journal of Econometrics\/}~{\em 219\/}(1), 137--170.

\bibitem[\protect\citeauthoryear{Miao, Phillips, and Su}{Miao, Phillips and
  Su}{2023}]{miao2023high}
Miao, K., P.~C. Phillips, and L.~Su (2023).
\newblock High-dimensional vars with common factors.
\newblock {\em Journal of Econometrics\/}~{\em 233\/}(1), 155--183.

\bibitem[\protect\citeauthoryear{Moon and Weidner}{Moon and
  Weidner}{2015}]{MoonWeidner2015}
Moon, H.~R. and M.~Weidner (2015).
\newblock Linear regression for panel with unknown number of factors as
  interactive fixed effects.
\newblock {\em Econometrica\/}~{\em 83\/}(4), 1543--1579.

\bibitem[\protect\citeauthoryear{Moon and Weidner}{Moon and
  Weidner}{2018}]{moon2018nuclear}
Moon, H.~R. and M.~Weidner (2018).
\newblock Nuclear norm regularized estimation of panel regression models.
\newblock {\em arXiv preprint arXiv:1810.10987\/}.

\bibitem[\protect\citeauthoryear{Noack and Rothe}{Noack and
  Rothe}{2024}]{noack_bias-aware_2024}
Noack, C. and C.~Rothe (2024).
\newblock Bias-{Aware} {Inference} in {Fuzzy} {Regression} {Discontinuity}
  {Designs}.
\newblock {\em Econometrica\/}~{\em 92\/}(3), 687--711.

\bibitem[\protect\citeauthoryear{Onatski}{Onatski}{2010}]{Onatski2010}
Onatski, A. (2010).
\newblock Determining the number of factors from empirical distribution of
  eigenvalues.
\newblock {\em The Review of Economics and Statistics\/}~{\em 92\/}(4),
  1004--1016.

\bibitem[\protect\citeauthoryear{Onatski}{Onatski}{2012}]{Onatski2012}
Onatski, A. (2012).
\newblock Asymptotics of the principal components estimator of large factor
  models with weakly influential factors.
\newblock {\em Journal of Econometrics\/}~{\em 168\/}(2), 244--258.

\bibitem[\protect\citeauthoryear{Pesaran}{Pesaran}{2006}]{Pesaran2006estimation}
Pesaran, M.~H. (2006).
\newblock Estimation and inference in large heterogeneous panels with a
  multifactor error structure.
\newblock {\em Econometrica\/}~{\em 74\/}(4), 967--1012.

\bibitem[\protect\citeauthoryear{Pillai and Yin}{Pillai and
  Yin}{2012}]{pillai2012edge}
Pillai, N.~S. and J.~Yin (2012).
\newblock Edge universality of correlation matrices.
\newblock {\em The Annals of Statistics\/}, 1737--1763.

\bibitem[\protect\citeauthoryear{Recht, Fazel, and Parrilo}{Recht, Fazel and
  Parrilo}{2010}]{RechtFazelParrilo2010}
Recht, B., M.~Fazel, and P.~A. Parrilo (2010).
\newblock Guaranteed minimum-rank solutions of linear matrix equations via
  nuclear norm minimization.
\newblock {\em SIAM review\/}~{\em 52\/}(3), 471--501.

\bibitem[\protect\citeauthoryear{Robertson and Sarafidis}{Robertson and
  Sarafidis}{2015}]{robertson2015iv}
Robertson, D. and V.~Sarafidis (2015).
\newblock {IV} estimation of panels with factor residuals.
\newblock {\em Journal of Econometrics\/}~{\em 185\/}(2), 526--541.

\bibitem[\protect\citeauthoryear{Robinson}{Robinson}{1988}]{robinson_root-n-consistent_1988}
Robinson, P.~M. (1988).
\newblock Root-{N}-{Consistent} {Semiparametric} {Regression}.
\newblock {\em Econometrica\/}~{\em 56\/}(4), 931--954.

\bibitem[\protect\citeauthoryear{Rohde and Tsybakov}{Rohde and
  Tsybakov}{2011}]{RohdeTsybakov2011}
Rohde, A. and A.~B. Tsybakov (2011).
\newblock Estimation of high-dimensional low-rank matrices.
\newblock {\em The Annals of Statistics\/}~{\em 39\/}(2), 887--930.

\bibitem[\protect\citeauthoryear{Soshnikov}{Soshnikov}{2002}]{Soshnikov2002}
Soshnikov, A. (2002).
\newblock {A note on universality of the distribution of the largest
  eigenvalues in certain sample covariance matrices}.
\newblock {\em Journal of Statistical Physics\/}~{\em 108\/}(5), 1033--1056.

\bibitem[\protect\citeauthoryear{van~de Geer, Bühlmann, Ritov, and
  Dezeure}{van~de Geer, Bühlmann, Ritov and
  Dezeure}{2014}]{van_de_geer_asymptotically_2014}
van~de Geer, S., P.~Bühlmann, Y.~Ritov, and R.~Dezeure (2014).
\newblock On asymptotically optimal confidence regions and tests for
  high-dimensional models.
\newblock {\em The Annals of Statistics\/}~{\em 42\/}(3), 1166--1202.

\bibitem[\protect\citeauthoryear{Wang, Su, and Zhang}{Wang, Su and
  Zhang}{2022}]{wang2022low}
Wang, Y., L.~Su, and Y.~Zhang (2022).
\newblock Low-rank panel quantile regression: Estimation and inference.
\newblock {\em arXiv preprint arXiv:2210.11062\/}.

\bibitem[\protect\citeauthoryear{Westerlund, Petrova, and Norkute}{Westerlund,
  Petrova and Norkute}{2019}]{westerlund2019cce}
Westerlund, J., Y.~Petrova, and M.~Norkute (2019).
\newblock {CCE} in fixed-{$T$} panels.
\newblock {\em Journal of Applied Econometrics\/}~{\em 34\/}(5), 746--761.

\bibitem[\protect\citeauthoryear{Wolfers}{Wolfers}{2006}]{Wolfers2006}
Wolfers, J. (2006).
\newblock Did unilateral divorce laws raise divorce rates? a reconciliation and
  new results.
\newblock {\em American Economic Review\/}~{\em 96}, 1802--1820.

\bibitem[\protect\citeauthoryear{Yang}{Yang}{2019}]{yang2019edge}
Yang, F. (2019).
\newblock Edge universality of separable covariance matrices.
\newblock {\em Electron. J. Probab\/}~{\em 24\/}(123), 1--57.

\bibitem[\protect\citeauthoryear{Yata}{Yata}{2021}]{yata_optimal_2021}
Yata, K. (2021).
\newblock Optimal decision rules under partial identification.
\newblock {\em arXiv preprint arXiv:2111.04926\/}.

\bibitem[\protect\citeauthoryear{Zeleneev}{Zeleneev}{2019}]{zeleneev2019identification}
Zeleneev, A. (2019).
\newblock Identification and estimation of network models with nonparametric
  unobserved heterogeneity.
\newblock {\em Department of Economics, Princeton University\/}.

\bibitem[\protect\citeauthoryear{Zhang and Zhang}{Zhang and
  Zhang}{2014}]{zhang_confidence_2014}
Zhang, C.-H. and S.~S. Zhang (2014).
\newblock Confidence intervals for low dimensional parameters in high
  dimensional linear models.
\newblock {\em Journal of the Royal Statistical Society: Series B (Statistical
  Methodology)\/}~{\em 76\/}(1), 217--242.

\bibitem[\protect\citeauthoryear{Zhu}{Zhu}{2019}]{zhu2019well}
Zhu, Y. (2019).
\newblock How well can we learn large factor models without assuming strong
  factors?
\newblock {\em arXiv preprint arXiv:1910.10382\/}.

\end{thebibliography}


\newpage