EconBase
← Back to paper

Robust Estimation and Inference in Panels with Interactive Fixed Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

126,851 characters · 15 sections · 113 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Estimation and Inference in Panels with Interactive Fixed Effects

\thispagestyle{empty} \setcounter{page}{0}

abstractWe consider estimation and inference for a regression coefficient in panels with interactive fixed effects (i.e., with a factor structure). We demonstrate that existing estimators and confidence intervals (CIs) can be heavily biased and size-distorted when some of the factors are weak. We propose estimators with improved rates of convergence and bias-aware CIs that remain valid uniformly, regardless of factor strength. Our approach applies the theory of minimax linear estimation to form a debiased estimate, using a nuclear norm bound on the error of an initial estimate of the interactive fixed effects. Our resulting bias-aware CIs take into account the remaining bias caused by weak factors. Monte Carlo experiments show substantial improvements over conventional methods when factors are weak, with minimal costs to estimation accuracy when factors are strong.

\vskip 3cm

Introduction

In this paper, we consider a linear panel regression model of the form

align[align omitted — 111 chars of source]

where $Y_{it}, X_{it}, Z_{1,it}, \ldots, Z_{K,it} \in \mathbb{R}$ are the observed outcome variable and covariates for units $i=1,\ldots,N$ and time periods $t=1,\ldots,T$. The error components $\Gamma_{it} \in \mathbb{R}$ and $U_{it} \in \mathbb{R}$ are unobserved, and the regression coefficients $\beta, \delta_1, \ldots, \delta_K \in \mathbb{R}$ are unknown. The parameter of interest is $\beta \in \mathbb{R}$, the coefficient on $X_{it}$. We are interested in “large panels”, where both $N$ and $T$ are relatively large.

The error component $U_{it}$ is modelled as a mean-zero random shock that is uncorrelated with the regressors $X_{it}$ and $Z_{k,it}$ and that is at most weakly autocorrelated across $i$ and over $t$. By contrast, the error component $\Gamma_{it}$ can be correlated with $X_{it}$ and $Z_{k,it}$ and can also be strongly autocorrelated across $i$ and over $t$. Of course, further restrictions on $\Gamma_{it}$ are required to allow estimation and inference on $\beta$. For example, the additive fixed effect model imposes that $\Gamma_{it} = \alpha_i + \gamma_t$, where $\alpha_i$ accounts for any omitted variable that is constant over time, and $\gamma_t$ for any omitted variable that is constant across units. Instead of this additive fixed effect model we consider the so-called interactive fixed effect model, where

align[align omitted — 101 chars of source]

Here, the $\lambda_{ir}$ and $f_{tr} $ can either be interpreted as unknown parameters or as unobserved shocks. This model for $\Gamma_{it} $ is also known as a factor model, with factor loadings $\lambda_{ir}$ and factors $f_{tr}$. We will use the terms factor and interactive fixed effect interchangeably. The number of factors $R$ is unknown, but is assumed to be small relative to $N$ and $T$. The interactive fixed effect model is attractive because it introduces enough restrictions to allow estimation and inference on $\beta$ while still incorporating or approximating a large class of data generating processes (DGPs) for $ \Gamma_{it} $.

The existing econometrics literature on panel regressions with interactive fixed effects is quite large. Since the seminal work of Pesaran2006estimation and bai2009panel, developing tools for estimation and inference on $\beta$ in model (ref)-(ref) under large $N$ and large $T$ asymptotics has been a primary focus of this literature. Specifically, Pesaran2006estimation introduces the common correlated effects (CCE) estimator, which uses cross-sectional averages of the observed variables as proxies for the unobserved factors. bai2009panel derives the large $N$, $T$ properties of the least-squares (LS) estimator that jointly minimizes the sum of squared residuals over the regression coefficients, factors, and factor loadings.\footnote{This estimator was first introduced by Kiefer1980.}

bai2009panel shows that, under appropriate assumptions, the LS estimator for the regression coefficients is $\sqrt{NT}$-consistent and asymptotically normally distributed as both $N$ and $T$ grow to infinity. One of the key assumptions imposed for this result is the so-called “strong factor assumption”, which requires all the factor loadings $ \lambda_{ir}$ and factors $f_{tr}$ to have sufficient variation across $i$ and over $t$, respectively. If the strong factor assumption is violated, then the LS estimator for $ \lambda_{ir}$ and $f_{tr}$ may be unable to pick up the true loadings and factors correctly, because the “weak factors”\footnote{ See, for example, Onatski2010,Onatski2012 for a discussion and formalization of the notion of weak factors. } in $ \Gamma_{it}$ cannot be distinguished from the noise in $U_{it}$. This can lead to substantial bias and misleading inference, due to omitted variables bias from $\Gamma_{it}$ that is not picked up by the estimator.

figure[figure omitted — 920 chars of source]

To illustrate how this can lead to problems with conventional estimates and CIs for $\beta$, Figure (ref) presents a subset of the results of our Monte Carlo study.\footnote{A detailed description of the numerical experiment is provided in Section (ref).} When the factors are nonexistent (panel a) or strongly identified (panel d), the distribution of the LS estimator (in blue) is centered at the true parameter value $\beta$ (equal to $0$ in this case). However, when the factors are present but weak enough that they are difficult to estimate (panels b and c), the LS estimator is heavily biased and non-normally distributed. In our Monte Carlo study in Section (ref), we show that this indeed leads to severe coverage distortion, with conventional CIs based on the LS estimator having almost zero coverage.

In this paper, we address this issue by developing new tools for estimation and inference on $\beta$ in the model (ref). We develop a debiased estimator along with a bound on the remaining bias, which we use to construct a bias-aware confidence interval. As illustrated in Figure (ref), our debiased estimator (shown in red) substantially decreases the bias of the LS estimator when factors are weak, leading to a large improvement in overall estimation error. In addition, this improved performance under weak factors does not come at a substantial cost to performance when factors are strong or nonexistent: our debiased estimator performs similarly to the LS estimator in these cases. Importantly, our CI requires only an upper bound on the number of factors: we show that it is valid uniformly over a large class of DGPs that allows for weak, strong or nonexistent factors up to a specified upper bound on the number of factors. We derive rates of convergence that hold uniformly over this class of DGPs, and we show that our estimator achieves a faster uniform rate of convergence than existing approaches when weak factors are allowed. In the case where $N$ and $T$ grow at the same rate, our estimator achieves the parametric $\sqrt{NT}$ rate.

Our debiasing approach uses a preliminary estimate $\hat\Gamma_{\rm pre}$ of the individual effect matrix $\Gamma$ along with a bound $\hat C$ on the nuclear norm $\Vert \Gamma - \hat \Gamma_{\rm pre}\Vert_*$ of its estimation error. Letting $\tilde \Gamma:= \Gamma - \hat \Gamma_{\rm pre}$, we then consider the augmented outcomes

align*[align* omitted — 147 chars of source]

Treating $\tilde \Gamma_{it}$ as nuisance parameters satisfying a convex constraint $\Vert \tilde \Gamma \Vert_{*} \le \hat C$, we derive linear weights $A_{it}$ such that the estimator $\sum_{i=1}^N\sum_{t=1}^TA_{it}\tilde Y_{it}$ for $\beta$ optimally uses this constraint, using the theory of minimax linear estimators ibragimov_nonparametric_1985,donoho1994statistical,armstrong2018optimal. In particular, the resulting weights $A_{it}$ control the remaining omitted variables bias $\sum_{i=1}^N\sum_{t=1}^TA_{it}\tilde \Gamma_{it}$ due to possible weak factors in $\tilde \Gamma=\Gamma-\hat\Gamma_{\rm pre}$ not picked up by the initial estimate $\hat\Gamma_{\rm pre}$.

A key step in deriving our CI is the construction of the preliminary estimator $\hat\Gamma_{\rm pre}$ and bound $\hat C$ on the nuclear norm of its estimation error. Our CI is bias-aware: it uses the bound $\hat C$ to explicitly take into account any remaining bias in the debiased estimator. Our bound is feasible once an upper bound on the number of factors is specified. In our Monte Carlo study, we find that, while our CIs are often conservative, they are about as wide as an “oracle” CI that uses an infeasible critical value to correct the coverage of a CI based on the standard LS estimator.

While our results allow for arbitrary sequences of weak factors, our conditions on other aspects of the model are similar to bai2009panel and MoonWeidner2015. An important condition is that the covariate of interest $X_{it}$ must not itself be entirely explained by a low dimensional factor model. For example, in a panel where $X_{it}$ is the minimum hourly wage in state $i$ and year $t$, we would require that states change their minimum wage laws in different years, and that this is done sufficiently often to generate variation in $X_{it}$ that cannot be explained by a small number of factors $f_t$. This rules out settings where $X_{it}$ is an indicator variable for a policy that affects a subset of the units and occurs only during a single time period: in this case, $X_{it}=\lambda_i\cdot f_t$ where $\lambda_i$ is an indicator variable for unit $i$ undergoing the policy change and $f_t$ is an indicator variable for periods after the policy change. See Section (ref) for formal conditions and further discussion.

A special case of the factor model is the grouped unobserved heterogeneity model considered by bonhomme2015grouped. In this model, $\Gamma_{it}=\alpha_{g(i),t}$, where $g(\cdot)$ is an unknown function mapping individuals $i$ to a group index $g(i)\in \{1,\ldots, R\}$. This takes the form of the factor model ((ref)) with $\lambda_{ir}=1$ if $g(i)=r$ and $0$ otherwise, and with $f_{tr}=\alpha_{r,t}$. The strong factor assumption corresponds to the strong group separation assumption imposed in this literature (e.g., Assumption 2(b) in bonhomme2015grouped) which imposes that the group means $\alpha_{r,\cdot}=(\alpha_{r,1},\ldots,\alpha_{r,T})'$ are sufficiently far away for different groups $r$. Our results apply in this setting and allow for this assumption to be relaxed. An interesting question for future research is whether it is possible to modify our approach to take advantage of the additional structure in the grouped unobserved heterogeneity model.

{ \bf Related literature}

The papers by Pesaran2006estimation and bai2009panel mentioned previously have motivated a large follow up literature on large $N$ and $T$ analysis of panel models with interactive effects. bai2016econometric provides a review with further references. Another literature has proposed alternative estimation methods along with asymptotic analysis in the regime with $T$ fixed and $N$ increasing. This includes the quasi-difference approach of HoltzEakin-Newey-Rosen1988 and generalized method of moments approaches of AhnLeeSchmidt2001,AhnLeeSchmidt2013. More recent papers analyzing the fixed $T$ large $N$ regime include robertson2015iv, juodis2018fixed, westerlund2019cce, higgins2021fixed, juodis2022linear. None of these papers provide inference methods that remain valid when factors are weak or rank-deficient (e.g.\ $f=0$). Chamberlain2009 derive estimators that satisfy a Bayes-minimax property over a certain class of priors in a finite sample setting that includes a version of the model ((ref)). This Bayes-minimax property does not, however, translate to a guarantee on coverage or estimation error under weak factors.

A special case of the violation of the strong factor assumption is when some factor are equal to zero, while all other factors are strong; the inference results of bai2009panel are usually robust towards this specific violation of the strong factor assumption MoonWeidner2015. This robustness, however, does not carry over to more general weak factors in the DGP of $\Gamma_{it}$, as illustrated by Figure (ref).

The problem of weak factors is related to the problem of omitted variable bias of LASSO estimators in high dimensional regression that is the focus of debiased LASSO estimators belloni_inference_2014,javanmard_confidence_2014,van_de_geer_asymptotically_2014,zhang_confidence_2014. Just as LASSO estimators omit variables with coefficients that are large enough to cause omitted variables bias but too small to distinguish from zero, weak factors in $\Gamma$ can be difficult to estimate, leading to omitted variables bias in conventional estimates of $\beta$. Our approach to using minimax linear estimation to debias an initial estimate mirrors the approach of javanmard_confidence_2014 to debiasing the LASSO. We discuss this connection further in Section (ref). hirshberg_augmented_2020 provide a general discussion and further references for minimax linear debiasing; we refer to this general approach as augmented linear estimation following their terminology. Minimax linear estimation itself goes back at least to ibragimov_nonparametric_1985, with further results on this approach and its optimality properties in donoho1994statistical, armstrong2018optimal and yata_optimal_2021, among others. The particular form of the minimax estimator used for debiasing in our setup follows from a formula given in armstrong_bias-aware_2020.

Requiring $\Gamma_{it}$ to have the factor structure (ref) is equivalent to requiring the matrix of unobserved effects $\Gamma$ to have rank at most $R$, i.e., having ${\rm rank}(\Gamma) \leq R$. Bounding the nuclear norm of $\tilde \Gamma$ or $\Gamma$ instead can also be seen as a convex relaxation of this requirement. Similar convexifications of the rank constraint have been widely used in the matrix completion literature (e.g., RechtFazelParrilo2010 and Hastieetal2015 for recent surveys), and for reduced rank regression estimation (e.g., RohdeTsybakov2011). In the econometrics literature, the numerous applications of this idea include, for example, estimation of pure factor models BaiNg2017, estimation of panel regression models with homogeneous moon2018nuclear,beyhum2019square and heterogeneous coefficients chernozhukov2019inference, estimation of treatment effects athey_matrix_2021,fernandez2021low, and many others.\footnote{For example, recent economic applications of nuclear norm and related penalization methods also include latent community detection alidaee2020recovering,ma2022detecting, quantile regression belloni2019high,wang2022low,feng_2023, and estimation of panel threshold models and high-dimensional VARs (miao2020panel and miao2023high).} However, none of these papers obtain asymptotically valid CIs or improved rates of convergence under weak factors.

In recent work, chetverikov2022spectral propose an estimator that, like ours, achieves a faster rate of convergence than conventional approaches under weak factors.\footnote{The main focus of chetverikov2022spectral is the grouped effects model of bonhomme2015grouped, which is a special case of the interactive fixed effects setting we consider here. However, the authors extend their results to the general interactive fixed effects setting.} While chetverikov2022spectral allow for weak factors in some of their estimation results, they assume strong factors when constructing CIs. The estimation approach in chetverikov2022spectral also differs from our approach by using modelling assumptions that place a factor structure on the covariate matrix $X$.

Our focus is on allowing for weak factors without imposing additional assumptions on the error term $U$, such as homoskedasticity or full independence from the individual effects $\Gamma$ and regressor $X$. Such additional structure allows for further identifying information by making it easier to distinguish between the error term $U$ and the individual effects $\Gamma$, leading to a fundamentally different analysis. zhu2019well derives asymptotic upper and lower bounds for estimators and CIs in a setting with possible weak factors under homoskedastic and fully independent errors. The estimators and CIs constructed by zhu2019well take advantage of the additional structure of zhu2019well's setting, making them inapplicable in ours. However, the lower bounds derived by zhu2019well are immediately relevant: they show that no CI can be asymptotically valid under weak factors while mimicking the performance of the CI of bai2009panel when factors are strong.

As discussed above, our assumptions rule out the case where $X_{it}$ is an indicator variable for a policy that affects a subset of units starting in the same time period. Recent papers that analyze such settings include ferman_synthetic_2021 and arkhangelsky_synthetic_2021. The fact that $X_{it}$ is collinear with the confounding factor model in this setting presents a fundamental identification issue that requires placing additional conditions on the model. In contrast to this literature, our goal is to leverage variation in $X_{it}$ that cannot be explained by a low dimensional factor model in settings where such variation exists.

beyhum2022factor, fan2022learning, and bai2023approximate consider estimation and inference in various settings under a regime in which a lower bound on the strength of the factors can decrease with $N$ and $T$, but is large enough that factors can be consistently estimated. This is analogous to the “semi-strong” regime in weak instrument and related settings; see andrews2012estimation. While the semi-strong regime requires careful theoretical analysis, the fact that factors can be consistently estimated leads to asymptotically unbiased and normal estimators for the main effect $\beta$. Our results apply to semi-strong and strong regimes as well, while also allowing for weak factor regimes in which factors cannot be consistently estimated.

Finally, cox2024weak develops tools for inference in low-dimensional factor models with weak identification. In cox2024weak, the primary objects of interests are the covariance of the factors and the loadings. The baseline model in cox2024weak does not include observed covariates, whereas we focus on estimation and inference on $\beta$, the coefficient on $X_{it}$, exclusively.\footnote{cox2024weak mentions that observed covariates could, in principle, be incorporated in his framework as long as they are uncorrelated with the unobserved effects, which is a primary worry in the panel literature.}

The rest of this paper is organized as follows. Section (ref) introduces the framework and describes construction of the debiased estimator and bias-aware CI. Section (ref) provides implementation details. Section (ref) provides formal statistical guarantees. Section (ref) considers numerical and empirical illustrations. A supplementary appendix contains all proofs and additional results for the numerical and empirical illustrations.

Construction of robust estimates and confidence intervals

Setup

We consider a panel setting in which we observe a scalar outcome $Y_{it}$, a scalar covariate $X_{it}$ of interest and additional control covariates $\{Z_{k,it}\}_{k=1}^K$ for $i=1,\ldots,N$, $t=1,\ldots, T$, which follow the regression model ((ref)). The error term $U_{it}$ is assumed to be mean zero conditional on $X$, $\{Z_{k,it}\}_{k=1}^K$ and $\Gamma$,\footnote{We note that this requires strict exogeneity and in particular rules out using lagged outcomes as covariates. We leave extensions to models with lagged outcomes as a topic for future research.} but we allow for heteroskedasticity, which may depend on $X_{it}$ and $\Gamma_{it}$, as well as some weak dependence. We write the model in matrix notation as

align[align omitted — 133 chars of source]

where $Z$ denotes the three dimensional array $\{Z_{k,it}\}$ and we define $Z\cdot \delta = \sum_{k=1}^K Z_{k}\delta_k$ where $Z_k$ denotes the matrix with $i,t$-th element $Z_{k,it}$. We use $\lambda$ to denote the $N\times R$ matrix of loadings $\lambda_{ir}$ and $f$ to denote the $T\times R$ matrix of factors $f_{tr}$, so that ((ref)) can be written in matrix form as $\Gamma=\lambda f'$.

We are interested in the coefficient $\beta$ of $X_{it}$, which can be interpreted as the effect of a treatment variable $X_{it}$ in a constant treatment effects model (we discuss extensions to heterogeneous treatment effects in Remark (ref)). For concreteness, we use panel notation, and we refer to $i$ and $t$ as individuals and time periods respectively. However, we allow for other settings such as network data in which $i$ and $t$ both index individuals in a network. While we will assume a low rank structure on $\Gamma$, we allow for arbitrary dependence between the covariate $X_{it}$ and the individual effect $\Gamma_{it}$.

A key ingredient in our approach is an initial estimate $\hat\Gamma$ of $\Gamma$ and a bound on its estimator in the nuclear norm, which holds with probability approaching one:

align[align omitted — 164 chars of source]

Here, $\|\cdot\|_*$ denotes the nuclear norm of the argument matrix, and $\hat C \geq 0$ is a known or estimated constant. We describe our estimate $\hat\Gamma$ and bound $\hat C$ in Section (ref), and we state a formal result giving conditions under which the bound holds with probability approaching one in Section (ref). This bound depends on an upper bound for the number of factors $R$, which must be specified a priori. Importantly, our approach does not require specifying the exact number of factors: some of the factors may be zero, in addition to the possibility of being “weak” in the sense of being close to zero. We emphasize that obtaining an computable upper bound $\hat C$ that enables construction of our CI is itself one of the main technical contributions of this paper.\footnote{As we discuss further in Section (ref), a tighter CI can be constructed by bounding the difference between the estimate $\hat\Gamma$ and $\Gamma+P_{\lambda} U$, where $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$, with $M^+$ denoting the Moore–Penrose inverse of a matrix $M$. The implementation in Section (ref) is for the tighter CI that uses these arguments. For ease of exposition, however, we focus on using the bound (ref) directly in the remainder of this section.}

remarkAlthough the main focus of this paper is on models with the linear factor structure (ref), the methodology presented in this section applies to general interactive fixed effects models as long as it is possible to construct a preliminary estimator $\hat \Gamma$ and a bound $\hat C$ satisfying (ref). For example, we conjecture that our method can also be extended to nonlinear factor models with $\Gamma_{it} = g(\lambda_i,f_t)$, where $g(\cdot,\cdot)$ is some unknown function (e.g., zeleneev2019identification,freeman2023linear). As noted, for example, in fernandez2021low, such $\Gamma$ can be approximated by a low-rank matrix with a (slowly) growing rank $R$. Hence, we expect that our method can be applied in this setting with $\hat \Gamma$ constructed using a growing $R$ and $\hat C$ adjusted for the low-rank approximation error (if needed), in the same spirit as sieve approximations are used in nonparametric estimation.

Augmented linear estimators and CIs

We first define a class of estimators and CIs, indexed by an $N\times T$ matrix $A$. We then provide a choice of the matrix $A$, based on finite sample optimality in an idealized setting. Our class of estimators is given in the following definition.

definitionLet $A=A(X,Z)$ be an $N \times T$ matrix of weights $A_{it} \in \mathbb{R}$ that can depend on the matrix $X$ and array $Z$. Let $\hat\Gamma$ be an initial estimate of $\Gamma$, and let $\widetilde Y=Y-\hat\Gamma$. The augmented linear estimator with weight matrix $A$ and initial estimate $\hat\Gamma$ is given by \begin{align} \hat\beta_A &:= \sum_{i=1}^N \sum_{t=1}^T A_{it}\widetilde Y_{it} =\langle A, \widetilde Y \rangle_F. \end{align}

Here, $\langle \cdot, \cdot \rangle_F$ denotes the entry-wise inner product between the argument matrices.

The estimator $\hat\beta_A=\langle A,\widetilde Y \rangle_F$ applies a linear estimator after an initial estimation step in which the initial estimate $\hat\Gamma$ is subtracted from the outcome $Y$. This mirrors applications of this idea in other settings going back to javanmard_confidence_2014; see hirshberg_augmented_2020 for references (the term “augmented linear estimation” is used in the latter paper).

To analyze this class of estimators, note that subtracting the initial estimate from both sides of the equation ((ref)) gives

align[align omitted — 93 chars of source]

(recall that $\widetilde Y=Y-\hat\Gamma$ and $\widetilde \Gamma=\Gamma-\hat\Gamma$). Our choice of the matrix $A$ will be motivated by a heuristic in which we consider the model ((ref)) with $\widetilde Y$ as the observed outcome and $\widetilde \Gamma$ a nuisance parameter such that $U$ is mean zero conditional on $X,Z$ and $\tilde \Gamma$, with the bound ((ref)) interpreted as a deterministic bound that holds with $\hat C$ nonrandom. This heuristic is not literally true, since $\widetilde\Gamma$ depends on $U$ through the estimation error in the initial estimate $\hat\Gamma$. Nonetheless, the CIs and estimators we obtain will be asymptotically valid and consistent respectively, under conditions that we give in Section (ref).

Following this heuristic, we consider the decomposition

align[align omitted — 172 chars of source]

where

align[align omitted — 226 chars of source]

Under the heuristic where $\tilde\Gamma$ is nuisance parameter in the model ((ref)), $\operatorname{bias}_{\beta,\delta,\widetilde \Gamma}$ gives the bias of the estimator $\hat\beta_{A}$ conditional on $X,Z$ and $\tilde \Gamma$. In reality, $\operatorname{bias}_{\beta,\delta,\widetilde \Gamma}$ does not literally give the bias or conditional bias of $\hat\beta_A$, since conditioning on $\widetilde\Gamma=\Gamma-\hat\Gamma$ means conditioning on an information set that depends on $Y$ through the preliminary estimate $\hat\Gamma$. We nonetheless refer to $\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ as a bias term, following our heuristic.

Let $\widehat{\operatorname{se}}$ be an estimate of the standard deviation of $\langle A,U \rangle_F=\sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it}$. For example, to allow for arbitrary heteroskedasticity in $U_{it}$ while imposing independence across $i$ and $t$, we can use $\widehat{\operatorname{se}}=\sqrt{\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{it}^2}$ where $\hat U_{it}$ denotes residuals from an initial regression. If $\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ were zero, then we could form a CI by adding and subtracting a normal critical value times $\widehat{\operatorname{se}}$. To take into account the possibility that $\operatorname{bias}_{\beta,\delta,\Gamma}(\hat\beta_A)$ will in general be nonnegligible in our setting, we use the bound ((ref)) to obtain an upper bound on the bias term. In particular, when ((ref)) holds, we have $\left| \operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A) \right| \le \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$, where for general $C \geq 0$ we define

align[align omitted — 751 chars of source]

Here, for the second equality we used that the supremum over $\beta$ and $\delta$ is unbounded unless $\langle A, X \rangle_F=1$ and $\langle A, Z_k \rangle_F=0$, and for the final step we used that the nuclear norm $ \| \cdot \|_*$ is dual to the spectral norm, which we denote by $s_1(\cdot)$ since it is equal to the largest singular value of the argument matrix. We refer to $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ as the worst-case bias of the estimator $\hat\beta_A$ (again, this terminology reflects the heuristic in which $\tilde\Gamma$ is treated as a nuisance parameter in ((ref)) rather than estimation error from the initial estimate $\hat\Gamma$).

Note that, whereas $\operatorname{bias}_{\beta,\delta,\widetilde\Gamma}(\hat\beta_A)$ depends on the unknown matrix of individual effects $\Gamma$ through the matrix $\widetilde\Gamma=\Gamma-\hat\Gamma$, $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ is feasible to compute once a bound $\hat C$ is given. Taking into account the possible bias leads to a bias-aware CI:

align[align omitted — 186 chars of source]

To motivate this CI, note that the probability that the lower endpoint is greater than $\beta$ is

align*[align* omitted — 521 chars of source]

where the last step assumes that $\sum_{i=1}^N\sum_{t=1}^TA_{it}U_{it}$ is approximately normally distributed with zero mean and standard deviation close to $ \widehat{\operatorname{se}}$. We provide formal justifications for this later. By a similar argument, the probability that the upper endpoint is less than $\beta$ can be bounded by $\alpha/2$, and these calculations together imply that the coverage of our CI is approximately at least $1-\alpha$.

remarkIn principle, our approach can be extended to a heterogeneous treatment effect model where the constant coefficient $\beta$ is replaced by an individual specific coefficient $\beta_{it}$ that is allowed to vary with $i$ and $t$. In particular, if a bound on the nuclear norm of the matrix of coefficients $\beta_{it}$ or on the error of preliminary estimates of these coefficients is available in addition to such a bound for $\Gamma$, we can use minimax linear debiasing to estimate a linear functional of the individual specific effects $\beta_{it}$. For example, the linear functional $\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T \beta_{it}$ gives the average treatment effect of a one-unit change in $X_{it}$ over the $NT$ units in a setting where $\beta_{it}$ is interpreted as the causal effect of a change in the variable $X_{it}$. Deriving a computable bound on the nuclear norm error of an initial estimate of the coefficients $\beta_{it}$ in this case is nontrivial, however, and we leave this question for future research.

Choice of weights $A=(A_{it})$

As described in the last subsection, one can construct valid confidence intervals for $\beta$ of the form (ref) for any choice of weight matrix $A$, subject to weak regularity conditions. To get a simple baseline procedure, we compute weights that are optimal in an idealized setting where $U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$ independently of $X,Z$ and $\tilde \Gamma$ (again, this involves invoking the heuristic of treating $\tilde\Gamma$ as a nuisance parameter in ((ref)) rather than estimation error from a preliminary estimate). In this idealized setting, $\hat\beta_A$ is then normally distributed with variance $\sigma^2\sum_{i=1}^N\sum_{t=1}^TA_{it}^2=\sigma^2\|A\|_F^2$ (where $\|\cdot\|_F$ denotes the Frobenius norm), and with bias ranging from $-\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ to $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$. Thus, if we choose worst-case MSE under i.i.d. normal errors as our criterion function for the weights, then the optimal weights are obtained by minimizing $\left( \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) \right)^2 + \sigma^2 \|A\|_F^2 $. By substituting the formula for $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_A)$ from ((ref)), we obtain the following baseline choice of weights, indexed by a tuning parameter ${b}$ that corresponds to $\hat C/\sigma$.

definitionFor ${b}>0$, define the “optimal” $N\times T$ weight matrix by \begin{align*} A^*_{{b}} &:= \operatorname*{argmin}_{A \in \mathbb{R}^{N \times T} } \, {b}^2 s_1(A)^2 + \|A\|_F^2 \;\; \qquads.t.\quad $\langle A,X \rangle_F = 1$ and $\langle A,Z_k\cdot \delta \rangle_F = 0$, \end{align*} Here, the constraint $\langle A,Z_k\cdot \delta \rangle_F = 0$ is imposed for all $k \in \{1,\ldots,K\}$.

Heuristically, we expect that a good choice of ${b}$ will correspond to $\hat C/\sigma$ such that the bound $\hat C$ on the nuclear norm holds with high probability. Conveniently, our nuclear norm bound in the exact factor model in Section (ref) scales with the standard deviation $\sigma$ in the homoskedastic case, which gives us a simple and feasible choice of the tuning parameter ${b}$.

We emphasize again that while the definition of $A^*_{b}$ is motivated by the idealized setting $U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$, we do {\it not} assume that the error terms $U_{it}$ satisfy this strong assumption. Choosing $A=A^*_{b}$ to construct the estimator $\hat\beta_A $ and the confidence intervals (ref) under more general error distributions just means that the resulting estimates and confidence intervals will not be optimal (in finite samples), but we will nevertheless show them to be consistent and valid, respectively.

remarkWhile we have used MSE to motivate our baseline choice of weights $A^*_{b}$, one could use other criteria corresponding to different weights on bias and variance. For example, optimizing CI length when $\hat C/\sigma=b$ would give the criterion ${b} s_1(A)+ z_{1-\alpha}\|A\|_F$. If $\beta$ gives the net welfare gain of an all-or-nothing policy change, then one can target minimax welfare regret as in ishihara_evidence_2021 and yata_optimal_2021. In our Monte Carlo simulations however, we find that the exact choice of criterion has little effect on performance.

Practical implementation

The definition of $A^*_{b}$ is a convex optimization problem that can easily be solved numerically for any given input $X$, $Z$, ${b}$. Using results from armstrong_bias-aware_2020, it follows that $A^*_{b}$ can also be computed using the residuals of a nuclear norm regularized regression of $X$ on $Z_1,\ldots,Z_K$ and a matrix of individual effects. When there are no additional covariates $Z$, this nuclear norm regularized regression simplifies further: it can be solved by computing the singular value decomposition of $X$, and then performing soft thresholding on the singular values. The resulting weights $A^*_{b}$ obtained from the residuals of this regression replace the largest singular values of $X$ with a constant. We provide details in Appendix (ref).

In addition to giving alternative methods for computing the weights $A^*_{b}$, these results provide some intuition for these weights. The residuals from this nuclear norm regularized regression of $X$ on $Z_1,\ldots, Z_K$ and the individual effects “partial out” potential correlation of $X$ with the estimation error $\widetilde\Gamma$, similar to the estimator of robinson_root-n-consistent_1988 in the partially linear model. When there are no additional covariates $Z$, this amounts to removing the largest singular values of $X$ and replacing them with a constant.

To summarize, we can compute an estimator $\hat\beta_{A}$ using Definition (ref) using any matrix of weights $A$. We can also compute a CI $\left\{ \hat\beta_A \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2} \widehat{\operatorname{se}} \right] \right\}$ as in ((ref)), once we have a standard error $\widehat{\operatorname{se}}$ and an upper bound $\hat C$ for the nuclear norm of the error in the initial estimate of $\Gamma$. Definition (ref) gives us a heuristic for computing a reasonable choice of the matrix $A$, once we have an initial choice of ${b}$ for the ratio $\hat C/\sigma$ of the nuclear norm bound to variance of $U_{it}$.

Thus, to apply our approach, we need an initial choice ${b}$ to compute the weights $A^*_{b}$ using Definition (ref). We also need a robust upper bound $\hat C$ such that the bound ((ref)) holds with high probability. Finally, we need a robust standard error $\widehat{\operatorname{se}}$. Our CI then takes the form in ((ref)) with $A=A^*_{b}$ and the given bound $\hat C$ and standard error $\widehat{\operatorname{se}}$. In Section (ref), we give details of these choices, as well as how to compute the initial estimate of $\Gamma$.

Implementation

In this section, we describe the implementation of our approach. Our approach relies on bounds for the nuclear norm of the initial estimate of $\Gamma$, derived formally in Section (ref). As explained in Section (ref), a tighter CI can be derived using a more nuanced argument that bounds the difference between $\hat\Gamma$ and $\Gamma+P_{\lambda} U$, where $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$, and $M^+$ denotes the Moore–Penrose inverse of a matrix $M$. In particular, we show that the bound $\hat C\approx 2Rs_1(U)$ can be used. The bound $R$ on the number of factors must be specified by the researcher, similar to other methods in this literature bai2009panel. Furthermore, the weights $A^*_{b}$ are designed to be optimal when $U_{it}\stackrel{iid}{\sim} N(0,\sigma^2)$, which leads to the approximation $s_1(U)/\sigma\approx \sqrt{N}+\sqrt{T}$ geman1980limit. We therefore use ${b}={b^*}:=2R(\sqrt{N}+\sqrt{T})$ as our default choice to calibrate $\hat C/\sigma$ when computing the weights in Definition (ref). We then use an upper bound $\hat C$ that is valid under heteroskedasticity when computing $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{b^*}})$ in the construction of the CI.

Our initial estimate is formed in two steps. First, we form the least squares estimate $\hat\Gamma_{\rm LS}$. We then apply our debiasing approach to get an estimate of the coefficients $\beta$ and $\delta$ and form an $\hat\Gamma_{\rm pre}$ by applying least squares to estimate $\Gamma$, with $\beta$ and $\delta$ fixed at this initial debiased estimate. This estimate $\hat\Gamma_{\rm pre}$ is then used as the initial estimate in our procedure. Thus, our procedure involves applying our debiasing approach twice. This appears to be necessary to get an initial estimate $\hat\Gamma_{\rm pre}$ the best possible nuclear norm bounds on the estimation error.

Below we provide the details of our implementation algorithm.\footnote{Implementation of this algorithm in R is also available at \url{https://github.com/chenweihsiang/PanelIFE/tree/main}. We thank Chen-Wei Hsiang for his excellent assistance in preparing this R package.}

algorithm[algorithm omitted — 2,625 chars of source]
remark[Behavior of estimator under strong factors] In Section (ref), we show that, in the absence of weak factors, $\overline{\operatorname{bias}}_{\hat C}(\hat\beta_{A^*_{{b^*}}})$ becomes negligible relative to $\widehat{\operatorname{se}}$ in the construction of the CI in (ref). Thus, the CI \begin{align} \hat\beta_{A^*_{{b^*}}} \pm z_{1-\alpha/2}\widehat{\operatorname{se}}, \end{align} which uses our bias-corrected estimator but ignores bias when computing the critical value, will have correct asymptotic coverage in the strong factor case. In our Monte Carlos provided in Section (ref), we find that this CI is (i) comparable to alternative non-robust CIs in terms of the length and (ii) despite its non-robustness, substantially less size-distorted if there is a weak factor(s) since it is based on the debiased estimator. In settings where the bias-aware CI ((ref)) is too wide to yield precise inference, we recommend reporting the CI ((ref)) alongside the bias-aware CI ((ref)) as a compromise between ignoring weak factors and a fully robust approach.
remark[Choice of $R$ and $\varepsilon$] The quantity $\varepsilon$ is used in the bound $\hat C = (2 + \varepsilon) R s_1(\hat U_{\rm pre})$ on $\|\widetilde \Gamma\|_*$ needed to compute the CI in the final step. While $\varepsilon>0$ is necessary for theoretical guarantees, in our Monte Carlos, we find that we get good coverage when choosing $\varepsilon=0$. In contrast, the choice of $R$ has a substantive effect on the CI, both through the bound $\hat C = (2 + \varepsilon) R s_1(\hat U_{\rm pre})$ and through the point estimate. Since the number of weak factors cannot be determined from the data, the researcher must specify an a priori bound on the total number of factors $R$. Nonetheless, the data can be informative about the number of strong or semi-strong factors, which provides a lower bound for $R$. We recommend forming an estimate $\hat R_s$ of the number of strong factors using one of the standard methods (e.g., BaiNg2002,Onatski2010,AhnHorenstein2013) and using this as a starting point for examining the sensitivity of the results to the choice of $R$. For example, by taking $R = \hat R_s + 1$, the researcher allows for the potential presence of an additional weak factor (or $R - \hat R_s$ weak factors in general for a bigger $R$).
remark[Lindeberg condition] The asymptotic validity of the CI depends on asymptotic normality of the stochastic term $\langle A, U \rangle_F$ where $A=A^*_{{b^*}}$ is a non-random matrix of weights. This, in turn, depends on a Lindeberg condition on the weights $A$. To ensure that this holds, we can modify our optimization procedure for computing the weights $A=A^*_{{b^*}}$ by imposing a bound on the Lindeberg weights \begin{align} \operatorname{Lind}(A)=\frac{\max_{1\le i\le N, 1\le t\le T} A_{it}^2}{\sum_{i=1}^N\sum_{t=1}^T A_{it}^2}. \end{align} A similar approach to showing asymptotic validity is taken in javanmard_confidence_2014 in a different setting. To make this approach practical, we need guidance on what makes $\operatorname{Lind}(A)$ “small enough to use the central limit theorem” in a given sample size. A formal answer to this question is elusive, due to the difficulty of obtaining finite sample bounds on approximation error in the central limit theorem that are practically useful. As a heuristic, we can use comparisons to other settings where the central limit theorem is used. For example, the sample mean $\bar W=\frac{1}{n}\sum_{i=1}^n W_i$ with $n$ observations corresponds to an estimator with Lindeberg constant $(1/n)^2/[n\cdot (1/n)^2]=1/n$. If we are comfortable using the normal approximation in such a setting with, say, $n=50$, then we can impose a bound $\operatorname{Lind}(A)\le 1/50$. noack_bias-aware_2024 provide some discussion of these issues in a related setting involving inference in fuzzy regression discontinuity. In our Monte Carlos, we find that $\operatorname{Lind}(A)$ is very small for the weights used in Algorithm (ref) once $N$ and $T$ are larger than, say, 20. Thus, imposing a bound on these weights does not appear to be necessary in practice in the data generating processes we have examined.
remark[Standard error] The standard error $\widehat{\operatorname{se}}^2=\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{{\rm pre},it}^2$ assumes that $U_{it}$ is uncorrelated across $i$ and $t$, but allows for heteroskedasticity. Such an assumption will be reasonable if $\Gamma_{it}$ captures all of the dependence in errors for the outcome. However, incorporating all dependence in $\Gamma_{it}$ may lead to an unnecessarily conservative choice of the upper bound $R$ on the number of factors, leading to a wider CI. To avoid such conservative bounds on $\Gamma$, one can incorporate any dependence that is not directly correlated with $X_{it}$ into the error term $U_{it}$, and allow for such dependence when constructing the standard error. For example, to allow for (arbitrary) time dependence of $U_{it}$ (while maintaining uncorrelatedness across $i$), one could simply use clustered standard errors (see, e.g., arellano1987computing,hansen2007asymptotic) \begin{align*} \widehat{\operatorname{se}}^2=\sum_{i=1}^N \left(\sum_{t=1}^T A_{it}\hat U_{{\rm pre},it}\right)^2. \end{align*}

Asymptotic results

This section gives formal asymptotic results for the estimators and CIs given in Sections (ref) and (ref). We consider the following decomposition of our regression model:

align[align omitted — 208 chars of source]

Here, $\widetilde\Gamma$ and $\widetilde U$ an be any $N\times T$ matrices chosen compatibly so that $\widetilde \Gamma + \widetilde U = \Gamma-\hat\Gamma_{\rm pre} + U$. While our discussion so far has focused on the case where $\widetilde \Gamma = \Gamma - \hat \Gamma_{\rm pre}$ and $\widetilde U = U$, it turns out that allowing for other choices of $\widetilde\Gamma$ and $\widetilde U$ allows for an improvement in the width of our CI.

To formally state asymptotic results that allow for weak factors and an unknown error distribution, we introduce some additional notation. We consider uniform-in-the-underlying distribution asymptotics over a set $\mathcal{P}$ of distributions $P$ for $\Gamma$ and $X,Z_1,\ldots,Z_K,U$ and a set $\Theta$ of parameters $\theta=(\beta,\delta')'$. While we treat $\Gamma,X,Z_1,\ldots,Z_k$ as random variables determined by the unknown probability distribution $P$ for notational purposes, we note that a fixed design setting in which $\Gamma,X,Z_1,\ldots,Z_k$ are non-random (sequences of) matrices can be incorporated by considering a set $\mathcal{P}$ that places a probability one mass on a given value of $\Gamma,X,Z_1,\ldots,Z_k$. We use $\mathbb{P}_{P,\theta}$ to denote probability under the given distribution $P$ and parameters $\theta$. Formally, we consider large $N$, large $T$ asymptotics in which $N=N_n\to\infty$ and $T=T_n\to\infty$, and we consider sequences of distributions $\mathcal{P}=\mathcal{P}_n$ and parameter spaces $\Theta=\Theta_n$. Asymptotic statements are then taken in the sequence $n$. However, we suppress the dependence on an index sequence $n$ in order to save on notation. For a sequence of vectors or matrices $A_{N,T}=A_{N,T}(\theta,P)$ of fixed dimension (which may depend on $\theta,P$), we use the notation $A_{N,T}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T})$ when, for every $\varepsilon>0$, there exists $C_\varepsilon$ such that

align*[align* omitted — 158 chars of source]

and we use the notation $A_{N,T}=o_{\Theta,\mathcal{P}}(r_{N,T})$ when, for every $\varepsilon>0$, we have

align*[align* omitted — 147 chars of source]

We use the notation $A_{N,T}\asymp_{\Theta,\mathcal{P}}r_{N,T}$ when $A_{N,T}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T})$ and $A_{N,T}^{-1}=\mathcal{O}_{\Theta,\mathcal{P}}(r_{N,T}^{-1})$. We use the notation $A_{N,T} \underset{\Theta,\mathcal{P}}{\overset{d}{\to}} \mathcal{L}$ to denote the statement

align*[align* omitted — 176 chars of source]

where $F_{\mathcal{L}}$ denotes the cdf of the probability law $\mathcal{L}$.

General results

We first show asymptotic validity of the CI ((ref)) under the following high level assumption imposed on the augmented model (ref) and the weights $A_{it}$.

assumption\begin{enumerate}[(i)] • $\inf_{\theta\in\Theta,P\in\mathcal{P}}\mathbb{P}_{\theta,P}\left( \|\widetilde \Gamma\|_* \le \hat C \right)\to 1$; • $\frac{\langle A, \widetilde U \rangle_F}{\widehat{\operatorname{se}}} \underset{\Theta,\mathcal{P}}{\overset{d}{\to}} N(0,1)$. \end{enumerate}
theoremSuppose that Assumption (ref) holds. Then \begin{align*} \liminf \inf_{\theta\in\Theta,P\in\mathcal{P}} \mathbb{P}_{\theta,P}\left( \beta \in \left\{ \hat\beta_A \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta_A) + z_{1-\alpha/2}\widehat{\operatorname{se}} \right] \right\} \right) \ge 1-\alpha. \end{align*}

Primitive conditions

We now apply these results to the initial estimate and bound given in Section (ref), under the assumption of a linear factor model for $\Gamma$. We allow for a side condition on the Lindeberg weights $\operatorname{Lind}(A)$ defined in ((ref)), as described in Remark (ref). Let $A^*_{{b}, c}$ be defined in the same way as $A^*_{{b}}$, with the modification that we impose the constraint $\operatorname{Lind}(A)\le c$:

align[align omitted — 275 chars of source]

In particular, the weights used in Algorithm (ref) are given by $A^*_{{b^*},\infty}=A^*_{{b^*}}$, and the weights $A^*_{{b^*},c}$ with $c<\infty$ correspond to the modification described in Remark (ref). In our asymptotic theory we will require $c=c_{NT}$ to converge to zero as $N,T \rightarrow \infty$.

We impose the following conditions.

assumption[Factor Model] Suppose that $\text{rank}\left(\Gamma\right) \le R$, i.e., $\Gamma = \lambda f'$ for some $N \times R$ matrix $\lambda$ and some $T \times R$ matrix $f$, with probability one for all $P\in\mathcal{P}$ and the following conditions hold: \leavevmode \begin{enumerate}[(i)] • Write $W$ for $X,Z_1,\ldots, Z_K$ and $W\cdot \gamma=X\beta + \sum_{k=1}^KZ_k\delta_k$ where $\gamma=(\beta,\delta')'$. We assume that there exists $\underline s^2 > 0$ such that \begin{align*} \min_{\gamma \in \mathbb R^{K+1}: \left\Vert \gamma\right\Vert = 1}\frac{1}{NT} \sum_{r = 2 R + 1}^{\min\{N,T\}} s_r^2 (W \cdot \gamma) \ge \underline s^2 \end{align*} with probability approaching 1 uniformly over $P \in \mathcal P$; • $s_1(X) = \mathcal{O}_{\Theta,\mathcal{P}} \left(\sqrt{NT}\right)$, $s_1(Z_k) = \mathcal{O}_{\Theta,\mathcal{P}} \left(\sqrt{NT}\right)$ for $k \in \{1, \ldots, K\}$, and $s_1(U) \asymp_{\Theta,\mathcal{P}} \max \{\sqrt{N},\sqrt{T}\} $; • $\langle X, U \rangle_F = \mathcal{O}_{\Theta,\mathcal{P}} (\sqrt{NT})$ and $\langle Z_k, U \rangle_F = \mathcal{O}_{\Theta,\mathcal{P}} (\sqrt{NT})$ for $k \in \{1, \ldots, K\}$; • $\left(s_1(U) - s_{r}(U)\right)/s_1(U) = o_{\Theta,\mathcal{P}}(1)$ for any fixed positive integer $r$; • For any sequence of matrices $A=A_{N,T}(X,Z)$ that is a function of $X,Z_1,\ldots,Z_k$, we have $\langle A, U \rangle_F = \mathcal{O}_{\Theta,\mathcal{P}}(\|A\|_F)$. \end{enumerate}

Assumption (ref)(ref) is a generalized non-collinearity condition, which requires that there is enough variation in the regressors after concentrating out $2R$ arbitrary factors. It is closely related to Assumption A of bai2009panel, but our version here avoids mentioning the unobserved factor loadings. The same generalized non-collinearity assumption is imposed in MoonWeidner2015. The assumption would be violated if some linear combination $W \cdot \gamma$ of the covariates were to have rank smaller or equal to $2R$. In particular, “low-rank regressors” are ruled out by this condition. Intuitively, Assumption (ref)((ref)) holds provided that $X_{it}$ and $Z_{it}$ have non-collinear idiosyncratic components. This intuition is formalized in Appendix (ref), where we also show that, when $X$ is the only regressor in the model (i.e.,\ $K=0$), then Assumption (ref)(ref) and (ref) below already guarantee that Assumption (ref)((ref)) holds.

Assumption (ref)((ref)) places mild bounds on $X$ and $Z_k$. For example, if the second moments of $X_{it}$ are (uniformly) bounded, then $ \mathbb{E}[s_{1}(X)^2] \leq \mathbb{E}[\| X \|_F^2] = \sum_{i=1}^N \sum_{t=1}^T \mathbb{E}[X_{it}^2] = \mathcal O_{\Theta,\mathcal{P}}(NT), $ which, by Markov's inequality, implies $s_{1}(X) = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$. In addition Assumption (ref)((ref)) also places a rate restriction on $s_1(U)$ that will hold as long as $U_{it}$ does not exhibit too much dependence over $i$ and $t$. This rate for $s_1(U)$ is closely related to Assumption (ref)((ref)), which is discussed below. Assumption (ref)((ref)) again holds as long as $U_{it}$ does not exhibit too much dependence over $i$ and $t$, and is uncorrelated with $X_{it}$ and $Z_{it}$. Finally, Assumption (ref)((ref)) holds as long as $U$ is mean zero given $X$ and $Z$ and satisfies bounds on dependence and second moments.

Assumption (ref)((ref)) is a high level assumption on the first few singular values of $U$ (note that $r$ is fixed as $N$ and $T$ converge to infinity). The singular values of $U$ are the square roots of the eigenvalues of $UU'$. The random matrix theory literature shows that, if $U$ is an appropriate noise matrix, the largest few eigenvalues of $UU'$ converge to the Tracy-Widom law, after appropriate rescaling: if $N$ and $T$ grow at the same rate, then each of the largest eigenvalues of $UU'$ grows at rate $N$, while the gaps between them grow at rate $N^{1/3}$. Johnstone2001 establish the Tracy-Widom law for the largest eigenvalues of $UU'$, for the case of i.i.d.\ normal error $U_{it}$. The subsequent literature has shown the universality of this result for more general error distributions, see e.g.\ Soshnikov2002, pillai2012edge and yang2019edge.

We also place conditions on the matrix $X$ requiring that there is sufficient variation after controlling for individual effects and the additional covariates $Z$. The constant $c=c_{N,T}$ in the following assumption is the one that appears in our construction of $A^*_{b,c}$.

assumptionFor all $P\in\mathcal{P}$, there exists uniformly bounded $\pi=\pi_P$ and random matrices $H$ and $V$ such that $X=Z\cdot \pi + H + V$ and the following conditions hold: \begin{enumerate}[(i)] • $\left\Vert V\right\Vert_F \asymp_{\Theta,\mathcal{P}} \sqrt{NT}$, $s_1 (V) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$; • $\left\Vert H\right\Vert_F = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$ and $\langle H,V \rangle_F = \mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$; • $\|Z_k\|_F=\mathcal O_{\Theta,\mathcal{P}}(\sqrt{NT})$ and $\langle Z_k, V \rangle_F = \mathcal O_{\Theta,\mathcal{P}} (\sqrt{N T})$ for $k \in \{1, \ldots, K\}$; • $\left(\bf{ Z{}' Z}\right)^{-1} = \mathcal O_{\Theta,\mathcal{P}}\left(\frac{1}{NT}\right)$ where ${\bf Z} = [{\rm vec}(Z_1), \ldots, {\rm vec}(Z_K)]$; • $\max_{i,t} V_{it}^2 = o_{\Theta,\mathcal{P}} (NT c_{N,T})$ and $\max_{i,t} Z_{k,it}^2 = o_{\Theta,\mathcal{P}} \left( (NT)^2 c_{N,T}\right)$ for $k \in \{1,\ldots,K\}$. \end{enumerate}

Assumption (ref) uses a decomposition of $X_{it}$ that depends on an individual effect $H_{it}$ and a random variable $V_{it}$ that is approximately independent and uncorrelated with $Z_{1,it}\ldots,Z_{k,it}$ as well as being approximately uncorrelated with the individual effect $H_{it}$. Importantly, the individual effect $H_{it}$ can be arbitrarily correlated with $\Gamma_{it}$ and with the variables $Z_{k,it}$. Note also that we do not place any assumptions on the rank or nuclear norm of the matrix $H_{it}$.

Part ((ref)) holds under a tail bound on $V_{it}$ and $Z_{k,it}$. For example, if $V_{it}$ are (uniformly) sub-Gaussian then $\max_{i,t} V_{it}^2 = \mathcal O_{\Theta,\mathcal{P}} ( \log (N + T) )$, and the condition $\max_{i,t} V_{it}^2 = o_{\Theta,\mathcal{P}} (NT c_{N,T})$ is satisfied provided that $N T c_{N,T} / \log (N + T) \rightarrow \infty$. The only other requirement on $c_{N,T}$ is the requirement that $c_{N,T} \max\{N,T\}\to 0$ in Theorem (ref) below. Thus, our results allow for a range of choices of $c_{N,T}$.

Define $P_\lambda = \lambda (\lambda' \lambda)^+ \lambda$ where $M^+$ denotes the Moore–Penrose inverse of a matrix $M$.

theoremLet $\hat \Gamma_{\rm pre}$ be defined in Algorithm (ref), with the modification described in Remark (ref). Suppose that Assumption (ref) holds, and that Assumption (ref) holds as stated and with $Z_k$ and $X$ interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$. Then, for any $\varepsilon > 0$, Assumption (ref)(ref) holds with \begin{enumerate}[(i)] • $\widetilde \Gamma = \Gamma - \hat \Gamma_{\rm pre}$ and $\hat C = 3 R s_1 (\hat U_{\rm pre}) (1 + \varepsilon) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$; • $\widetilde \Gamma = \Gamma + P_\lambda U - \hat \Gamma_{\rm pre}$ and $\hat C = 2 R s_1 (\hat U_{\rm pre}) (1 + \varepsilon) = \mathcal O_{\Theta,\mathcal{P}} (\max \{\sqrt{N},\sqrt{T}\})$. \end{enumerate}

Theorem (ref) is the main novel technical result that allows us to construct a feasible CI. It provides an explicit bound on the nuclear norm error of our initial estimate. As we show in the proof of Theorem (ref) below, the term $\langle A, P_\lambda U\rangle_F$ is asymptotically negligible under our assumptions. Thus, redefining the target parameter to be $\Gamma + P_\lambda U$ instead of $\Gamma$ and using the bound in part (ii) of the theorem does not affect the construction of the CI. This leads to a shorter CI using the bound in part (ii) compared to using the bound in part (i). For this reason, we use the bound in part (ii) in the implementation described in Section (ref) and in our formal coverage results below.

We now turn to the rate of convergence of the debiased estimator and the coverage of the CI. The proofs of these theorems use the nuclear norm bounds in Theorem (ref).

theoremLet $\hat \beta=\hat\beta_{A^{*}_{{b^*},c}}$ be defined in Algorithm (ref), with the modification described in Remark (ref). Suppose that Assumption (ref) holds, and that Assumption (ref) holds as stated and with $Z_k$ and $X$ interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$. Then \begin{align*} \hat\beta-\beta=\mathcal{O}_{\Theta,\mathcal{P}}(1/\min\{N,T\}). \end{align*}

To obtain primitive conditions for a central limit theorem and asymptotic validity of the confidence interval, we impose that the errors are independent, but not necessarily identically distributed, conditional on $X$, $Z$ and $\Gamma$.

assumptionThere exist constants $\underline \sigma>0$ and $\eta>0$ such that, for all $P\in\mathcal{P}$, $U_{it}$ is independent over $i,t$ conditional on $W,\Gamma$ and, for all $i,t$, \begin{align*} \mathbb{E}_P [U_{it}|W,\Gamma]=0, \quad \mathbb{E}_P [U_{it}^2|W,\Gamma]>\underline\sigma^2, \quad \mathbb{E}_P [U_{it}^{4}|W,\Gamma]<1/\eta. \end{align*}
theoremLet $\hat \beta=\hat\beta_{A^{*}_{{b^*},c}}$ and $\hat C = 2Rs_1(\hat U_{\rm pre})(1+\varepsilon)$ be defined in Algorithm (ref), with the modification described in Remark (ref) for $c=c_{N,T}$ with $c_{N,T} \max \{N,T\}\to 0$. Suppose that Assumptions (ref)((ref))-((ref)) hold, and that Assumption (ref) holds as stated and with $Z_k$ and $X$ interchanged for each $k=1,\ldots,K$, for the given sequence $c=c_{N,T}$, and that Assumption (ref) holds. Let $\widehat{\operatorname{se}}^2=\sum_{i=1}^N\sum_{t=1}^T A_{it}^2\hat U_{it}^2$ where $A=A^{*}_{{b^*},c}$ and $\hat U_{it}$ is the residual from the least squares estimator. Then \begin{align*} \hat\beta-\beta=\mathcal{O}_{\Theta,\mathcal{P}}(1/\min\{N,T\}) \end{align*} and \begin{align*} \liminf \inf_{\theta\in\Theta,P\in\mathcal{P}} \mathbb{P}_{\theta,P}\left( \beta \in \left\{ \hat\beta \pm \left[ \overline{\operatorname{bias}}_{\hat C}(\hat\beta) + z_{1-\alpha/2}\widehat{\operatorname{se}} \right] \right\} \right) \ge 1-\alpha. \end{align*}

Strong factor case

Numerous studies on the estimation of panel regressions with unobserved factors assume that these factors are “strong” or “semi-strong”. This assumption implies that the unobserved error structure, $\Gamma_{it} + U_{it}$ in model (ref), viewed as an $N \times T$ matrix, contains an $R = \operatorname{rank}(\Gamma)$ factor component $\Gamma = \lambda f'$ with singular values that asymptotically diverge faster than the largest singular value, $s_1(U)$, of the idiosyncratic error part $U$. Specifically, as $N, T \rightarrow \infty$, the ratio $s_1(U) / s_R(\Gamma)$ approaches zero (at a certain rate) under the semi-strong (strong) factor assumptions. Both Pesaran2006estimation and bai2009panel impose conditions that imply strong factors in this sense, as do many subsequent papers.

A key motivation for the estimation approach in this paper is to avoid assuming strong factors, instead providing an inference method that remains uniformly valid regardless of factor strength. Nevertheless, it is natural to consider how our approach behaves when factors are, in fact, strong, if only to facilitate comparison with much of the existing literature. The following theorem, therefore, extends Theorem (ref) to accommodate the case of strong factors.

theoremLet $\hat \Gamma_{\rm pre}$ be defined in Algorithm (ref), with the modification described in Remark (ref). Suppose that the hypotheses of Theorem (ref) hold, and furthermore, assume that $s_1(U) / s_R(\Gamma)= o_{\Theta,\mathcal{P}} (1)$. Then Assumption (ref)(ref) holds with $\widetilde \Gamma = \Gamma + M_{\lambda} U P_{f} + P_{\lambda} U M_{f} - \hat \Gamma_{\rm pre}$ and $\hat C = o_{\Theta, \mathcal P} \left( s_1(U) \right) = o_{\Theta, \mathcal P} (\max \{\sqrt{N},\sqrt{T}\}) $, where $P_f = f (f' f)^+ f'$, $M_\lambda = \mathbb I_N - P_\lambda$, and $M_f = \mathbb I_T - P_f$.

Theorem (ref) additionally requires $s_1(U) / s_R(\Gamma)= o_P(1)$, i.e., that the factors are semi-strong or strong, and also that $R = \operatorname{rank}(\Gamma)$. This implies that $\hat \Gamma_{\rm pre}$ converges to $\Gamma + M_{\lambda} U P_{f} + P_{\lambda} U M_{f}$ in second leading order (see e.g.\ Lemma S.3 in the supplement to MoonWeidner2015). Theorem (ref) states that with this change of target matrix, the nuclear norm bound $\hat C$ is smaller than $ s_1(U) \asymp_{\Theta,\mathcal{P}} {(\max \{\sqrt{N},\sqrt{T}\})}.$ This implies that, when $N$ and $T$ grow at the same rate, the worst-case bias $\overline{\operatorname{bias}}_{\hat C}(\hat\beta)$ used in the construction of the CI in Theorem (ref) becomes negligible relative to the standard error $\widehat{\operatorname{se}}$. At the same time, the extra term $\langle A, M_{\lambda} U P_{f} + P_{\lambda} U M_{f} \rangle_F$ is asymptotically negligible by exactly the same arguments given in the proof of Theorem (ref) for the negligibility of the term $\langle A, P_\lambda U \rangle_F$. As a result, in the considered regime, the debiased estimator is asymptotically unbiased and the (non-bias) aware CI (ref) is asymptotically valid.

remarkTheorem (ref) shows that the bias term is asymptotically negligible when $N$ and $T$ grow at the same rate and all $R$ factors are strong. This justifies the CI ((ref)) discussed in Remark (ref) in the strong factor setting. More generally, in the case where $R_w$ factors are weak and $R_s=R-R_w$ factors are strong, we conjecture that Theorem (ref) could be extended to show that the bias-aware CI ((ref)) is valid with $\hat C = 2 R_w s_1 (\hat U_{\rm pre}) (1 + \varepsilon)$. In other words, we conjecture that the worst-case bias of our estimator only depends on the number of weak factors.

Comparison to other results in the literature

Our debiasing approach leads to the faster rate $\min \{N,T\}$ compared to the rate $\min \{\sqrt{N},\sqrt{T}\}$ for $\hat \beta_{\rm LS}$ (see, e.g., MoonWeidner2015). While our results appear to be the first to demonstrate a $\min\{N,T\}$ rate of convergence under the conditions above, recent papers have proposed estimators that use additional structure to construct estimators that achieve the same or better rates. chetverikov2022spectral impose a factor structure on $X$, which corresponds to imposing a low-rank assumption on the matrix $H$ in our Assumption (ref). They use this assumption to construct an estimator that, like ours, achieves a $\min\{N, T\}$ rate under weak factors. zhu2019well imposes homoskedastic and independent errors in addition to a factor structure on $X$, and shows that this allows for a faster $\sqrt{NT}$ rate of convergence, even under weak factors.

While robust to weak factors, our CI will be wider than a CI based on the strong factor asymptotics in bai2009panel. Ideally, one would like to form a CI that is adaptive to the strength of factors. Such a CI would be robust to weak factors, while being asymptotically equivalent to the CI in bai2009panel when factors are strong. However, as shown by zhu2019well, such an adaptive CI cannot be obtained, even if one imposes homoskedastic errors and additional structure on the covariate matrix $X$. Thus, while there may be some room for efficiency gains over our CI, one must allow for some increase in CI length relative to the CI in bai2009panel in order to allow for weak factors.

As discussed in the introduction, our debiasing approach is analogous to the approach to debiasing the LASSO taken in javanmard_confidence_2014 and, more broadly, other papers in the debiased LASSO literature such as belloni_inference_2014, van_de_geer_asymptotically_2014 and zhang_confidence_2014. Interestingly, this analogy extends to the rates of convergence in our asymptotic results. The debiased lasso applies to a high dimensional regression model with $s$ nonzero coefficients and $n$ observations. The resulting estimator has bias of order $s/n$, up to log terms, and variance $1/n$. Note that $s$ is the dimension of the constraint set for the unknown parameter, while $n$ is the total number of observations. In our setting, the debiased estimator has bias of order $\max\{N,T\}/(NT)$ and variance $1/(NT)$. The set of matrices $\Gamma$ with rank at most $R$ has dimension of order $\max\{N,T\}$ so, just as with the debiased lasso, the bias term is of the same order of magnitude as the ratio of the dimension of the constraint set to the total number of observations. In the debiased lasso setting, one can justify a CI that ignores bias by assuming that $s$ increases slowly enough relative to $n$ for the order $s/n$ bias term to be asymptotically negligible relative to the order $1/\sqrt{n}$ standard deviation term. Unfortunately, this cannot occur in our setting even if $R=1$, since the bias term is of order $\max\{N,T\}/(NT)$ which is always of at least the same order of magnitude as the standard deviation $1/\sqrt{NT}$. This necessitates our bias-aware approach.

Numerical Evidence

Simulation Study

We consider the following design:

align*[align* omitted — 150 chars of source]

where $\kappa_r$ controls the strength of factor $f_{tr}$, and $R$ stands for the number of factors. In addition, $\lambda_i$, $f_t$, $U_{it}$ and $V_{it}$ are all mutually independent across both $i$, $t$, and $(i,t)$, and

align*[align* omitted — 239 chars of source]

In the designs considered below, we fix $(\beta, \sigma_U^2, \sigma_V^2) = (0,1,1)$ and vary $N$, $T$, the number of factors $R$, and their strengths controlled by $\kappa_r$. The number of simulations in all of the considered designs is 5000. As before, we are interested in estimation of and inference on $\beta$.

In Tables (ref)-(ref), we report the bias, standard deviation, and rmse for the benchmark LS estimator of bai2009panel and for the proposed debiased estimator in various designs with 1 and 2 factors.\footnote{Note that the CCE estimator of Pesaran2006estimation would not work in these designs, regardless of whether the factors are strong or not, because the cross-sectional averages of $\lambda_{ir}$ equal zero.} We also report the size of the corresponding tests (with $5\%$ nominal size) and the average length of the CIs (with $95\%$ nominal coverage). For simplicity and for brevity of the reported results, we assume that the number of factors is known. We consider a case when the number of factors is overspecified in Appendix (ref).

The LS estimator is heavily biased and the associated tests and CIs are heavily size distorted unless all the factors are strong. At the same time, the proposed estimator effectively reduces the “weak factors” bias without inflating the variance. As a result, the potential efficiency gains from using the debiased estimator can be very large when there is a weak factor, especially for larger sample sizes (see Appendix (ref) for additional simulation results). Importantly, even if all the factors are strong, the debiased estimator performs comparably to the LS estimator.

When weak factors are present, the LS CIs can have zero coverage because they are (i) centered around the biased LS estimator and (ii) too short. Hence, the average length of the LS CIs is not a proper benchmark to compare the average length of the bias-aware CIs. To provide a relevant comparison, we also construct identification robust CIs by inverting the (absolute value of the) LS based t-statistic using appropriate identification robust critical values (instead of $z_{1-\alpha/2}$). Specifically, for a given design (here, for fixed $N$, $T$, and $R$), we (numerically) compute the least favorable (over $\kappa$) critical value for the absolute value of the t-statistic based on the LS estimator. We also construct analogous CIs by inverting the (absolute value of the) t-statistic based on the debiased estimator using the corresponding least favorable critical values. We refer to such CIs as the LS and debiased oracle CIs (because they are based on unknown design-specific least favorable critical values) and report their average length denoted by “length*” in the tables below.

Notice that the average length of the LS oracle CIs (the “length*” column under the LS heading) is at least comparable to but mostly significantly greater than the actual length of the bias-aware CIs (the “length” column under the debiased heading), especially for larger sample sizes (again, see Appendix (ref) for additional simulation results). Thus, the bias-aware CI outperforms the LS CI once one corrects the LS CI to compensate for its severe undercoverage.

Another important comparison is between the actual length and oracle length of the bias-aware CI. Throughout most of the designs, the oracle length of the bias-aware CI is slightly less than half the length of the actual bias-aware CI that we compute. This gives a bound on how conservative our CI is: our bias-aware critical value cannot be decreased by more than a factor of about two without sacrificing coverage in these Monte Carlos. There are two possible sources of this conservativeness: (1) the bound in Theorem (ref) may be conservative or (2) there may be some additional structure in the initial error or its correlation with the data that our nuclear norm debiasing method does not exploit. While further improving the CI using the proof techniques in this paper appears difficult, we cannot rule out these possibilities. On the other hand, it is possible that these Monte Carlos overstate the conservativeness of our bias-aware CI: there may be other DGPs for which our bias-aware critical value cannot be decreased without sacrificing coverage.

Despite the simplicity of the design considered in this section, the presented findings seem to be characteristic of more complicated and settings. Specifically, in Appendix (ref), we consider a design with an additional covariate and non-Gaussian, heteroskedastic, and serially correlated errors and establish qualitatively similar results regardless of whether the correct number of factors is known or overspecified.

Empirical Illustration

In this section, we illustrate the finite sample properties of the proposed estimator and confidence intervals in a numerical experiment calibrated to imitate an actual empirical setting. Specifically, we calibrate our experiment based on the seminal studies of the effects of unilateral divorce law reforms on the US divorce rates by Friedberg1998 and Wolfers2006, subsequently revisited by KimOka2014 and MoonWeidner2015 in the context of interactive fixed effects models.

For simplicity of the experiment, as a benchmark, we use the following static specification also considered in Friedberg1998 and Wolfers2006

align*[align* omitted — 91 chars of source]

where $Y_{it}$ denotes the annual divorce rate (per 1,000 persons) in state $i$ in year $t$, and $X_{it}$ is a dummy variable indicating if state $i$ had a unilateral divorce law in year $t$. Following Friedberg1998 and Wolfers2006, we also control for state-specific quadratic time trends and time effects.

We follow KimOka2014 and use their data to construct a balanced panel with $N = 48$ states and $T = 33$ years. As in MoonWeidner2015, first we profile out the individual trends and time effects from $Y_{it}$ and $X_{it}$ to form the projected model

align*[align* omitted — 70 chars of source]

and obtain the estimates $\hat \beta$ and $\hat \sigma_{U^\perp}^2$. We also extract the first principal component of the matrix of regressors $X^\perp$ denoted by $\Gamma^{X^\perp} = \lambda_i^{X^\perp} f_t^{X^\perp}$.

In our numerical experiment, we fix $X^\perp$, $\{\lambda_i^{X^\perp}\}_{i=1}^N$, and $\{f_t^{X^\perp}\}_{t=1}^T$, and consider the following DGP

align*[align* omitted — 115 chars of source]

where we introduce an additional factor $f_{t}^{X^\perp}$ and a parameter $\kappa$ controlling the strength of $f_{t}^{X^\perp}$. For every repetition, we draw $U_{it}^{\perp}$ as iid $N(0, \hat \sigma_U^2)$ and treat all the other parts of the DGP as fixed.

As before, we compare the LS estimator and inference performance with the proposed approach for various values of $\kappa$. Both approaches use the correctly specified number of factors $R = 1$. The results are based on 5,000 simulations and provided in Table (ref). We report the same statistics as in Section (ref).

The results are qualitatively similar to the results in Section (ref). The LS estimator is heavily biased when the factor is weak, and the standard tests and confidence intervals are severely size distorted. Compared to the LS estimator, the debiased estimator has a substantially smaller bias, standard deviation, and rmse when the factor is weak. It also performs competitively if the factor is strong. The LS CIs are much shorter than the bias-aware CIs but have very poor coverage. The oracle CIs based on the LS estimator have the correct coverage and are also considerably wider than the naive CIs and comparable with the bias-aware CIs. Again, the oracle CIs based on the debiased estimator are considerably shorter than the bias-aware CIs and LS oracle CIs, indicating that there is a potential scope for improvement.

Overall, our empirically calibrated simulation study shows that the presence of a weak factor can lead to poor performance of conventional estimators and inference procedures in an actual empirical setting. It also demonstrates that in such settings, the gains from using the debiased estimator could be substantial.

Finally, we also report estimation and inference results for the actual data set. For consistency with the numerical experiment above, we focus on the same single covariate $X_{it}$. In Appendix (ref), we also consider a specification with dynamic treatment effects as in Wolfers2006. Similarly to KimOka2014 and MoonWeidner2015, we estimate

align*[align* omitted — 125 chars of source]

for various values of $R$ using the LS and the debiased approaches and construct 95% CIs for $\beta$. As before, we first profile out the individual trends and time effects, and then use the residual outcomes and regressors as inputs for the LS and debiased estimators.

The results are provided in Table (ref). We construct three types of CIs based on the debiased estimator using $\hat C$ as in Remark (ref) with different values of $R_w$. The first one is constructed assuming that there are no weak factors among $R$ factors ($R_w = 0$), i.e., under the same assumption under which the standard LS CI is valid. This is the same CI as introduced in Remark (ref). In this application, these CIs are as short as or even shorter than the LS CIs. Thus, if the researcher wants to obtain shorter CIs at the cost of non-robustness to the potential presence of weak factors, they can still do that using our debiased estimator. As pointed out in Remark (ref) and documented in Section (ref), such CIs are still likely to have much better coverage than the LS ones when there is a weak factors since they are based on the debiased estimator.

The second type of CIs is constructed assuming that among $R$ factors there is up to one weak factor ($R_w = 1$). The corresponding bias-aware CIs are substantially wider than the non-robust ones. However, as the numerical experiment considered earlier in this sections suggests, this is how wide identification robust CIs appear to have to be in this setting. In the considered application, we find that the potential presence of one weak factor is likely to be sufficient to nullify the significance of the previously obtained non-robust estimates.

Finally, we also report our bias-aware CIs as in (ref) corresponding to $R_w = R$. These CIs are uniformly valid regardless of the strength of identification of the factors.

The results for a specification with dynamic treatment effects are qualitatively similar and provided in Appendix (ref).

table[table omitted — 3,455 chars of source]

\newgeometry{left=1.5cm,right=1.5cm,top=3cm,bottom=2cm}

landscape\begin{table}[H] \caption{Simulation results for the experiment in Section (ref), $N = 100$, $T = 50$, $R = 2$} \begin{scriptsize} \begin{center} \begin{tabular}{ l| c c c c c c c c c c| c c c c c c c c c c} & \multicolumn{10}{c}{LS} & \multicolumn{10}{c}{Debiased}\\ \diagbox[width=12mm,height=8mm]{$\kappa_1$}{$\kappa_2$}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}\\ \midrule \multicolumn{21}{c}{bias}\\ \midrule 0.00 & -0.000&0.016&0.031&0.035&0.018&0.008&0.004&0.002&0.001&0.000&-0.000&0.006&0.010&0.010&0.005&0.002&0.001&0.000&0.000&-0.000\\ 0.05 & 0.016&0.033&0.049&0.060&0.052&0.036&0.030&0.027&0.026&0.025&0.006&0.012&0.017&0.018&0.014&0.010&0.009&0.008&0.007&0.007\\ 0.10 & 0.031&0.049&0.066&0.080&0.083&0.067&0.057&0.051&0.050&0.049&0.010&0.017&0.022&0.025&0.022&0.018&0.015&0.014&0.014&0.013\\ 0.15 & 0.035&0.060&0.080&0.097&0.108&0.099&0.082&0.072&0.070&0.068&0.010&0.018&0.025&0.029&0.028&0.024&0.020&0.018&0.017&0.017\\ 0.20 & 0.018&0.052&0.083&0.108&0.123&0.114&0.089&0.068&0.064&0.060&0.005&0.014&0.022&0.028&0.029&0.024&0.018&0.014&0.013&0.012\\ 0.25 & 0.008&0.036&0.067&0.099&0.114&0.088&0.054&0.032&0.028&0.025&0.002&0.010&0.018&0.024&0.024&0.017&0.010&0.006&0.005&0.005\\ 0.30 & 0.004&0.030&0.057&0.082&0.089&0.054&0.027&0.015&0.012&0.010&0.001&0.009&0.015&0.020&0.018&0.010&0.005&0.003&0.002&0.002\\ 0.40 & 0.002&0.027&0.051&0.072&0.068&0.032&0.015&0.007&0.005&0.004&0.000&0.008&0.014&0.018&0.014&0.006&0.003&0.001&0.001&0.001\\ 0.50 & 0.001&0.026&0.050&0.070&0.064&0.028&0.012&0.005&0.004&0.002&0.000&0.007&0.014&0.017&0.013&0.005&0.002&0.001&0.001&0.000\\ 1.00 & 0.000&0.025&0.049&0.068&0.060&0.025&0.010&0.004&0.002&0.000&-0.000&0.007&0.013&0.017&0.012&0.005&0.002&0.001&0.000&-0.000\\ \midrule \multicolumn{21}{c}{std}\\ \midrule 0.00 & 0.009&0.009&0.012&0.019&0.019&0.013&0.011&0.011&0.010&0.010&0.013&0.013&0.014&0.015&0.015&0.014&0.014&0.014&0.014&0.014\\ 0.05 & 0.009&0.009&0.010&0.014&0.021&0.015&0.012&0.011&0.011&0.011&0.013&0.013&0.013&0.015&0.015&0.014&0.014&0.014&0.014&0.014\\ 0.10 & 0.012&0.010&0.010&0.012&0.019&0.020&0.014&0.013&0.013&0.013&0.014&0.013&0.014&0.015&0.016&0.015&0.015&0.014&0.014&0.014\\ 0.15 & 0.019&0.014&0.012&0.012&0.017&0.025&0.021&0.019&0.019&0.020&0.015&0.015&0.015&0.016&0.016&0.017&0.016&0.016&0.016&0.016\\ 0.20 & 0.019&0.021&0.019&0.017&0.025&0.043&0.045&0.040&0.039&0.039&0.015&0.015&0.016&0.016&0.018&0.020&0.020&0.019&0.019&0.019\\ 0.25 & 0.013&0.015&0.020&0.025&0.043&0.064&0.055&0.037&0.034&0.032&0.014&0.014&0.015&0.017&0.020&0.023&0.021&0.018&0.018&0.017\\ 0.30 & 0.011&0.012&0.014&0.021&0.045&0.055&0.036&0.021&0.019&0.019&0.014&0.014&0.015&0.016&0.020&0.021&0.018&0.016&0.016&0.016\\ 0.40 & 0.011&0.011&0.013&0.019&0.040&0.037&0.021&0.016&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.018&0.016&0.016&0.016&0.016\\ 0.50 & 0.010&0.011&0.013&0.019&0.039&0.034&0.019&0.015&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.018&0.016&0.016&0.015&0.015\\ 1.00 & 0.010&0.011&0.013&0.020&0.039&0.032&0.019&0.015&0.015&0.015&0.014&0.014&0.014&0.016&0.019&0.017&0.016&0.016&0.015&0.015\\ \midrule \multicolumn{21}{c}{rmse}\\ \midrule 0.00 & 0.009&0.019&0.033&0.039&0.026&0.015&0.012&0.011&0.010&0.010&0.013&0.014&0.017&0.018&0.016&0.014&0.014&0.014&0.014&0.014\\ 0.05 & 0.019&0.034&0.050&0.061&0.056&0.039&0.032&0.029&0.028&0.027&0.014&0.017&0.021&0.023&0.021&0.017&0.016&0.016&0.016&0.016\\ 0.10 & 0.033&0.050&0.066&0.081&0.085&0.070&0.058&0.053&0.051&0.050&0.017&0.021&0.026&0.029&0.027&0.023&0.021&0.020&0.020&0.020\\ 0.15 & 0.039&0.061&0.081&0.098&0.109&0.102&0.085&0.075&0.073&0.071&0.018&0.023&0.029&0.033&0.033&0.029&0.026&0.024&0.024&0.023\\ 0.20 & 0.026&0.056&0.085&0.109&0.125&0.122&0.099&0.079&0.075&0.071&0.016&0.021&0.027&0.033&0.034&0.032&0.027&0.024&0.023&0.022\\ 0.25 & 0.015&0.039&0.070&0.102&0.122&0.109&0.077&0.049&0.044&0.041&0.014&0.017&0.023&0.029&0.032&0.028&0.023&0.019&0.018&0.018\\ 0.30 & 0.012&0.032&0.058&0.085&0.099&0.077&0.045&0.026&0.023&0.021&0.014&0.016&0.021&0.026&0.027&0.023&0.018&0.016&0.016&0.016\\ 0.40 & 0.011&0.029&0.053&0.075&0.079&0.049&0.026&0.017&0.016&0.016&0.014&0.016&0.020&0.024&0.024&0.019&0.016&0.016&0.016&0.016\\ 0.50 & 0.010&0.028&0.051&0.073&0.075&0.044&0.023&0.016&0.015&0.015&0.014&0.016&0.020&0.024&0.023&0.018&0.016&0.016&0.016&0.015\\ 1.00 & 0.010&0.027&0.050&0.071&0.071&0.041&0.021&0.016&0.015&0.015&0.014&0.016&0.020&0.023&0.022&0.018&0.016&0.016&0.015&0.015\\ \bottomrule \bottomrule \end{tabular} \end{center} \end{scriptsize} \end{table} \begin{table}[H] \caption{Simulation results for the experiment in Section (ref), $N = 100$, $T = 50$, $R = 2$} \begin{scriptsize} \begin{center} \begin{tabular}{ l| c c c c c c c c c c| c c c c c c c c c c} & \multicolumn{10}{c}{LS} & \multicolumn{10}{c}{Debiased}\\ \diagbox[width=12mm,height=8mm]{$\kappa_1$}{$\kappa_2$}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}&{0.00}&{0.05}&{0.10}&{0.15}&{0.20}&{0.25}&{0.30}&{0.40}&{0.50}&{1.00}\\ \midrule \multicolumn{21}{c}{size}\\ \midrule 0.00 & 7.6&52.3&88.3&80.0&42.8&17.9&10.6&6.7&6.3&5.9&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.05 & 52.3&96.2&99.8&99.6&95.9&88.3&81.6&73.7&70.8&68.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.10 & 88.3&99.8&100.0&100.0&100.0&99.7&99.2&98.6&98.4&98.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.15 & 80.0&99.6&100.0&100.0&99.8&99.5&98.7&97.9&97.4&96.6&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.20 & 42.8&95.9&100.0&99.8&98.3&93.1&87.3&80.0&76.9&73.8&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.25 & 17.9&88.3&99.7&99.5&93.1&76.1&60.5&45.2&39.9&36.2&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.30 & 10.6&81.6&99.2&98.7&87.3&60.5&38.9&23.3&18.9&16.2&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.40 & 6.7&73.7&98.6&97.9&80.0&45.2&23.3&11.3&9.2&7.9&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 0.50 & 6.3&70.8&98.4&97.4&76.9&39.9&18.9&9.2&7.4&6.4&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ 1.00 & 5.9&68.0&98.0&96.6&73.8&36.2&16.2&7.9&6.4&5.7&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0&0.0\\ \midrule \multicolumn{21}{c}{length}\\ \midrule 0.00 & 0.031&0.031&0.032&0.034&0.037&0.038&0.039&0.039&0.039&0.039&0.257&0.257&0.258&0.260&0.261&0.262&0.262&0.262&0.262&0.262\\ 0.05 & 0.031&0.031&0.032&0.033&0.036&0.038&0.038&0.039&0.039&0.039&0.257&0.257&0.258&0.259&0.261&0.262&0.262&0.262&0.262&0.262\\ 0.10 & 0.032&0.032&0.032&0.032&0.034&0.037&0.039&0.039&0.039&0.039&0.258&0.258&0.258&0.260&0.261&0.262&0.263&0.263&0.263&0.263\\ 0.15 & 0.034&0.033&0.032&0.032&0.033&0.036&0.039&0.040&0.040&0.040&0.260&0.259&0.260&0.261&0.263&0.264&0.265&0.265&0.265&0.265\\ 0.20 & 0.037&0.036&0.034&0.033&0.034&0.037&0.042&0.045&0.045&0.046&0.261&0.261&0.261&0.263&0.265&0.266&0.267&0.267&0.268&0.268\\ 0.25 & 0.038&0.038&0.037&0.036&0.037&0.043&0.049&0.052&0.052&0.052&0.262&0.262&0.262&0.264&0.266&0.268&0.268&0.269&0.269&0.269\\ 0.30 & 0.039&0.038&0.039&0.039&0.042&0.049&0.053&0.054&0.054&0.054&0.262&0.262&0.263&0.265&0.267&0.268&0.269&0.269&0.269&0.269\\ 0.40 & 0.039&0.039&0.039&0.040&0.045&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.267&0.269&0.269&0.269&0.269&0.270\\ 0.50 & 0.039&0.039&0.039&0.040&0.045&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.268&0.269&0.269&0.269&0.270&0.270\\ 1.00 & 0.039&0.039&0.039&0.040&0.046&0.052&0.054&0.055&0.055&0.055&0.262&0.262&0.263&0.265&0.268&0.269&0.269&0.270&0.270&0.270\\ \midrule \multicolumn{21}{c}{length*}\\ \midrule 0.00 & 0.345&0.346&0.353&0.373&0.406&0.420&0.424&0.427&0.427&0.428&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.115\\ 0.05 & 0.346&0.346&0.349&0.360&0.391&0.416&0.424&0.427&0.428&0.429&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.115\\ 0.10 & 0.353&0.349&0.348&0.354&0.375&0.409&0.424&0.430&0.432&0.432&0.114&0.114&0.114&0.115&0.115&0.115&0.115&0.115&0.115&0.116\\ 0.15 & 0.373&0.360&0.354&0.354&0.365&0.399&0.428&0.441&0.443&0.445&0.115&0.115&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116\\ 0.20 & 0.406&0.391&0.375&0.365&0.371&0.410&0.459&0.491&0.497&0.503&0.115&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116\\ 0.25 & 0.420&0.416&0.409&0.399&0.410&0.475&0.536&0.568&0.573&0.576&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116&0.116\\ 0.30 & 0.424&0.424&0.424&0.428&0.459&0.536&0.579&0.595&0.598&0.599&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.116&0.116&0.117\\ 0.40 & 0.427&0.427&0.430&0.441&0.491&0.568&0.595&0.604&0.606&0.607&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117\\ 0.50 & 0.427&0.428&0.432&0.443&0.497&0.573&0.598&0.606&0.608&0.609&0.115&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117\\ 1.00 & 0.428&0.429&0.432&0.445&0.503&0.576&0.599&0.607&0.609&0.610&0.115&0.115&0.116&0.116&0.116&0.116&0.117&0.117&0.117&0.117\\ \bottomrule \bottomrule \end{tabular} \end{center} \end{scriptsize} \end{table}

\restoregeometry

table[table omitted — 1,543 chars of source]
table[table omitted — 911 chars of source]

{2pt}

thebibliography\bibitem[\citeauthoryear{Ahn and Horenstein}{Ahn and Horenstein}{2013}]{AhnHorenstein2013} Ahn, S. C. and A. R. Horenstein (2013). \newblock Eigenvalue ratio test for the number of factors. \newblock {\em Econometrica\/} {\em 81\/}(3), 1203--1227. \bibitem[\citeauthoryear{Ahn, Lee, and Schmidt}{Ahn, Lee and Schmidt}{2001}]{AhnLeeSchmidt2001} Ahn, S. C., Y. H. Lee, and P. Schmidt (2001). \newblock {GMM} estimation of linear panel data models with time-varying individual effects. \newblock {\em Journal of Econometrics\/} {\em 101\/}(2), 219--255. \bibitem[\citeauthoryear{Ahn, Lee, and Schmidt}{Ahn, Lee and Schmidt}{2013}]{AhnLeeSchmidt2013} Ahn, S. C., Y. H. Lee, and P. Schmidt (2013). \newblock Panel data models with multiple time-varying individual effects. \newblock {\em Journal of Econometrics\/} {\em 174\/}(1), 1--14. \bibitem[\citeauthoryear{Alidaee, Auerbach, and Leung}{Alidaee, Auerbach and Leung}{2020}]{alidaee2020recovering} Alidaee, H., E. Auerbach, and M. P. Leung (2020). \newblock Recovering network structure from aggregated relational data using penalized regression. \newblock {\em arXiv preprint arXiv:2001.06052\/}. \bibitem[\citeauthoryear{Andrews and Cheng}{Andrews and Cheng}{2012}]{andrews2012estimation} Andrews, D. W. and X. Cheng (2012). \newblock Estimation and inference with weak, semi-strong, and strong identification. \newblock {\em Econometrica\/} {\em 80\/}(5), 2153--2211. \bibitem[\citeauthoryear{Arellano}{Arellano}{1987}]{arellano1987computing} Arellano, M. (1987). \newblock Computing robust standard errors for within-groups estimators. \newblock {\em Oxford Bulletin of Economics & Statistics\/} {\em 49\/}(4). \bibitem[\citeauthoryear{Arkhangelsky, Athey, Hirshberg, Imbens, and Wager}{Arkhangelsky, Athey, Hirshberg, Imbens and Wager}{2021}]{arkhangelsky_synthetic_2021} Arkhangelsky, D., S. Athey, D. A. Hirshberg, G. W. Imbens, and S. Wager (2021, December). \newblock Synthetic {Difference}-in-{Differences}. \newblock {\em American Economic Review\/} {\em 111\/}(12), 4088--4118. \bibitem[\citeauthoryear{Armstrong and Koles{\'a}r}{Armstrong and Koles{\'a}r}{2018}]{armstrong2018optimal} Armstrong, T. B. and M. Koles{\'a}r (2018). \newblock Optimal inference in a class of regression models. \newblock {\em Econometrica\/} {\em 86\/}(2), 655--683. \bibitem[\citeauthoryear{Armstrong, Kolesár, and Kwon}{Armstrong, Kolesár and Kwon}{2020}]{armstrong_bias-aware_2020} Armstrong, T. B., M. Kolesár, and S. Kwon (2020). \newblock Bias-{Aware} {Inference} in {Regularized} {Regression} {Models}. \newblock {\em arXiv:2012.14823 [econ, stat]\/}. \bibitem[\citeauthoryear{Athey, Bayati, Doudchenko, Imbens, and Khosravi}{Athey, Bayati, Doudchenko, Imbens and Khosravi}{2021}]{athey_matrix_2021} Athey, S., M. Bayati, N. Doudchenko, G. Imbens, and K. Khosravi (2021). \newblock Matrix {Completion} {Methods} for {Causal} {Panel} {Data} {Models}. \newblock {\em Journal of the American Statistical Association\/} {\em 0\/}(0), 1--15. \bibitem[\citeauthoryear{Bai}{Bai}{2009}]{bai2009panel} Bai, J. (2009). \newblock Panel data models with interactive fixed effects. \newblock {\em Econometrica\/} {\em 77\/}(4), 1229--1279. \bibitem[\citeauthoryear{Bai and Ng}{Bai and Ng}{2002}]{BaiNg2002} Bai, J. and S. Ng (2002, January). \newblock Determining the number of factors in approximate factor models. \newblock {\em Econometrica\/} {\em 70\/}(1), 191--221. \bibitem[\citeauthoryear{Bai and Ng}{Bai and Ng}{2017}]{BaiNg2017} Bai, J. and S. Ng (2017). \newblock Principal components and regularized estimation of factor models. \newblock {\em arXiv preprint arXiv:1708.08137\/}. \bibitem[\citeauthoryear{Bai and Ng}{Bai and Ng}{2023}]{bai2023approximate} Bai, J. and S. Ng (2023). \newblock Approximate factor models with weaker loadings. \newblock {\em Journal of Econometrics\/} {\em 235\/}(2), 1893--1916. \bibitem[\citeauthoryear{Bai and Wang}{Bai and Wang}{2016}]{bai2016econometric} Bai, J. and P. Wang (2016). \newblock Econometric analysis of large factor models. \newblock {\em Annual Review of Economics\/} {\em 8}, 53--80. \bibitem[\citeauthoryear{Belloni, Chen, Madrid Padilla, and Wang}{Belloni, Chen, Madrid Padilla and Wang}{2023}]{belloni2019high} Belloni, A., M. Chen, O. H. Madrid Padilla, and Z. Wang (2023). \newblock High-dimensional latent panel quantile regression with an application to asset pricing. \newblock {\em The Annals of Statistics\/} {\em 51\/}(1), 96--121. \bibitem[\citeauthoryear{Belloni, Chernozhukov, and Hansen}{Belloni, Chernozhukov and Hansen}{2014}]{belloni_inference_2014} Belloni, A., V. Chernozhukov, and C. Hansen (2014). \newblock Inference on {Treatment} {Effects} after {Selection} among {High}-{Dimensional} {Controls}. \newblock {\em The Review of Economic Studies\/} {\em 81\/}(2), 608--650. \bibitem[\citeauthoryear{Beyhum and Gautier}{Beyhum and Gautier}{2019}]{beyhum2019square} Beyhum, J. and E. Gautier (2019). \newblock Square-root nuclear norm penalized estimator for panel data models with approximately low-rank unobserved heterogeneity. \newblock {\em arXiv preprint arXiv:1904.09192\/}. \bibitem[\citeauthoryear{Beyhum and Gautier}{Beyhum and Gautier}{2022}]{beyhum2022factor} Beyhum, J. and E. Gautier (2022). \newblock Factor and factor loading augmented estimators for panel regression with possibly nonstrong factors. \newblock {\em Journal of Business & Economic Statistics\/}, 1--12. \bibitem[\citeauthoryear{Billingsley}{Billingsley}{1995}]{billingsley1995probability} Billingsley, P. (1995). \newblock {\em Probability and measure\/} (3rd ed.). \newblock A Wiley-Interscience publication. John Wiley and Sons. \bibitem[\citeauthoryear{Bonhomme and Manresa}{Bonhomme and Manresa}{2015}]{bonhomme2015grouped} Bonhomme, S. and E. Manresa (2015). \newblock Grouped patterns of heterogeneity in panel data. \newblock {\em Econometrica\/} {\em 83\/}(3), 1147--1184. \bibitem[\citeauthoryear{Chamberlain and Moreira}{Chamberlain and Moreira}{2009}]{Chamberlain2009} Chamberlain, G. and M. J. Moreira (2009). \newblock Decision theory applied to a linear panel data model. \newblock {\em Econometrica\/} {\em 77\/}(1), 107--133. \bibitem[\citeauthoryear{Chernozhukov, Hansen, Liao, and Zhu}{Chernozhukov, Hansen, Liao and Zhu}{2019}]{chernozhukov2019inference} Chernozhukov, V., C. Hansen, Y. Liao, and Y. Zhu (2019). \newblock Inference for heterogeneous effects using low-rank estimation of factor slopes. \bibitem[\citeauthoryear{Chetverikov and Manresa}{Chetverikov and Manresa}{2022}]{chetverikov2022spectral} Chetverikov, D. and E. Manresa (2022). \newblock Spectral and post-spectral estimators for grouped panel data models. \newblock {\em arXiv preprint arXiv:2212.13324\/}. \bibitem[\citeauthoryear{Cox}{Cox}{2024}]{cox2024weak} Cox, G. F. (2024). \newblock Weak identification in low-dimensional factor models with one or two factors. \newblock {\em The Review of Economics and Statistics\/} {\em 03}, 1--45. \bibitem[\citeauthoryear{Donoho}{Donoho}{1994}]{donoho1994statistical} Donoho, D. L. (1994). \newblock Statistical estimation and optimal recovery. \newblock {\em The Annals of Statistics\/}, 238--270. \bibitem[\citeauthoryear{Fan and Liao}{Fan and Liao}{2022}]{fan2022learning} Fan, J. and Y. Liao (2022). \newblock Learning latent factors from diversified projections and its applications to over-estimated and weak factors. \newblock {\em Journal of the American Statistical Association\/} {\em 117\/}(538), 909--924. \bibitem[\citeauthoryear{Fan}{Fan}{1951}]{fan1951maximum} Fan, K. (1951). \newblock Maximum properties and inequalities for the eigenvalues of completely continuous operators. \newblock {\em Proceedings of the National Academy of Sciences\/} {\em 37\/}(11), 760--766. \bibitem[\citeauthoryear{Feng}{Feng}{2023}]{feng_2023} Feng, J. (2023). \newblock Nuclear norm regularized quantile regression with interactive fixed effects. \newblock {\em Econometric Theory\/}, 1–31. \bibitem[\citeauthoryear{Ferman and Pinto}{Ferman and Pinto}{2021}]{ferman_synthetic_2021} Ferman, B. and C. Pinto (2021). \newblock Synthetic controls with imperfect pretreatment fit. \newblock {\em Quantitative Economics\/} {\em 12\/}(4), 1197--1221. \bibitem[\citeauthoryear{Fern{\'a}ndez-Val, Freeman, and Weidner}{Fern{\'a}ndez-Val, Freeman and Weidner}{2021}]{fernandez2021low} Fern{\'a}ndez-Val, I., H. Freeman, and M. Weidner (2021). \newblock Low-rank approximations of nonseparable panel models. \newblock {\em The Econometrics Journal\/} {\em 24\/}(2), C40--C77. \bibitem[\citeauthoryear{Freeman and Weidner}{Freeman and Weidner}{2023}]{freeman2023linear} Freeman, H. and M. Weidner (2023). \newblock Linear panel regressions with two-way unobserved heterogeneity. \newblock {\em Journal of Econometrics\/} {\em 237\/}(1), 105498. \bibitem[\citeauthoryear{Friedberg}{Friedberg}{1998}]{Friedberg1998} Friedberg, L. (1998). \newblock Did unilateral divorce raise divorce rates? {E}vidence from panel data. \newblock {\em American Economic Review\/}, 608--627. \bibitem[\citeauthoryear{Geman}{Geman}{1980}]{geman1980limit} Geman, S. (1980). \newblock A limit theorem for the norm of random matrices. \newblock {\em The Annals of Probability\/} {\em 8\/}(2), 252--261. \bibitem[\citeauthoryear{Hansen}{Hansen}{2007}]{hansen2007asymptotic} Hansen, C. B. (2007). \newblock Asymptotic properties of a robust variance matrix estimator for panel data when t is large. \newblock {\em Journal of Econometrics\/} {\em 141\/}(2), 597--620. \bibitem[\citeauthoryear{Hastie, Tibshirani, and Wainwright}{Hastie, Tibshirani and Wainwright}{2015}]{Hastieetal2015} Hastie, T., R. Tibshirani, and M. Wainwright (2015). \newblock {\em Statistical learning with sparsity: the lasso and generalizations}. \newblock CRC press. \bibitem[\citeauthoryear{Higgins}{Higgins}{2021}]{higgins2021fixed} Higgins, A. (2021). \newblock Fixed {$T$} estimation of linear panel data models with interactive fixed effects. \newblock {\em arXiv preprint arXiv:2110.05579\/}. \bibitem[\citeauthoryear{Hirshberg and Wager}{Hirshberg and Wager}{2020}]{hirshberg_augmented_2020} Hirshberg, D. A. and S. Wager (2020). \newblock Augmented minimax linear estimation. \newblock arXiv: 1712.00038. \bibitem[\citeauthoryear{Holtz-Eakin, Newey, and Rosen}{Holtz-Eakin, Newey and Rosen}{1988}]{HoltzEakin-Newey-Rosen1988} Holtz-Eakin, D., W. Newey, and H. S. Rosen (1988). \newblock Estimating vector autoregressions with panel data. \newblock {\em Econometrica\/} {\em 56\/}(6), 1371--95. \bibitem[\citeauthoryear{Horn and Johnson}{Horn and Johnson}{2013}]{horn2013matrix} Horn, R. A. and C. R. Johnson (2013). \newblock {\em Matrix Analysis\/} (2 ed.). \newblock Cambridge University Press. \bibitem[\citeauthoryear{Ibragimov and Khas’minskii}{Ibragimov and Khas’minskii}{1985}]{ibragimov_nonparametric_1985} Ibragimov, I. and R. Khas’minskii (1985). \newblock On {Nonparametric} {Estimation} of the {Value} of a {Linear} {Functional} in {Gaussian} {White} {Noise}. \newblock {\em Theory of Probability & Its Applications\/} {\em 29\/}(1), 18--32. \bibitem[\citeauthoryear{Ishihara and Kitagawa}{Ishihara and Kitagawa}{2021}]{ishihara_evidence_2021} Ishihara, T. and T. Kitagawa (2021). \newblock Evidence {Aggregation} for {Treatment} {Choice}. \newblock {\em arXiv preprint arXiv:2108.06473\/}. \bibitem[\citeauthoryear{Javanmard and Montanari}{Javanmard and Montanari}{2014}]{javanmard_confidence_2014} Javanmard, A. and A. Montanari (2014). \newblock Confidence {Intervals} and {Hypothesis} {Testing} for {High}-{Dimensional} {Regression}. \newblock {\em Journal of Machine Learning Research\/} {\em 15\/}(82), 2869--2909. \bibitem[\citeauthoryear{Johnstone}{Johnstone}{2001}]{Johnstone2001} Johnstone, I. (2001). \newblock {On the distribution of the largest eigenvalue in principal components analysis}. \newblock {\em Annals of Statistics\/} {\em 29\/}(2), 295--327. \bibitem[\citeauthoryear{Juodis and Sarafidis}{Juodis and Sarafidis}{2018}]{juodis2018fixed} Juodis, A. and V. Sarafidis (2018). \newblock Fixed {$T$} dynamic panel data estimators with multifactor errors. \newblock {\em Econometric Reviews\/} {\em 37\/}(8), 893--929. \bibitem[\citeauthoryear{Juodis and Sarafidis}{Juodis and Sarafidis}{2022}]{juodis2022linear} Juodis, A. and V. Sarafidis (2022). \newblock A linear estimator for factor-augmented fixed-{$T$} panels with endogenous regressors. \newblock {\em Journal of Business & Economic Statistics\/} {\em 40\/}(1), 1--15. \bibitem[\citeauthoryear{Kiefer}{Kiefer}{1980}]{Kiefer1980} Kiefer, N. (1980). \newblock {A time series-cross section model with fixed effects with an intertemporal factor structure}. \newblock {\em Unpublished manuscript, Department of Economics, Cornell University\/}. \bibitem[\citeauthoryear{Kim and Oka}{Kim and Oka}{2014}]{KimOka2014} Kim, D. and T. Oka (2014). \newblock Divorce law reforms and divorce rates in the usa: An interactive fixed-effects approach. \newblock {\em Journal of Applied Econometrics\/}. \bibitem[\citeauthoryear{Ma, Su, and Zhang}{Ma, Su and Zhang}{2022}]{ma2022detecting} Ma, S., L. Su, and Y. Zhang (2022). \newblock Detecting latent communities in network formation models. \newblock {\em The Journal of Machine Learning Research\/} {\em 23\/}(1), 13971--14031. \bibitem[\citeauthoryear{Miao, Li, and Su}{Miao, Li and Su}{2020}]{miao2020panel} Miao, K., K. Li, and L. Su (2020). \newblock Panel threshold models with interactive fixed effects. \newblock {\em Journal of Econometrics\/} {\em 219\/}(1), 137--170. \bibitem[\citeauthoryear{Miao, Phillips, and Su}{Miao, Phillips and Su}{2023}]{miao2023high} Miao, K., P. C. Phillips, and L. Su (2023). \newblock High-dimensional vars with common factors. \newblock {\em Journal of Econometrics\/} {\em 233\/}(1), 155--183. \bibitem[\citeauthoryear{Moon and Weidner}{Moon and Weidner}{2015}]{MoonWeidner2015} Moon, H. R. and M. Weidner (2015). \newblock Linear regression for panel with unknown number of factors as interactive fixed effects. \newblock {\em Econometrica\/} {\em 83\/}(4), 1543--1579. \bibitem[\citeauthoryear{Moon and Weidner}{Moon and Weidner}{2018}]{moon2018nuclear} Moon, H. R. and M. Weidner (2018). \newblock Nuclear norm regularized estimation of panel regression models. \newblock {\em arXiv preprint arXiv:1810.10987\/}. \bibitem[\citeauthoryear{Noack and Rothe}{Noack and Rothe}{2024}]{noack_bias-aware_2024} Noack, C. and C. Rothe (2024). \newblock Bias-{Aware} {Inference} in {Fuzzy} {Regression} {Discontinuity} {Designs}. \newblock {\em Econometrica\/} {\em 92\/}(3), 687--711. \bibitem[\citeauthoryear{Onatski}{Onatski}{2010}]{Onatski2010} Onatski, A. (2010). \newblock Determining the number of factors from empirical distribution of eigenvalues. \newblock {\em The Review of Economics and Statistics\/} {\em 92\/}(4), 1004--1016. \bibitem[\citeauthoryear{Onatski}{Onatski}{2012}]{Onatski2012} Onatski, A. (2012). \newblock Asymptotics of the principal components estimator of large factor models with weakly influential factors. \newblock {\em Journal of Econometrics\/} {\em 168\/}(2), 244--258. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2006}]{Pesaran2006estimation} Pesaran, M. H. (2006). \newblock Estimation and inference in large heterogeneous panels with a multifactor error structure. \newblock {\em Econometrica\/} {\em 74\/}(4), 967--1012. \bibitem[\citeauthoryear{Pillai and Yin}{Pillai and Yin}{2012}]{pillai2012edge} Pillai, N. S. and J. Yin (2012). \newblock Edge universality of correlation matrices. \newblock {\em The Annals of Statistics\/}, 1737--1763. \bibitem[\citeauthoryear{Recht, Fazel, and Parrilo}{Recht, Fazel and Parrilo}{2010}]{RechtFazelParrilo2010} Recht, B., M. Fazel, and P. A. Parrilo (2010). \newblock Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization. \newblock {\em SIAM review\/} {\em 52\/}(3), 471--501. \bibitem[\citeauthoryear{Robertson and Sarafidis}{Robertson and Sarafidis}{2015}]{robertson2015iv} Robertson, D. and V. Sarafidis (2015). \newblock {IV} estimation of panels with factor residuals. \newblock {\em Journal of Econometrics\/} {\em 185\/}(2), 526--541. \bibitem[\citeauthoryear{Robinson}{Robinson}{1988}]{robinson_root-n-consistent_1988} Robinson, P. M. (1988). \newblock Root-{N}-{Consistent} {Semiparametric} {Regression}. \newblock {\em Econometrica\/} {\em 56\/}(4), 931--954. \bibitem[\citeauthoryear{Rohde and Tsybakov}{Rohde and Tsybakov}{2011}]{RohdeTsybakov2011} Rohde, A. and A. B. Tsybakov (2011). \newblock Estimation of high-dimensional low-rank matrices. \newblock {\em The Annals of Statistics\/} {\em 39\/}(2), 887--930. \bibitem[\citeauthoryear{Soshnikov}{Soshnikov}{2002}]{Soshnikov2002} Soshnikov, A. (2002). \newblock {A note on universality of the distribution of the largest eigenvalues in certain sample covariance matrices}. \newblock {\em Journal of Statistical Physics\/} {\em 108\/}(5), 1033--1056. \bibitem[\citeauthoryear{van de Geer, Bühlmann, Ritov, and Dezeure}{van de Geer, Bühlmann, Ritov and Dezeure}{2014}]{van_de_geer_asymptotically_2014} van de Geer, S., P. Bühlmann, Y. Ritov, and R. Dezeure (2014). \newblock On asymptotically optimal confidence regions and tests for high-dimensional models. \newblock {\em The Annals of Statistics\/} {\em 42\/}(3), 1166--1202. \bibitem[\citeauthoryear{Wang, Su, and Zhang}{Wang, Su and Zhang}{2022}]{wang2022low} Wang, Y., L. Su, and Y. Zhang (2022). \newblock Low-rank panel quantile regression: Estimation and inference. \newblock {\em arXiv preprint arXiv:2210.11062\/}. \bibitem[\citeauthoryear{Westerlund, Petrova, and Norkute}{Westerlund, Petrova and Norkute}{2019}]{westerlund2019cce} Westerlund, J., Y. Petrova, and M. Norkute (2019). \newblock {CCE} in fixed-{$T$} panels. \newblock {\em Journal of Applied Econometrics\/} {\em 34\/}(5), 746--761. \bibitem[\citeauthoryear{Wolfers}{Wolfers}{2006}]{Wolfers2006} Wolfers, J. (2006). \newblock Did unilateral divorce laws raise divorce rates? a reconciliation and new results. \newblock {\em American Economic Review\/} {\em 96}, 1802--1820. \bibitem[\citeauthoryear{Yang}{Yang}{2019}]{yang2019edge} Yang, F. (2019). \newblock Edge universality of separable covariance matrices. \newblock {\em Electron. J. Probab\/} {\em 24\/}(123), 1--57. \bibitem[\citeauthoryear{Yata}{Yata}{2021}]{yata_optimal_2021} Yata, K. (2021). \newblock Optimal decision rules under partial identification. \newblock {\em arXiv preprint arXiv:2111.04926\/}. \bibitem[\citeauthoryear{Zeleneev}{Zeleneev}{2019}]{zeleneev2019identification} Zeleneev, A. (2019). \newblock Identification and estimation of network models with nonparametric unobserved heterogeneity. \newblock {\em Department of Economics, Princeton University\/}. \bibitem[\citeauthoryear{Zhang and Zhang}{Zhang and Zhang}{2014}]{zhang_confidence_2014} Zhang, C.-H. and S. S. Zhang (2014). \newblock Confidence intervals for low dimensional parameters in high dimensional linear models. \newblock {\em Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/} {\em 76\/}(1), 217--242. \bibitem[\citeauthoryear{Zhu}{Zhu}{2019}]{zhu2019well} Zhu, Y. (2019). \newblock How well can we learn large factor models without assuming strong factors? \newblock {\em arXiv preprint arXiv:1910.10382\/}.