EconBase
← Back to paper

Fixed-order PCA: Theory for Overestimated Factor Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,039 characters · 13 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Fixed-order PCA: Theory for Overestimated Factor Models

\onehalfspacing

abstractWe develop asymptotic theory for principal component analysis (PCA) of a high-dimensional factor model in which the working dimension $R$ is fixed and only required to satisfy $R \ge r$, where $r$ is the true number of factors. Building on anisotropic local laws from random matrix theory, we show that the “extra” empirical eigencomponents beyond the $r$-th are asymptotically noise-governed, incoherent, and nearly orthogonal to the factor loadings. We introduce two rotations, an expanded $r\times R$ map $H'$ and a compressed $R\times r$ map $H^{+}$, and establish consistency of the estimated factors under both. As an application, we analyze a factor-augmented regression for treatment-effect inference and prove $\sqrt{T}$-asymptotic normality for every fixed $R \ge r$. These results provide a theoretical underpinning for the common empirical practice of adopting a conservative upper bound on the number of factors, and shift the analytical burden from consistent dimension selection to the milder requirement of bounding $r$ from above. {\bf Key words:} factor model, fixed-$R$ PCA, treatment effect, overestimated rank

\onehalfspacing

Introduction

Factor models provide a parsimonious description of high-dimensional data by decomposing an observation matrix $X \in \mathbb{R}^{N\times T}$ as

equation[equation omitted — 51 chars of source]

where $B \in \mathbb{R}^{N\times r}$ is the loading matrix, $F \in \mathbb{R}^{T\times r}$ collects the latent factors, and $U \in \mathbb{R}^{N\times T}$ collect the idiosyncratic errors. Such models have become central in economics, finance, genomics, and signal processing, and principal component analysis (PCA) is by far the most widely used estimator. The asymptotic theory developed by {connor1986performance,bai03} and others typically rests on a critical prerequisite: the investigator must first consistently estimate the true number of factors $r$. Under strong-factor assumptions and clear eigenvalue gaps, a variety of methods, such as the information-criteria-based procedure BN02 and the eigenvalue-ratio test AH, can consistently recover $r$. In practice, however, the signal is rarely so clean. In such settings, these methods can be unstable and may fail to detect weak factors while also selecting spurious ones. As a result, downstream inference may be adversely affected, particularly when the number of factors is underestimated.

Motivated by this fragility, empirical researchers often proceed by selecting a working number of factors $R$ that is intended to exceed the true $r$, and then base estimation and inference on this potentially over-specified model. Variants of this approach appear widely in the applied literature, with $R$ typically chosen by rule of thumb or guided by domain knowledge. Despite its prevalence, the statistical properties of this practice remain largely uncharacterized. This paper provides a theoretical foundation for such over-specification.

Contributions

We develop a unified spectral theory for fixed-order PCA, wherein the working dimension $R$ is an arbitrary, user-specified integer satisfying $R \ge r$. Our contributions are threefold.

(1) Spectrum and eigenvectors beyond the true factor dimension. Building on anisotropic local laws from random matrix theory knowles2017anisotropic, we characterize the empirical singular values $\lambda_{r+1}(X), \ldots, \lambda_{R}(X)$ and their associated singular vectors. We show that the “extra” singular values track those of the noise matrix $U$ up to a sharp remainder, and that the overestimated singular vectors are incoherent (also known as localized) with respect to any direction independent of $U$ and are near-orthogonal to the factor space. {Our random-matrix-based analysis achieves a strictly sharper alignment rate between the overestimated principal-component directions and the factor loadings than the standard Davis--Kahan / $\sin\Theta$ benchmark.} These facts provide the technical basis for the rest of the paper.

(2) Estimation of the full factor space. Because the fitted factor matrix $\widehat F$ has $R$ columns rather than $r$, the usual notion of a rotation matrix must be extended. We introduce two rotations, an expanded $r \times R$ rotation $H$ and a compressed $R \times r$ generalized inverse $H^{+}$, and establish convergence rates for $\widehat F - FH'$ and $\widehat F H^{+} - F$. The former is faster because it only requires $F$ to lie in the column space of $\widehat F$, whereas the latter must further extract each of the $r$ factor directions from the $R$-dimensional fitted space.

(3) Robust factor-augmented inference. As an illustrative application, we study treatment-effect estimation in a factor-augmented regression, where the number of latent confounders is unknown and consistent selection of $r$ is not invoked. We show that $\sqrt{T}$-asymptotic normality of the resulting estimator persists for every fixed $R \ge r$ when $r\ge 1$. Adding extra principal components may increase residual variation but does not affect the signal at first order, because the extra components are asymptotically orthogonal to the factor space and, since $R - r$ is fixed, their inclusion inflates the variance by {an asymptotically negligible factor of $1+O(1/T)$}.

These results imply that, within the asymptotic regime considered here, overestimating $r$ by a bounded amount is {asymptotically harmless}, whereas underestimation can have first-order consequences. Fixed-order PCA therefore provides a conservative and transparent procedure.

Related Literature

Robustness to overspecification of the factor number. The consequences of overspecifying factor dimension have been examined in several settings. MW11 showed that inference in interactive-effects models remains valid when the number of factors is overspecified, using perturbation theory for linear operators Kato95. They established conditions under which the inclusion of overestimated eigenvectors does not affect the first order behavior of the parameter of interest. Their operator-perturbation route is more flexible across estimator families, while the random-matrix-based analysis that we adopt in this paper specializes to PCA but yields sharper rates when applicable (e.g., Theorem (ref)(iii)). In addition, we allow the factors to be much weaker in the sense that the top singular values of $X$ can grow much slower than $NT$, which is a strictly broader regime than the strong-factor scaling employed in MW11. barigozzi2020consistent established consistency of the low-rank component under overestimation through a trimmed PCA estimator that enforces incoherence by post-processing. fan2022learning and choi2023inference propose diversified-projection estimators that are robust to overspecification; their approach requires prespecified weights that are independent of the idiosyncratic noise, typically obtained from external characteristics or a held-out sub-sample. Our analysis shows that, under the setting of (ref) together with the incoherence conditions of knowles2017anisotropic, vanilla PCA is itself robust to overspecification, with no trimming, weighting, or external information required; the incoherence that those papers manufacture by post-processing or by external weights is here a generic property of noise-side singular vectors, supplied by the local law.

Entrywise eigenvector perturbation. A parallel methodological line fan2018ell, cape2019two, abbe2020entrywise,fan2021recent develops entrywise (sup-norm) perturbation bounds for the leading singular subspace of low-rank, incoherent signal matrices, sharpening the Davis--Kahan $\sin\Theta$ rate by exploiting incoherence of the signal. Importantly, their results only apply to the first $r$ eigenvectors. In contrast, the setting $R>r$ requires a different apparatus: the empirical eigencomponents that govern overspecification lie inside the noise bulk, where no spectral gap is available, so the deterministic perturbation arguments developed in this literature for the leading subspace do not directly apply. We draw instead on the anisotropic local law of knowles2017anisotropic, which supplies the analogous entrywise control of noise singular vectors in this no-gap regime; this random-matrix-based control is the natural counterpart of the entrywise perturbation bounds derived in the cited works, and is what enables our finer analysis of the extra components.

Spectral methods with weak or mixed signals. More broadly, our results complement a growing literature on the spectrum of factor and spiked-covariance models with weak or mixed signal strengths onatski2010determining, onatski2012asymptotics, wang2017asymptotics, freyaldenhoven2022factor, uematsu2022estimation, bai2023approximate, giglio2023prediction. Our theoretical results are developed under a weak-factor regime, in which the signal strength is required only to diverge with the sample size and is not constrained to the strong-factor scaling of bai03; this places our analysis squarely within the asymptotic setting considered by these works and makes the resulting spectral characterizations directly relevant to the inferential and prediction problems they study. Relative to that literature, we provide a single framework that simultaneously handles signal and noise eigencomponents in an overestimated spectrum, with explicit rates that readily plug into other factor-model inference problems. Our approach takes a different route than the factor-number selection BN02,onatski2010determining, ahn2013panel: rather than asking when these procedures recover $r$ exactly, we ask what inference can be performed given an arbitrary upper bound, which is the practically relevant object once the analyst recognizes that a sharp $\widehat r$ is unavailable.

Connection to double machine learning. The factor-augmented inferential application we develop in Section (ref) sits within the partialling-out / {double machine learning (DML)} tradition belloni2014inference, chernozhukov2016double, hansen2018fac, in which a low-dimensional treatment effect is estimated from a Neyman-orthogonal moment after residualizing both the outcome and the treatment with respect to high-dimensional nuisance components. Two features distinguish our setting. First, the nuisance is the latent factor space spanned by $F$, recovered by PCA from the panel $X$ rather than {constructed by selecting or fitting a function of observed controls; the two nuisances of canonical DML --- the conditional expectations of the outcome and the treatment given the controls --- share a common latent factor space in our setting, so a single PCA step suffices to residualize both equations, and the “double” structure of DML survives in the residualization rather than in the ML estimation step}; the standard {DML} “product of nuisance rates” condition is replaced by {a factor-strength condition stated explicitly in Assumption (ref)(ii)}. Second, {in the factor-augmented regression of Section (ref), a structural independence assumption between the regression errors and the factor--noise pair $(F,U)$} supplies the conditional independence that cross-fitting is engineered to deliver in the i.i.d.\ setting; consequently, neither sample splitting nor data-driven tuning of the working dimension $R$ is required, and any fixed $R\ge r$ with $R-r$ bounded yields valid $\sqrt{T}$ inference.

Organization and Notation

The remainder of the paper is organized as follows. Section (ref) introduces the model, fixed-order PCA, and our main spectral results. Section (ref) applies these results to inference on a treatment coefficient in a factor-augmented regression. Section (ref) reports simulation evidence. Section (ref) concludes. All proofs are deferred to the supplementary material.

Throughout the paper, $\lambda_k(A)$ denotes the $k$-th largest singular value of a matrix $A$, $\|A\|$ its spectral norm, and $\|A\|_{\mathrm{F}}$ its Frobenius norm. For a full-column-rank matrix $A$, $P_A = A(A'A)^{-1}A'$; for an arbitrary matrix, $A^{+}$ denotes its Moore--Penrose pseudoinverse and $P_A = A A^{+}$. We write $a_n \asymp b_n$ if $c\,b_n \le a_n \le C\,b_n$ for some $0 < c \le C < \infty$, and $a_n \ll b_n$ (equivalently, $a_n = o(b_n)$) if $a_n/b_n \to 0$.

We write $\Phi(T) \prec \phi(T)$ (stochastic domination) if for every $K>0$ and $\delta>0$ there exists $C>0$ such that ${\mathrm{P}}\{\Phi(T) > T^{\delta}\, C \phi(T)\} \le T^{-K}$ for all sufficiently large $T$. This is stronger than the standard $O_P$ notation, which we reserve for ordinary boundedness in probability. We use $c$ and $C$ for generic positive constants whose values may change across displays, and $\varepsilon$ for an arbitrarily small positive constant.

Fixed-Order PCA: Model and Theory

Factor Model Setup

We consider the static approximate factor model in which, for each time period $t = 1, \ldots, T$, we observe an $N \times 1$ vector {$x_t$} satisfying \[ {x_t} = B f_t + u_t. \] Here $B$ is an $N \times r$ matrix of factor loadings, $f_t$ is an $r \times 1$ vector of latent common factors, and $u_t$ is an $N \times 1$ vector of idiosyncratic components. We assume throughout that $u_t$ is independent of $f_t$ and that both have mean zero.

Let {$X = (x_1, \ldots, x_T)$}, $F = (f_1, \ldots, f_T)'$, and $U = (u_1, \ldots, u_T)$. Then $X$ admits the matrix representation \[ X = M + U, \qquad M = B F', \] where $M \in \mathbb{R}^{N \times T}$ is the low-rank signal matrix and $U \in \mathbb{R}^{N \times T}$ collects idiosyncratic errors.

We model idiosyncratic components as $u_t = \Sigma_e^{1/2} e_t$, where $e_t$ is an $N \times 1$ vector of standardized shocks and $\Sigma_e$ is an $N \times N$ positive-definite matrix. In matrix form, \[ U = \Sigma_e^{1/2} E, \qquad E = (e_1, \ldots, e_T), \] and we denote by $\sigma_1, \ldots, \sigma_N$ the eigenvalues of $\Sigma_e$. {The matrix $\Sigma_e$ is allowed to be a general positive definite matrix subject to the spectral regularity conditions in Assumption (ref)(ii) below: {its eigenvalues are bounded and the associated deformed Marchenko--Pastur law is regular.} Together these conditions accommodate cross-sectional heteroskedasticity and weak cross-sectional dependence in the idiosyncratic components, but rule out approximate factor-like structures within $u_t$ that would create outlying spikes in the spectrum of $\Sigma_e$; the latter would correspond to additional, undetected factors and is best modelled by enlarging $r$.}

{Two features of this setup deserve emphasis. First, $B$ is treated as deterministic and $F$ as random throughout; statements such as “$\max_{i \le N}\|b_i\|=O(1)$” below should be read as deterministic bounds uniform in $N$. Second, all dependence in $U$ is channeled through the deterministic operator $\Sigma_e^{1/2}$ acting on entries $e_{i,t}$ that are independent across both $i$ and $t$ (Assumption (ref)(i)). This is the standard random-matrix setting in which the anisotropic local law of knowles2017anisotropic applies, but it is more restrictive than the classical approximate-factor-model framework of bai03, which permits weak serial as well as cross-sectional dependence in $u_t$.} {We emphasize this trade-off explicitly: on one hand, both cross-sectional and serial independence are required by the local-law we use to characterize the spectral theory of the extra eigenvectors. On the other hand, a more general HAC-type extension to weakly serially-dependent $u_t$ is the most empirically pressing direction for future work, which we list accordingly in Section (ref). The applications for which our setup is best suited are therefore those in which independence across $(i,t)$ is plausible by sampling design --- repeated cross-sections, randomized rollouts, large-panel snapshots in survey data. Meanwhile, in this paper the factor process $f_t$ is indeed allowed to be serially dependent. }

Because $B$ and $F$ are identified only up to a rotation, we fix a canonical normalization tied to the singular vectors of $M$. Let $M = \Xi_r L_r V_r'$ denote the SVD of $BF'$, write $S_B = N^{-1}B'B$ and $S_f = T^{-1}F'F$, and let $H_B,H_F\in\mathbb{R}^{r\times r}$ be the rotation matrices that simultaneously diagonalize $S_B^{1/2}S_f S_B^{1/2}$ and $S_f^{1/2} S_B S_f^{1/2}$, normalized so that

equation[equation omitted — 169 chars of source]

The construction of $H_B$ and $H_F$ from $(S_B,S_f)$ and the verification of (ref) are entirely algebraic; the explicit formulas and computation are deferred to Section (ref). The identities in (ref) are the only consequences used in the body of the paper.

Fixed-Order PCA

Throughout, the true number of latent factors $r = \dim(f_t) \ge 0$ is assumed to be fixed and bounded. When factors are not strong (the singular values of $M$ grow slower than $\sqrt{NT}$), consistently estimating $r$ typically requires stringent technical conditions and performs poorly in finite samples; information criteria and eigenvalue-ratio methods are well documented to under- or over-estimate $r$ in finite sample, depending on signal strength.

To avoid this instability, we adopt fixed-order PCA. Rather than attempting consistent recovery of $r$, the statistician specifies an integer $R$ (the working number of factors) and imposes only \[ R \ge r. \] The estimator is then constructed from the first $R$ empirical singular components of $X$, regardless of whether $R$ matches the true factor dimension.

{The requirement $R\ge r$ is still substantive: one needs a credible upper bound on the factor dimension, supplied for example by a conservative screen from information criteria, or by scientific constraints on the number of latent channels. When no such upper bound is available, fixed-order PCA should be viewed as a sensitivity analysis over a small grid of plausible $R$'s rather than as a replacement for factor-number learning.}

{From a machine-learning perspective, the working dimension $R$ is the tuning parameter of fixed-order PCA. The main results of this section can therefore be read as a robustness-to-tuning-parameter statement: valid inference holds for any fixed $R$ within a range bounded below by the unknown $r$ and bounded above by a constant, without data-driven optimization of $R$.}

Write the singular value decomposition of $X$ as \[ X = \widehat \Xi_R \widehat L_R \widehat V_R' + \widehat \Xi_{-R} \widehat L_{-R} \widehat V_{-R}', \] where $\widehat \Xi_R \in \mathbb{R}^{N \times R}$, $\widehat V_R \in \mathbb{R}^{T \times R}$, and $\widehat L_R \in \mathbb{R}^{R \times R}$ collect the top-$R$ left singular vectors, right singular vectors, and singular values, and $(\widehat \Xi_{-R}, \widehat L_{-R}, \widehat V_{-R})$ collect the remaining components. The PCA estimators of loadings and factors are \[ \widehat B = \sqrt{N}\,\widehat \Xi_R, \qquad \widehat F = \frac{1}{\sqrt{N}}\, X' \widehat \Xi_R = \frac{1}{\sqrt{N}}\,\widehat V_R \widehat L_R, \] and the low-rank component is estimated by singular-value thresholding, \[ \widehat M = \widehat B \widehat F' = \widehat \Xi_R \widehat L_R \widehat V_R' = X \widehat V_R \widehat V_R'. \] In particular, $\widehat S_f := T^{-1} \widehat F'\widehat F = (TN)^{-1} \widehat L_R^{2}$; since the top $R$ singular values of $X$ are positive almost surely under Assumption (ref), $\widehat S_f$ is invertible (almost surely).

A central contribution of this paper is to show that fixed-order PCA still consistently recovers the factor space spanned by $B$ even when $R$ strictly exceeds $r$. The additional empirical eigencomponents (those beyond the true factor dimension) do not contaminate the recovery of the true factor space; instead, they converge to well-characterized noise-governed directions. The leading $r$ components of fixed-order PCA asymptotically reconstruct the true factor space, while the remaining $R-r$ components behave in a controlled and predictable manner.

Asymptotic Behavior of the Extra Eigencomponents

The main objective of this section is to establish the consistency of PCA when $R > r$. We partition the empirical eigencomponents as \[ \widehat \Xi_R = (\widehat\Xi_r, \widehat\Xi_{-r}), \qquad \widehat V_R = (\widehat V_r, \widehat V_{-r}), \] where $\widehat\Xi_r$ denotes the leading $r$ left singular vectors, the usual spiked eigenvectors, and $\widehat\Xi_{-r}$ collects the remaining $R-r$ eigenvectors, which we refer to as the extra (or overestimated) eigenvectors. When $R = r$, we set $\widehat\Xi_{-r} = \emptyset$. The analogous decomposition applies to $\widehat V_R$.

The behavior of the extra eigenvectors is the key obstacle: $\widehat\Xi_{-r}$ carries no factor signal, and its asymptotic effect depends delicately on the noise $U$. For PCA-based inference to remain valid under overspecification, $\widehat\Xi_{-r}$ must be incoherent, in the sense that its entries spread approximately uniformly across coordinates rather than concentrating on a few positions; coherent eigenvectors (also known as localized eigenvectors), in contrast, would interact spuriously with deterministic directions of interest. We establish this incoherency for $\widehat\Xi_{-r}$ as Theorem (ref)(ii) below.

assumptionThe idiosyncratic shocks and their cross-sectional covariance satisfy: \begin{enumerate} • \text The entries $e_{i,t}$ of $e_t$ are independent random variables with $\mathbb{E}[e_{i,t}] = 0$ and $\mathbb{E}[e_{i,t}^2] = 1$. In addition, $e_{i,t}$ is subexponential, in the sense that $\|e_{i,t}\|_{\psi_1}:=\inf\{K>0: \mathbb E\exp(|e_{i,t}|/K)\leq 2\}<\infty$. • The eigenvalues of $\Sigma_e$ lie in $[c,C]$ for constants $0<c\le C<\infty$ that do not depend on $N$, and their associate deformed Marchenko--Pastur (MP) has regular edges and bulks, as defined in Definition (ref) in Appendix (ref). \end{enumerate}

We require the noise be subexponential-tailed, as defined in vershynin2020high. Condition (ii) allows us to develop a local-law of the spectrum of $X$ for factor models. The MP law of $U$ can be determined using its Stieltjes transform. Assuming it has regular edges and bulks, we show that the analysis of $X$ can proceed by leveraging the eigenvalue rigidity, edge spacing, and anisotropic incoherence properties of $U$ established in knowles2017anisotropic. In part (ii) of this assumption, the required regularity on edges and bulks for the eigenvalues of $\Sigma_e$ are standard, and we defer the detailed definition to Definition (ref). Here are two examples to satisfy this condition. Let the ordered eigenvalues of $\Sigma_e$ be $\sigma_1\geq \sigma_2\geq...\geq\sigma_N$, and their empirical spectral distribution be $\pi = \frac{1}{N}\sum_{i=1}^N \delta_{\sigma_i}$, where $\delta_{\sigma_i}$ denotes the Dirac measure at $\sigma_i$.

enumerate• Discrete limit. The measure $\pi$ is supported on $K$ fixed values $\{s_k\}_{k=1}^K$, and $\pi(s_k)$ converges to a limit as $T \to \infty$. • Continuous limit. There exists a measure $\pi_\infty$ supported on $[a, b]$ whose density is bounded in $[\tau, \tau^{-1}]$ for some $\tau > 0$, and $\pi$ converges weakly to $\pi_\infty$ as $T \to \infty$.

Therefore, Assumption (ref)(ii) is satisfied in standard examples: the homoskedastic case $\Sigma_e=I\sigma$, sparse perturbations of the identity, and $\Sigma_e$ with bounded condition number whose limiting spectral measure has positive density.

assumptionThe loadings, factors, and signal strength satisfy: \begin{enumerate} • {The loadings $B$ are deterministic with row norms uniformly bounded in $N$, and the factors $F$ are random with row norms bounded in probability:} \[ \max_{i \le N} \|b_i\| {\le C}, \qquad \max_{t \le T} \|f_t\| = O_P(1). \]$N/T \to \phi$ for some $\phi \in (0, \infty)$ with $\phi \ne 1$. • There exists a sequence $\nu_M \to \infty$ with $\nu_M = O(\sqrt{N})$ such that \[ c \nu_M \sqrt{T} \le \lambda_r(M) < \cdots < \lambda_1(M) \le C \nu_M \sqrt{T}, \] and \[ c \le \lambda_r(S_f) \le \cdots \le \lambda_1(S_f) \le C, \qquad S_f := \frac{1}{T} F'F. \] \end{enumerate}

The restriction $\phi \ne 1$ is inherited from knowles2017anisotropic: the anisotropic local law guarantees incoherence of the singular vectors of $U$ only when the aspect ratio is bounded away from the MP edge. {We do not claim that the conclusions fail at or near $\phi=1$; rather, the available local-law input used here does not provide the required singular-vector incoherence there.} To our knowledge, the extent to which this is a proof artifact versus a genuine phase transition for incoherence remains open. Recent work on the MP edge in the regime $\phi\to 1$ (see, e.g., the line of research developed by L.\ Erd\H{o}s and collaborators on rigidity at the soft and hard edges) studied the rate at which $\phi$ converges to one, and suggested that the natural fluctuation scale changes from the bulk $T^{-1/2}$ to an edge scale of order $T^{-2/3}$; whether this scale is compatible with the entrywise bounds we require for Theorem (ref) is an interesting question for future work, and would be the natural route to closing the $\phi\ne 1$ gap.

Part (iii) is the factor-strength condition: equivalently, $$\lambda_k(B'B) \asymp \nu_M^{2}, \quad k = 1, \ldots, r.$$ The sequence $\nu_M$ indexes the signal strength and governs the sharpness of all subsequent results. At one extreme, $\nu_M \asymp \sqrt{N}$ recovers the strong-factor setting of bai03; at the other, the results remain informative as long as $\nu_M \to \infty$, which is considerably weaker.

{The role of $\nu_M$ across the main results is as follows. The spectral statements in Theorem (ref)(ii)--(iii) require only $\nu_M\to\infty$ for the first $r$ singular values to be separated from the noise, while the factor-space recovery of Theorem (ref) adds the mild $T^{\varepsilon}=o(\nu_M)$ requirement, which is satisfied by any polynomial rate for $\nu_M\to\infty$. Furthremore, the inference theorem strengthens this to the product-of-rates condition $\sqrt{T}=o(\nu_M^{2})$ in Assumption (ref)(ii), the analogue of the DML rate condition.}

The factor-strength part of Assumption (ref) is imposed only when $r\ge 1$. Meanwhile we allow a special case $r=0$, that is, there is no factor present so $X=U$. In this case, statements about extra components reduce to the corresponding noise-only local-law statements, which we state in Corollary (ref) below. This case is interesting in applications where statisticians are concerned about the impact of confounding factors but are not sure whether they are present. Therefore, allowing $r=0$ as a special case makes the PCA-based inference be also robust to whether confounding factors are present.

The following theorem is the analytical foundation of the paper. It shows that the extra spectrum closely resembles the spectrum of the noise matrix, and that the extra eigenvectors are incoherent and nearly orthogonal to the factor directions. The theorem is stated for $r\ge 1$, the setting of direct interest for factor-model inference, whereas the case $r=0$ is recorded as Corollary (ref) below.

In the theorem we use the notion of stochastic domination $\prec$ to describe the rate of convergence, which is slightly stronger than the notion of $O_P(\cdot)$. While the definition is given at the end of Section 1, roughly speaking $X_T \prec a_T$ means $X_T= O_P(T^{\varepsilon}a_T)$ for an arbitrarily small $\varepsilon>0.$

theorem[Extra spectrum, $R > r$] Under Assumptions (ref) and (ref) with $r\ge 1$, the bounds in (i)--(iii) below all hold uniformly in $k \in \{1, \ldots, R-r\}$ on a single high-probability event. \begin{enumerate} • Singular values. For any $R > r \ge 1$ and $k = 1, \ldots, R-r$, \[ \lambda_k(U)^2 \ge \lambda_{r+k}(X)^2, \qquad \lambda_k(U)^2 - \lambda_{r+k}(X)^2 \prec \nu_M^{-2}\, T. \] • Incoherence of extra eigenvectors. For any sequence of unit vectors $\{\eta_i, \zeta_t: i\leq N, t\leq T\}$, which are independent of $U$, $\|\eta_i\|=1$ and $\|\zeta_t\|=1$, $i=1...N$, $t=1,...,T$, \[ \max_{i\leq N}\|\eta_i' \widehat\Xi_{-r}\| \prec \nu_M^{-1}, \qquad \max_{t\leq T}\|\zeta_t' \widehat V_{-r}\| \prec \nu_M^{-1}. \] • Near-orthogonality with factors. \[ \|B\|_{\mathrm{F}}^{-1}\,\| B' \widehat\Xi_{-r}\| \prec \nu_M^{-2}, \qquad \|F\|_{\mathrm{F}}^{-1}\,\|F' \widehat V_{-r}\| \prec \nu_M^{-2}. \] \end{enumerate}

The theorem admits the following geometric interpretation. The signal matrix $M$ contributes $r$ singular values of order $\nu_M\sqrt{T}$, while the noise matrix $U$ has Marchenko--Pastur spectrum on the order of $\sqrt{T}$. When $\nu_M \gg 1$, {the empirical singular spectrum of $X$ exhibits the so-called “canonical outlier-plus-bulk pattern": $r$ spike outliers at scale $\nu_M\sqrt{T}$, well separated from a Marchenko--Pastur bulk concentrated at scale $\sqrt{T}$. They contain essentially all of the factor information in $X$. In addition, the next $R-r$ empirical singular vectors $\widehat\Xi_{-r}$ of $X$ lie in the orthogonal complement of this signal subspace and are therefore noise-driven; their geometry is governed by the local law for $U$. The contribution of the signal to these extra subspaces is a second-order effect, captured by the $\nu_M^{-2}$ rate in Theorem (ref)(iii).

Part (ii) of Theorem (ref) implies $\|\eta'\widehat\Xi_{-r}\| \overset{P}{\to} 0$ for every unit vector $\eta$ independent of $U$, at rate $\nu_M^{-1}$. This is precisely the incoherence condition: were $\widehat\Xi_{-r}$ sparsely concentrated, say at the first coordinate, then taking $\eta = (1, 0, \ldots, 0)'$ would produce $\eta' \widehat\Xi_{-r} = 1$, contradicting the conclusion. In other words, even though the extra eigenvectors carry no signal, they are diffuse rather than coherent.

This theorem leads to three statistical insights. First, result (ii) imply the $\ell_{\infty}$ perturbation bounds for $R>r$, by setting $\eta$ and $\zeta$ as canonical basis vectors: $$ \|\widehat\Xi_{-r}\|_{\infty}= \max_{k}\|\eta_k' \widehat\Xi_{-r}\| =O_P( \nu_M^{-1}) $$ where $\eta_k'=(0,..,0,1,0...)$, taking value one on its $k$ th element (similarly we have bounds for $\|\widehat V_{-r}\|_{\infty}$). The related results on deterministic $\ell_\infty$ perturbation arguments of fan2018ell and abbe2020entrywise only apply to the first $r$ eigenvectors, but not for $\widehat\Xi_{-r}$ or $\widehat V_{-r}$. The technical challenge in extending $R=r$ to $R>r$ is that the extra singular vectors correspond to singular values without clear separation from neighboring singular values, so the standard eigen-gap arguments do not apply to them. Our Theorem (ref)(ii) leverages the random-matrix local law to analyze the entrywise structure of the extra singular vectors.

Secondly, the improvement from $\nu_M^{-1}$ in (ii) to $\nu_M^{-2}$ in (iii) is genuine: the extra eigenvectors are driven by $U$ and thus nearly orthogonal to the factor space. The gap between the signal and noise singular values is what drives the fast rate of convergence $\nu_M^{-2}$. The rate is also strictly sharper than what a Davis--Kahan / $\sin\Theta$ argument would deliver. The standard $\sin\Theta$ bound for a perturbed singular subspace yields only $O_P(\nu_M^{-1})$, with no improvement for the eigenvectors orthogonal to factors. By contrast, an arbitrary $\eta$ in (ii) carries no a priori alignment with the signal block, so only the leading-order incoherence is available, which explains why the rate in (ii) is slower.

Lastly, each extra empirical singular vector, denoted by $\widehat\xi_k$ as the $k$th column of $\widehat \Xi_R$ for $k>r$, admits the orthogonal decomposition $$\widehat\xi_k = \Xi_r x_k + \Xi_c y_k$$ where the signal-aligned coefficient $x_k\in\mathbb{R}^r$ shrinks at the strictly faster rate $\|x_k\|=O_P(\nu_M^{-2})$ while the noise-aligned coefficient $y_k\in\mathbb{R}^{N-r}$ inherits the entrywise incoherency of the underlying noise singular vectors. The difference in the rates, $\nu_M^{-1}$ for arbitrary deterministic directions in (ii) versus $\nu_M^{-2}$ for the factor block in (iii), is the geometric reason why over-estimation in PCA is asymptotically benign and robust for inference.

As a statistical application, Theorem (ref) is especially useful for inference with over-estimated factors, as formalized below.

corollary[$r \ge 1$] Suppose the assumptions of Theorem (ref) hold. Let $G_N$ and $G_T$ be any $N \times K$ and $T \times K$ matrices with $K = O(1)$, $\|G_N\|_{\mathrm{F}} + \|G_T\|_{\mathrm{F}} = O_P(\sqrt{T})$, and $G_N, G_T$ independent of $U$. Then \[ \frac{1}{NT}\,\| \widehat B' U G_T\| + \frac{1}{NT}\,\| \widehat F' U' G_N\| =O_P( T^{-1/2}\, \nu_M^{-1}). \]

To see that the rate in Corollary (ref) is sharp, consider the strong-factor case $\nu_M = \sqrt{T}$. Then $(NT)^{-1}\|\widehat F' U G_N\| = O_P( T^{-1})$, which is the same rate as if the true factors were used $(NT)^{-1}\|F' U G_N\|= O_P(T^{-1})$. So the price of replacing the population factor by its PCA estimator does not introduce extra first order variance. Such a statistical corollary plays a central role in Section (ref) for factor-augmented inference.

Lastly, the corollary below specifies the special case that $r=0$:

corollary[Boundary case $r=0$] Suppose Assumption (ref) holds, $N/T\to\phi\ne 1$, and $r=0$, so that $X=U=\Sigma_e^{1/2}E$. Then for every bounded integer $R$ and any unit-norm vectors $\eta\in\mathbb{R}^N,\zeta\in\mathbb{R}^T$ independent of $U$, \[ \|\eta'\widehat\Xi_R\|\;\prec\; T^{-1/2},\qquad \|\zeta'\widehat V_R\|\;\prec\;T^{-1/2}. \]

Each of the leading $R$ empirical singular vectors of $X$ is therefore incoherent with respect to any direction independent of $U$ at the standard noise-only rate. The bound follows from a more general result, which we state as Proposition (ref) in the appendix: when $r=0$, the empirical singular vectors of $X$ coincide with those of $U$, and no factor-strength condition is needed.

Asymptotic Behavior of the Overestimated Factor Space

We next turn to estimation of the low-rank signal and the factor space.

theorem[Low-rank recovery, $r \ge 1$] Under the assumptions of Theorem (ref), for each fixed $R \ge r$ with $R-r$ bounded, \[ \frac{1}{\sqrt{NT}}\,\| \widehat M - M\|_{\mathrm{F}} = O_P\bigl(T^{-1/2}\bigr). \] {The bound is uniform over $R\in\{r,r+1,\ldots,\bar R\}$ for any fixed $\bar R\ge r$.}

Theorem (ref) shows that the low-rank component is recovered at the standard parametric rate, uniformly over any bounded working dimension $R \ge r$. Two features of this statement are worth noting. The rate $T^{-1/2}$ matches the optimal rate achievable by a hypothetical oracle estimator that knows $r$: overestimation by a bounded amount does not slow down the rate of convergence for low-rank recovery. The extra components $\widehat\Xi_{-r}\widehat L_{-r}\widehat V_{-r}'$ included in $\widehat M$ have Frobenius norm of order $\sqrt{T}$, but Theorem (ref)(iii) ensures that this contribution is aligned with the noise direction rather than the signal direction. While alternative estimators based on hard-thresholding, soft-thresholding, or nuclear-norm minimization can also achieve parametric recovery, fixed-order PCA does so without any tuning parameter beyond the integer $R$ itself; this is useful in settings where cross-validation is unstable or computationally expensive.

We now address estimation of the factor space itself. In the classical case $R = r$, PCA consistently recovers $F$ up to an invertible $r \times r$ rotation. When $R > r$, the notion of rotation must be generalized, and we introduce two natural extensions.

Expanded rotation. Define the $r \times R$ matrix \[ H' = \frac{1}{N} B' \widehat B. \] The model $X = B F' + U$ immediately yields

equation[equation omitted — 77 chars of source]

Up to the statistical error $N^{-1} U' \widehat B$, the true factors $F$ are expanded by $H'$ into the larger space spanned by $\widehat F$.

Compressed rotation. Let $H^{+} = (HH')^{+} H$ denote the $R \times r$ Moore--Penrose inverse of $H$. Post-multiplying (ref) by $H^{+}$ and using $H' H^{+} = I_r$, valid on the high-probability event $\{\lambda_r(H) > 0\}$, gives

equation[equation omitted — 88 chars of source]

Thus $\widehat F$ is compressed by $H^{+}$ back to an object of the true dimension.

{The two rotations serve different purposes. The expanded rotation $H$ is appropriate when the goal is to align the columns of $\widehat F$ with $F$ without forcing a dimension match; this formulation applies when $\widehat F$ enters downstream as a linear regressor, since regression projects onto a span and is invariant to the choice of basis within that span. Let $\mathrm{col}(F)$ denote the linear space spanned by the columns of $F$. The identity (ref) exhibits the geometric content of the expansion: the $r$-dimensional factor span $\mathrm{col}(F)$ is approximately contained in the $R$-dimensional fitted span $\mathrm{col}(\widehat F)$, with $H'$ mapping a basis of the former into the latter.

The compressed rotation $H^{+}$ is appropriate when the goal is to identify the $r$ true factor directions within the larger $R$-dimensional fitted space; this formulation applies when the downstream object is an $r$-column object explicitly referencing the true factor dimension rather than the working dimension $R$. The identity (ref) performs the inverse extraction of $\mathrm{col}(F)$ from inside $\mathrm{col}(\widehat F)$; The two notions coincide when $R = r$, in which case both reduce to the standard $r\times r$ rotation matrix of the classical PCA literature.}

The next proposition records the non-degeneracy scale of $H$ uniformly over $R \ge r$. We write $\nu_{\min}:=\lambda_r(S_B)$ when $r\ge 1$, so that $\nu_{\min}\asymp \nu_M^2/N$.

proposition[Non-degeneracy of the rotation] Under the assumptions of Theorem (ref) with $r \ge 1$, there exists a constant $c_0 > 0$ depending only on $(c, C, \phi)$ such that, for every bounded $R \ge r$, \[ {\mathrm{P}}\{\lambda_r(H) \ge c_0\nu_{\min}^{1/2}\} \to 1, \qquad \|H^{+}\| = O_P(\nu_{\min}^{-1/2}). \]

Under strong factors, $\nu_{\min}\asymp 1$, so Proposition (ref) recovers the familiar bounded-inverse behavior $\|H^+\|=O_P(1)$. Under weak factors, the inverse rotation carries the additional scale $\nu_{\min}^{-1/2}$; this scale slows down the compressed-rotation convergence rate, as stated below.

theorem[Factor space, $r \ge 1$] Suppose the assumptions of Theorem (ref) hold and there exists $\varepsilon > 0$ with $T^{\varepsilon} = o(\nu_M)$. Then for every bounded $R \ge r$: \begin{enumerate} • Expanded rotation. \[ \frac{1}{\sqrt{T}}\,\| \widehat F - F H'\| = O_P(T^{-1/2}), \qquad \frac{1}{T} \left\| F' (\widehat F - F H') \right\| \prec T^{-1/2}\, \nu_M^{-1}. \] • Compressed rotation. \[ \frac{1}{\sqrt{T}}\,\| \widehat F H^{+} - F\| = O_P(\nu_M^{-1}), \qquad \frac{1}{T}\,\| F'(\widehat F H^{+} - F)\| \prec \nu_M^{-2}. \] • Inverse factor covariance. \[ H' \left(\frac{1}{T} \widehat F' \widehat F\right)^{-1} H = \left(\frac{1}{T} F' F\right)^{-1} + o_P(1). \] \end{enumerate}

We use the $\nu_M$ form here for direct comparison with (ii) and bai03. When $\nu_M=\sqrt{T}$ (strong factors) both $\frac{1}{\sqrt{T}}\,\| \widehat F - F H'\|$ and $\frac{1}{\sqrt{T}}\,\| \widehat F H^{+} - F\| $ converge at $T^{-1/2}$. When $\nu_M$ is slower than $\sqrt{T}$ (weaker factors), the compressed rate in (ii) is slower than the expanded rate in (i), illustrating that recovering the true factor space is a more difficult problem than aligning the true factor space in the over-estimated space.

In addition, {Part (iii) of Theorem (ref) shows that the inverse estimated factor covariance, properly sandwiched by $H$, is consistent for the true inverse factor covariance regardless of the working dimension $R \ge r$. These inverse covariance matrices are often used for factor-augmented inference when estimated factors are being used as regressors. Hence, result (iii) is the property on which the robust factor-augmented inference of the next section relies.}

Application to Treatment-Effect Inference

{As an illustrative application of the spectral results in Section (ref), we consider inference on a treatment coefficient $\beta$ in the factor-augmented system}

eqnarray[eqnarray omitted — 199 chars of source]

where $g_t$ is the treatment variable of interest. We consider the application where $g_t$ is correlated with $\eta_t$ (so that it is endogenous). Then we include an observed instrumental variable (IV) $z_t$, and $\mathbb E( \eta_t|\varepsilon_{z,t}, f_t)=0$. The latent factors $f_t$ are not directly observed; instead we observe the high-dimensional panel of controls

equation[equation omitted — 51 chars of source]

exactly as in Section (ref), so that $f_t$ can be recovered by PCA from $X$. The intercepts $(\mu_y,\mu_g,\mu_z)$ are unrestricted nuisance parameters that absorb marginal means of $(y_t,g_t,z_t)$ at no asymptotic cost. The innovations $(\eta_t,\varepsilon_{g,t})$ are allowed to be mutually correlated within a period.

{This setup places the application within the partialling-out / Frisch--Waugh--Lovell / Neyman-orthogonal residualized-regression tradition of DML chernozhukov2016double, with the distinguishing feature that the high-dimensional nuisance is the latent factor space spanned by $F$, recovered by PCA from the panel $X$, rather than a function of observed controls. The new content of this section is to show that fixed-order PCA is a valid nuisance estimator, delivering $\sqrt{T}$-asymptotic normality with the unadjusted Eicker--White sandwich variance, without sample splitting or data-driven selection of $R$.

Let $P_A= A(A'A)^{-1}A'$ be the projection matrix of $A$. First, we apply PCA to $X$ to extract $R$ factors, whose $T\times R$ matrix is denoted by $\widehat F$, as in Section (ref). Then we use the factor-augmented IV estimator of $\beta$, which is based on the residualized just-identified IV regression:

equation[equation omitted — 159 chars of source]

with $\widehat\varepsilon_y = (I - P_{[1_T,\widehat F]}) Y$, $\widehat\varepsilon_g = (I - P_{[1_T,\widehat F]}) G$, and $\widehat\varepsilon_z = (I - P_{[1_T,\widehat F]}) Z$ denoting residuals from regressing the outcome {$Y=(y_1,\ldots,y_T)'$}, the treatment {$G=(g_1,\ldots,g_T)'$}, and the instrument {$Z=(z_1,\ldots,z_T)'$} on a constant and the PCA factor estimator. The projection on $[1_T,\widehat F]$ absorbs the unobserved intercepts together with the latent-factor nuisance $\rho'f_t$, and $\widehat\beta$ is the second-stage IV slope on the residualized variables.

The key innovation of our estimator is that our analysis does not require consistent estimation of the true factor number; it suffices that $R\geq r$ so that the span of $\widehat F$ cover the span of $F$. The key structural assumption is that $(\eta_t,\varepsilon_{g,t},\varepsilon_{z,t})$ is i.i.d.\ across $t$ and independent of $(F,U)$, formalized in Assumption (ref)(i) below. Under i.i.d.\ errors, the autocovariances of the score $\varepsilon_{z,t}\eta_t$ vanish, so the long-run variance reduces to the contemporaneous expectation and the heteroskedasticity-only Eicker--White (HC$_0$) sandwich is the natural variance estimator; the analysis allows arbitrary contemporaneous heteroskedasticity in the score, but not serial dependence or factor-driven conditional variance dynamics.

remark[OLS as a special case]In the special case that the treatment variable $g_t$ is uncorrelated with $\eta_t$, our method collapses to the residual based on OLS estimator by setting $z_t= g_t$. Then $\widehat\beta$ is similar to the OLS estimator analyzed in belloni2014inference, but with the high-dimensional Lasso step replaced by the PCA step.
assumptionThe regression and instrument errors and signal strength satisfy: \begin{enumerate} • $(\eta_t,\varepsilon_{g,t},\varepsilon_{z,t})$ is independent and identically distributed across $t$, independent of $(F,U)$, with $\mathbb{E}[\eta_t] = \mathbb{E}[\varepsilon_{g,t}] = \mathbb{E}[\varepsilon_{z,t}] = 0$, the exclusion restriction $\mathbb{E}[\eta_t\,\varepsilon_{z,t}] = 0$, and the relevance condition $\gamma := \mathbb{E}[\varepsilon_{g,t}\,\varepsilon_{z,t}] \ne 0$. • $\sqrt{T} = o(\nu_M^{2})$. • $\mathbb{E}|\eta_t|^{8} + \mathbb{E}|\varepsilon_{g,t}|^{8} + \mathbb{E}|\varepsilon_{z,t}|^8 \le C$, with $\mathbb{E}[\eta_t^2] > 0$ and $\mathbb{E}[\varepsilon_{z,t}^2] > 0$. \end{enumerate}

Condition (ii) is the factor-strength requirement relevant for inference; it is equivalent to $\lambda_r(B'B) \gg T^{1/2}$ and is weaker than the standard strong-factor assumption. {Condition (iii) is a uniform moment bound that supports both Lyapunov's central limit theorem for $\sqrt{T}(\widehat\beta-\beta)$ and consistency of the heteroskedasticity-robust sandwich variance estimator.} When $z_t = g_t$, the relevance condition $\gamma = \mathbb{E}[\varepsilon_{g,t}^2] \ne 0$ holds automatically and the exclusion $\mathbb{E}[\eta_t\,\varepsilon_{z,t}]=0$ becomes the OLS exogeneity $\mathbb{E}[\eta_t\,\varepsilon_{g,t}]=0$.

theoremSuppose Assumptions (ref), (ref), and (ref) hold, with the convention that Assumption (ref)(ii) imposes no condition when $r=0$ since $\nu_M$ is undefined in that case. Define the residual $\widehat\eta_t = \widehat\varepsilon_{y,t} - \widehat\beta\,\widehat\varepsilon_{g,t}$ and the Eicker--White (HC$_0$) variance estimator \[ \widehat\sigma^{2} \;=\; \bigl(\widehat\varepsilon_z'\widehat\varepsilon_g/T\bigr)^{-1}\,\Bigl(\tfrac{1}{T}\sum_{t=1}^{T}\widehat\varepsilon_{z,t}^{\,2}\,\widehat\eta_{t}^{\,2}\Bigr)\,\bigl(\widehat\varepsilon_z'\widehat\varepsilon_g/T\bigr)^{-1}. \] Then for every fixed (bounded) integer $R$ with $0\leq r \leq R$, \[ \sqrt{T}\,\widehat\sigma^{-1}(\widehat\beta - \beta) \;\overset{d}{\longrightarrow}\; \mathcal N(0, 1). \] {The estimator $\widehat\sigma^2$ is consistent for the just-identified IV sandwich variance \[ \sigma^2 \;=\; \gamma^{-2}\,\mathbb{E}[\varepsilon_{z,t}^{2}\eta_t^{2}],\qquad \gamma\;=\;\mathbb{E}[\varepsilon_{g,t}\,\varepsilon_{z,t}]. \] In the OLS specialization $z_t=g_t$, this reduces to $\sigma^2=\mathbb{E}[\varepsilon_{g,t}^2]^{-2}\,\mathbb{E}[\varepsilon_{g,t}^2\eta_t^2]$, the partialled-out OLS sandwich.}

We briefly discuss {the variance-vs-bias tradeoff behind Theorem (ref). Including $R-r$ extra principal components removes $R-r$ extra degrees of freedom from the residualized regression, creating a finite-sample degrees-of-freedom cost of the familiar order $T/(T-R-1)= {1+O(R/T)}$ {; this is asymptotically negligible for fixed $R$ but should be read as a finite-sample variance cost when $R$ is moderate relative to $T$, as documented in Section (ref)}. Overestimation does not induce first-order bias: by Theorem (ref)(iii), the overestimated principal directions are asymptotically orthogonal to the factor space, so they do not remove confounding signal. }

Moving on to the variance, note that if $F$ were observed, the score would be $T^{-1/2}\sum_t\varepsilon_{z,t}\eta_t$. Replacing $F$ by $\widehat F$ introduces additional terms involving the residual errors and the estimated projection $P_{\widehat F}$. Because $\widehat F$ is a function of $(F,U)$ and the regression errors are independent of $(F,U)$, these terms can be controlled conditionally on the factor-estimation sample. The residual bounds in the supplement show that the effect of estimating the factor space is $o_P(1)$ after $\sqrt{T}$ normalization whenever $\sqrt{T}=o(\nu_M^2)$. Thus the feasible residualized score has the same first-order limit as the infeasible score based on the true factor space. Therefore, the net effect is no first-order bias and only a mild finite-sample efficiency cost relative to the consequences of underestimation.

{Theorem (ref) does not prescribe how large $R$ should be in finite samples; we defer concrete guidance to Section (ref), where variance inflation as a function of $R/T$ is examined empirically. A working dimension chosen slightly above the value returned by an information criterion or an eigenvalue-ratio test is a natural conservative default. In addition, we allow $R>r=0$ as a special case, which is the case when it is an unknown matter of fact to statisticians that there are no confounding factors. The result shows that it does not hurt to estimate $R>0$ “factors" in this case. }

remark[Connection to related work] {In canonical DML, cross-fitting is used to weaken the dependence between the estimated nuisance and the score evaluated on the same observations. Here that dependence is structurally absent: $\widehat F$ is a measurable function of $(F,U)$. By Assumption (ref), $(F,U)$, and therefore $\widehat F$, are independent of the regression and instrument errors, so no sample splitting is needed.} {The closest non-DML antecedent is MW11, who establish robustness-to-overspecification for interactive-fixed-effects panel models via Kato perturbation of a profile-likelihood objective. Two further contrasts are worth noting beyond the rate comparison given in Section (ref). First, the operator-perturbation route is more flexible across estimator families (e.g.\ quasi-MLE, GMM), while the local-law route specializes to PCA but yields sharper rates when applicable. Second, the variance-sandwich consistency in Theorem (ref)(iii) that underlies our unadjusted-sandwich CLT does not appear to be available from operator-perturbation alone.}

Simulations

We report Monte Carlo evidence on the finite-sample behavior of the fixed-order PCA inference of Section (ref), {focusing on the OLS specialization $z_t=g_t$ (Remark (ref))}. The data-generating process has $i = 1, \ldots, N$ and $t = 1, \ldots, T$, intercepts $\mu_g = 2$ and $\mu_y = 3$, and $\beta = 0$. Loadings are generated sparsely: \[ b_{i,k} \sim \mathcal N(0, 1) \text{ with probability } p_N = N^{-\alpha}, \qquad b_{i,k} = 0 \text{ otherwise}, \] so a non-zero $\alpha\in[0,1)$ controls the strength of the latent factors: at $\alpha = 0$ the loadings are dense and $\nu_M\asymp\sqrt N$ (the classical strong-factor case of bai03); as $\alpha$ grows the non-zero proportion $N^{-\alpha}$ shrinks and $\nu_M\asymp N^{(1-\alpha)/2}$ approaches the boundary $\nu_M\to\infty$ admitted by Assumption (ref)(iii). Idiosyncratic errors are $u_t = \Sigma_e^{1/2}e_t$ with $e_{i,t}$ i.i.d.\ $\mathcal N(0,1)$ across both $i$ and $t$, exactly as required by Assumption (ref)(i), and $\Sigma_e = \mathrm{diag}(D)$ with $D_{ii}\sim\mathrm{Uniform}(0.5,1.5)$. The regression errors $\varepsilon_{g,t}$, $\eta_t$, the factors $f_{k,t}$, the loadings $\rho_k$, and $\alpha_{g,k}$ are all i.i.d.\ standard normal and mutually independent.

For each replication, $\widehat\beta$ is computed by partialling out $[1_T, \widehat F]$ from $y_t$ and $g_t$ and running OLS on the resulting residuals, with $\widehat F$ from the SVD of $X$ as defined in Section (ref); standard errors are Eicker--White HC$_0$. We report the standardized statistic \[ t_\beta = \sqrt{T}\,\frac{\widehat\beta - \beta}{\widehat\sigma}, \] and assess the closeness of its empirical distribution under the null $\beta = 0$ to the standard-normal benchmark implied by Theorem (ref). The tables below report, across $1{,}000$ Monte Carlo replications per cell, the sample mean and sample standard deviation (sd) of $t_\beta$, its $0.025$ and $0.975$ quantiles, and the Kolmogorov--Smirnov (KS) $p$-value against $\mathcal N(0,1)$. The Monte Carlo standard error of the sample sd at this resolution is roughly $1/\sqrt{2{,}000}\approx 0.022$.

\paragraph{Choice of $(N, T, \alpha)$} The theoretical conditions of Section (ref) impose three joint requirements: (a) Assumption (ref)(ii) requires $N/T \to \phi$ with $\phi \in (0, \infty)$ and $\phi \neq 1$, so $N$ and $T$ should grow proportionally and the aspect ratio should be bounded away from one; (b) Assumption (ref)(ii) imposes the rate condition $\sqrt{T} = o(\nu_M^{2})$, which under the sparse-loading design becomes $T = o(N^{2 - 2\alpha})$ and binds non-trivially when $\alpha$ is close to $1/2$; (c) the working dimension $R$ should be a fixed bounded overestimate of the true $r$. We therefore organize the evidence around three axes that map to (a)--(c), with $N = 200$ throughout:

itemize• Experiment 1 fixes $\alpha = 0$ (strong factors, so the rate condition is slack) and varies $T \in \{100, 400, 800\}$, giving $N/T \in \{2, 0.5, 0.25\}$, all bounded away from one. This isolates the role of the aspect ratio. • Experiment 2 fixes $T = 400$ (so $N/T = 0.5$, well inside the theory) and varies $\alpha \in \{0.0, 0.1, 0.2, 0.3, 0.4\}$, sweeping factor strength from the classical regime toward the rate boundary. • Experiment 3 sets $r = 0$ (no factors), so the inference reduces to the partialled-out boundary case, and varies $T \in \{100, 400, 800\}$.

\paragraph{Experiment 1: varying the aspect ratio $N/T$} Table (ref) reports the diagnostics for $R \in \{1, 2, 3, 6, 12, 30\}$, spanning under-specification ($R < r$), correct specification ($R = r$), and over-specification. The contrast across $R$ is sharp. When $R < r$, the sd of $t_\beta$ explodes to between $4$ and $13$ and KS rejects at every $T$. Once $R \ge r$, the sd drops to between $1.01$ and $1.13$ and KS $p$-values typically exceed $0.4$; at $T \in \{400, 800\}$ the standard-normal approximation is quite accurate across the whole $R \ge r$ range, and at $T = 100$ it degrades only once $R$ approaches $T/3$, consistent with the intuition of variance-inflation. The aspect-ratio behaviour is symmetric: the $T = 100$ ($N/T = 2$) and $T = 400$ ($N/T = 0.5$) columns behave equivalently, in line with the symmetry of Assumption (ref)(ii) about $\phi = 1$. Figure (ref) displays the corresponding histograms for $R \in \{2, 3, 12\}$, with the $R = 2$ column ($R < r$) wildly diffuse and most mass off the plotting window.

figure[figure omitted — 468 chars of source]
table[table omitted — 1,842 chars of source]

\paragraph{Experiment 2: varying the factor strength $\alpha$} Table (ref) reports the diagnostics across $\alpha$ and $R \in \{1, 2, 3, 6, 12\}$. As $\alpha$ grows, $\nu_M\asymp N^{(1-\alpha)/2}$ shrinks, so the rate condition $\sqrt T = o(\nu_M^2)$ becomes $T = o(N^{2-2\alpha})$; the rate-condition margin $N^{2-2\alpha}/T$ decays as $\alpha$ increases. The simulations track this scale sharply. Under-specified rows ($R \in \{1, 2\}$) behave as in Experiment 1, with sd of $t_\beta$ between $8$ and $10$ across all $\alpha$. Once $R \ge r$, the sd becomes a clean monotone function of $\alpha$: $1.07$--$1.08$ at $\alpha \in \{0.0, 0.1\}$ (rate condition slack, KS $p \gtrsim 0.25$), $1.20$ at $\alpha = 0.2$, $1.31$ at $\alpha = 0.3$, and $1.83$ at $\alpha = 0.4$, with $0.025$/$0.975$ quantiles widening to $\pm 3.7$. The role of the rate condition $\sqrt T = o(\nu_M^2)$ is therefore not a proof artifact: it sharply governs the finite-sample quality of the standard-normal approximation.

table[table omitted — 2,177 chars of source]

\paragraph{Experiment 3: boundary case ($r = 0$)} Theorem (ref) is stated for $r \ge 1$; Table (ref) shows that the same Gaussian behavior extends to the no-factor boundary. At $T \in \{400, 800\}$ the standard-normal approximation is essentially exact for every $R \in \{0, 3, 6, 12, 30\}$ (sd within $1.02$--$1.10$, KS $p \gtrsim 0.4$); at $T = 100$, the approximation deteriorates only when $R$ approaches $T/3$.

table[table omitted — 1,412 chars of source]

Two operational takeaways follow. Since $\nu_M$ is unobserved in practice, the leading singular values of $X$ are a natural diagnostic for the rate condition $\sqrt T = o(\nu_M^2)$: if they do not visibly separate from the bulk, the inflation in Table (ref) should be expected. And since under-specification is catastrophic while modest over-specification is essentially free, a natural conservative default is to choose $R$ slightly above the value returned by an information criterion or eigenvalue-ratio test, and to interpret unstable estimates across $R \ge \widehat r$ as evidence that the factor-strength regime is unfavorable rather than that $R$ is too small.

Empirical Application

We illustrate fixed-order PCA on a single cross-section from the Health and Retirement Study Sonnega2014HRS, the panel survey at the empirical core of the structural retirement literature including French2011HRS, with respondents treated as i.i.d.\ draws. The substantive question is whether labor supply at the intensive-and-extensive margin---the choice variable $N_t$ in the canonical life-cycle model of French2011HRS---affects depressive symptoms after controlling for a low-dimensional latent state of vitality, socioeconomic resources, and underlying mental-health propensity that drives both labor-supply choices and contemporaneous well-being. We apply the partialled-out residualized regression of Section (ref).

{The key statistical insight from this application (as shown by results presented in Table (ref)) is the robustness profile of the first-order PCA estimator across $R$. The estimator stays positive and within $\pm 35\%$ of the saturating $R=N$ value across $R\in\{2,\ldots,28\}$ --- exactly the robustness-to-tuning-parameter property the fixed-order PCA framework is designed to deliver. In contrast, a procedure that demands consistent recovery of $r$ would have to commit to a single $\widehat r$, whereas our framework certifies inference uniformly over any $R$ in the upper-$R$ tail.}

We now provide detailed implementation of this application. The data are from the RAND HRS Longitudinal File 1992--2022 v1.0, restricted to the wave-14 (2018) interview. After complete-case filtering, the sample has $T=14{,}672$ respondents. Treatment is annual hours worked, \[ g_t \;=\; (\text{hours/week})_t \times (\text{weeks/year})_t \;+\; (\text{2nd-job hours/week})_t \times (\text{2nd-job weeks/year})_t, \] clipped to $[0, 5{,}000]$ and set to zero for non-workers. The marginal distribution is bimodal: a $62.3\%$ atom at $g_t=0$ (non-workers, including fully retired respondents and other labor-market non-participants) and a continuous mass concentrated near full-time, with ${\mathrm{P}}(g_t \in [1{,}000, 2{,}200)) = 20.5\%$ and ${\mathrm{P}}(g_t \ge 2{,}200) = 10.5\%$. The continuous specification strictly nests the binary retirement indicator: it preserves the extensive margin (the $g_t=0$ atom) while resolving the intensive-margin variation among bridge-job holders, partial retirees, and full-time workers that the structural retirement literature treats as economically distinct states. The outcome $y_t = \mathrm{cesd}_t$ is the eight-item Center for Epidemiologic Studies Depression score, integer in $[0,8]$ with higher values indicating more depressive symptoms (sample mean $1.53$, sample standard deviation $2.02$). The control panel $x_t \in \mathbb R^N$ collects {$N=28$} standardized fundamentals, including sex, ethnicity, education, marital and veteran status, chronic-condition and smoking indicators, cognitive scores and household variables. Each measures a distinct dimension of the underlying biopsychosocial state.

A preliminary spectral decomposition of the standardized panel returns shows that there are no single dominant factor: the first three principal components capture roughly $27\%$ of the total variation.

{To bring the IV machinery of Section (ref) to bear, we use the institutional cutoff at age 62 as the instrument: \[ z_t \;=\; \mathbf{1}\{\text{respondent $t$ is at least 62 years old}\}. \] Age 62 is the early Social Security claiming age and the empirically dominant discontinuity in U.S.\ retirement hazards. We prefer it to the alternative cutoff at age 65 because it isolates the labor-supply / Social Security channel from the Medicare insurance channel that turns on at 65. Conditional on the latent state $f_t$, $z_t$ shifts labor supply primarily through this institutional channel rather than through individual mental-health propensity, supporting the exclusion restriction in Assumption (ref)(i). The first-stage relationship is unambiguously strong: at $R=0$, regressing $g_t$ on $z_t$ gives a slope of $-0.98$ thousand hours/year ($t=-57.7$), and {the residualized first-stage Wald statistic remains in the range $|t|\in[40,55]$ across all working dimensions $R\in\{1,\ldots,28\}$ reported below}.\footnote{ {Estimating with the alternative cutoff $z_t=\mathbf{1}\{\text{age}_t\ge 65\}$ delivers qualitatively identical conclusions across all $R$: the same sign pattern, IV point estimates within roughly $5\%$ of the age-62 values, and the same significance ordering. We report the age-62 specification as primary because the absence of the Medicare channel makes the just-identified exclusion restriction more defensible.}}}

Table (ref) reports both the OLS estimator $\widehat\beta_{\mathrm{OLS}}$ ($z_t=g_t$) and the IV estimator $\widehat\beta_{\mathrm{IV}}$ ($z_t=\mathbf{1}\{\text{age}_t\ge 62\}$) of Theorem (ref) across {$R\in\{0,1,2,3,5,7,10,15,20,28\}$}, both with HC$_0$ standard errors and both reported per a $1{,}000$-hours-per-year increment.

{One scope qualification should be flagged before reading the numbers. The i.i.d.-across-respondents assumption used in our theory ignores household clustering present in HRS, where spouses appear as separate observations sharing many of the controls. We report the unweighted HC$_0$ inference for transparency with our theory but note that a household-clustered standard error would slightly inflate the reported standard errors. With this caveat, the IV estimator under the exclusion restriction in Assumption (ref)(i) does have a causal local-average-treatment-effect (LATE) interpretation in the sense of ImbensAngrist1994: $\widehat\beta_{\mathrm{IV}}$ identifies the effect of labor supply on depressive symptoms among the subpopulation of compliers, namely respondents whose retirement timing responds to crossing the {age-62 institutional threshold}. Off-panel confounders that are absorbed neither by $\widehat F$ nor by $z_t$ (for example, idiosyncratic preferences and unanticipated household shocks) cannot bias $\widehat\beta_{\mathrm{IV}}$ to first order so long as they are uncorrelated with the {age-62 cutoff} after factor adjustment.}

{Three patterns are visible. (i) Omitted-factor bias at $R=0$ is severe: OLS returns $\widehat\beta = -0.276$ ($t=-19.7$), the strong negative cross-sectional association between hours worked and depressive symptoms documented in the HRS-based literature DaveRashadSpasojevic2008,MandalRoe2008; once even one principal component is partialled out, the OLS estimate collapses to essentially zero (magnitude below $0.04$ across all $R\ge 1$). The unconditioned cross-sectional correlation is almost entirely accounted for by a single factor direction in $x_t$, and OLS without an instrument provides no informative signal once that factor is removed. (ii) The IV LATE, by contrast, is uniformly positive and stable across all $R\ge 1$: $\widehat\beta_{\mathrm{IV}}$ ranges between $+0.64$ and $+1.01$, taking the value $+0.660$ ($t=12.95$) at the saturating $R=N=28$. The sign flip relative to OLS at $R=0$ is the textbook signature of selection bias: respondents who endogenously work more hours are mentally healthier on average for unobserved reasons that the controls do not span. Once the age-62 instrument purges this selection, the IV LATE points the other way: among respondents whose labor supply is shifted by the institutional cutoff, an additional $1{,}000$ hours per year of work raises CES-D by roughly $0.66$ points (about a third of the outcome's standard deviation), consistent in sign with the IV-based retirement literature Bonsang2012,Coe2011,MazzonnaPeracchi2012,Insler2014 that finds retirement to be protective for mental health among compliers. (iii) HC$_0$ standard errors and the first-stage Wald statistic are essentially flat across $R\ge 1$ for both columns.}

table[table omitted — 1,418 chars of source]

Conclusion

This paper develops spectral theory for principal component analysis in factor models when the working number of factors $R$ is fixed and weakly dominates the true, unknown factor dimension $r$. Leveraging anisotropic local laws from random matrix theory, we show that the overestimated empirical eigencomponents are noise-governed, incoherent, and near-orthogonal to the factor space; that the low-rank signal and factor space are recovered at the usual parametric rates under suitably generalized (expanded and compressed) rotations; and that factor-augmented inference on a treatment coefficient remains asymptotically valid for any bounded $R \ge r$ when $r\ge 1$. These results formally justify a common empirical practice (deliberately overestimating or adopting a conservative upper bound on the number of factors) and shift the analytical burden from consistent factor-number selection to the structurally milder requirement of bounding $r$ from above.

Beyond the technical contributions, our results have a methodological implication for empirical practice. The dominant approach in factor-model inference treats consistent dimension selection as a logically prior step: estimate $r$, condition on it, and proceed. This paper shows that this step can be replaced by the weaker requirement of specifying an upper bound $R\ge r$. The benefit is robustness: an inferential procedure that depends only on $R\ge r$ is insulated from the finite-sample volatility of $\widehat r$, which is well documented to be substantial in the signal regimes where applied researchers most need reliable inference. The cost is a variance inflation that scales with $R/T$. For empirically relevant settings such as factor-augmented treatment-effect inference, factor-based forecasting, and large-panel principal-component regression, this trade-off can be favorable.

Several directions merit further investigation, listed in approximate order of substantive importance.} {First, the factor-augmented regression we consider assumes serially-uncorrelated residuals. A parallel extension to weakly serially-dependent residuals appears tractable so long as cross-sectional independence in $u_t$ is maintained. The genuinely difficult extension is to allow serial and cross-sectional dependence simultaneously in $u_t$, since that is precisely the regime in which the anisotropic local law of knowles2017anisotropic ceases to apply, and a fundamentally different random-matrix input would be required. Second, our analysis imposes $N/T \to \phi \ne 1$ to inherit incoherence from knowles2017anisotropic; removing this gap may require new random-matrix tools, perhaps through a fluctuation-scale argument at the Marchenko--Pastur edge. Finally, our framework restricts attention to bounded $R-r$; the regime in which $R$ grows slowly with $T$ would require sharper control of the cumulative variance contribution of the overestimated components and may be relevant for sieve-type applications.