Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,519 characters · 21 sections · 88 citation commands
Factor-Augmented Panel Regressions and Variance-Weighted Treatment Effects
\thispagestyle{empty} \setcounter{page}{0}
Linear panel regressions with interactive fixed effects have become a standard tool in empirical economics since the seminal contributions of pesaran2006estimation and bai2009panel. These estimators are designed for panels of $n$ units observed over $T$ time periods, and allow for unobserved unit- and time-specific heterogeneity that enters the outcome model multiplicatively, generalizing the additive two-way fixed effects structure. A large literature has since developed around these methods, studying their asymptotic properties, the selection of the number of factors, and extensions to nonlinear models. However, the theoretical properties of these estimators are almost exclusively derived under parametric assumptions: the data generating process is assumed to follow a linear factor model of the form $Y_{it} = X_{it}\beta + \lambda_i'f_t + \varepsilon_{it}$, and consistency is established for the true parameter $\beta$ of that model. A natural and largely unanswered question is what these estimators converge to when the parametric model is misspecified or when no parametric structure is imposed at all.
This paper answers that question for the principal components (PC) estimator of GreenawayMcGrevy201248 and the interactive fixed effects (IFE) estimator of bai2009panel. We show that both estimators have well-defined and interpretable probability limits under fully nonparametric assumptions on the data generating process. The key insight comes from combining the factor model literature with results on variance-weighted treatment effects from the cross-sectional literature. Angrist1998 showed that OLS with observed controls estimates a variance-weighted average of heterogeneous treatment effects, where the weights are proportional to the conditional variance of the regressor given the controls. We show that the same interpretation extends to the two-way panel setting with unobserved heterogeneity: both the PC and IFE estimators converge to a variance-weighted average of unit-time-specific treatment effects, where each weight is proportional to the conditional variance of the regressor given the unobserved unit- and time-specific heterogeneity. This result does not require any parametric restrictions on the conditional expectation model, but the number of estimated factors $R_{nT}$ is required to grow with the sample size. The reason this works is that a growing number of principal components effectively turns the truncated singular value decomposition (SVD) of the regressor matrix into a nonparametric estimator of its conditional mean given the unobserved heterogeneity: classical results in approximation theory show that any sufficiently smooth bivariate function admits a Hilbert-Schmidt expansion with rapidly decaying singular values, so that the rank-$R_{nT}$ truncation approximates the conditional mean arbitrarily well as $R_{nT} \to \infty$ birman1977estimates. By contrast, other popular approaches such as the common correlated effects (CCE) estimator of pesaran2006estimation or the standard two-way fixed effect (TWFE) estimator generally do not share this property.
Our result can be seen as the large $n,T$ analogue of findings in the DiD and TWFE literature that have attracted considerable attention since 2018. 2018arXiv180308807D, Goodman2021, SUN2020, and CALLAWAY2020, among others, showed that TWFE estimators with staggered treatment adoption can assign negative weights to some treatment effects, and proposed alternative estimators that deliver non-negative weights. In that literature, parallel trends assumptions are required throughout because $T$ is treated as fixed, and the identification problem is intrinsically harder with limited time periods. Our setting is different: with $T$ growing, the factor structure can be estimated nonparametrically, and the variance-weighted average can be recovered without any parallel trends assumption. Nevertheless, the conceptual message is similar: pooled regression estimators target specific weighted averages of heterogeneous effects, and non-negativity of those weights is essential for interpreting the estimates.
There is also a growing literature that, while allowing for additive unobserved intercepts, relaxes the assumption of common slope coefficients, see e.g.\ LU2023694 and KeaneNeal2020. Some papers allow for rich heterogeneity in both slopes and intercepts, see e.g.\ chernozhukov2019inference and unknown2024estimation. All of these approaches impose parametric structure on how the unit-time-specific treatment effects vary with the unobserved heterogeneity.
In the extension section, we informally discuss estimating user-specified weighted averages of these effects without such restrictions. We also discuss the limitation that our variance-weighted average interpretation is specific to the case of a single regressor. When multiple regressors are included, contamination bias arises as documented by 10.1257/aer.20221116 for cross-sectional regressions and by de2023two for TWFE regressions with several treatments. Finally, we explain that due to the nonparametric nature of our results, inference on the variance-weighted average is generally not possible without additional assumptions, as the asymptotic distribution of both estimators is dominated by approximation bias.
The remainder of the paper is structured as follows. Section (ref) defines the variance-weighted estimands for cross-sectional and panel data settings. Section (ref) introduces the PC and IFE estimators, discusses the difficulties of controlling for unobserved two-way heterogeneity, and states the main consistency results. Section (ref) contains further discussion and extensions, covering inference, estimation of targeted treatment effects with user-specified weights, and the case of multiple regressors. Section (ref) explores the finite sample properties of the methods using simulations. Section (ref) concludes. Appendix (ref) contains all proofs.
Notation. For any matrix $A$ we write $\|A\|$ for the spectral norm and $\|A\|_2$ for the Frobenius norm.
In this section we introduce the variance-weighted treatment effect framework. We begin with the cross-sectional setting, where the key ideas are most transparent. We then extend the framework to panel data with unobserved heterogeneity. The results reviewed here are well established in the literature. Our purpose in this section is purely to set up the notation and conceptual apparatus needed for the two-way panel setting in Section (ref).
Consider a random sample $(Y_i, X_i, C_i)$, $i = 1, \ldots, n$. Here $Y_i \in \mathbb{R}$ is an outcome variable, $X_i \in \mathbb{R}$ is a scalar regressor of interest, and $C_i \in \mathbb{R}^p$ is a vector of control variables. We are interested in the effect of $X_i$ on $Y_i$ while controlling for $C_i$. We begin with the simplest case where no controls are present. Assume only that $(Y_i, X_i)$ are identically distributed with finite second moments. Let $\widehat{\beta}$ denote the coefficient on $X_i$ from the OLS regression of $Y_i$ on $(1, X_i)'$. Then
as $n \to \infty$.\footnote{When the regressor is binary, $X_i \in \{0, 1\}$, this simplifies to $\beta^* = \mathbb{E}\left(Y_i \,\big|\, X_i = 1\right) - \mathbb{E}\left(Y_i \,\big|\, X_i = 0\right)$, which has a causal interpretation as an average treatment effect if $X_i$ is randomly assigned.} Now suppose we wish to control for the observed covariates $C_i$. For each value of the controls, we can define the conditional linear projection coefficient
This measures the linear relationship between $Y_i$ and $X_i$ holding $C_i$ fixed.\footnote{Again, for binary $X_i$, $B(c)$ can be expressed as the conditional average treatment effect $B(c) = \mathbb{E}\left(Y_i \,\big|\, X_i = 1,\, C_i = c\right) - \mathbb{E}\left(Y_i \,\big|\, X_i = 0,\, C_i = c\right)$.} In general, $B(c)$ varies with $c$, that is, the effect of $X_i$ on $Y_i$ is heterogeneous. Rather than estimating $B(c)$ for each value of $c$ separately, it is natural to consider a pooled estimand that summarizes the overall relationship. To this end, we define the orthogonal decomposition
Here $X_i^\parallel$ is the conditional mean of $X_i$ given $C_i$, while $X_i^\perp$ is the residual variation in $X_i$ that remains after conditioning on the controls. Note that $\mathbb{E}\left(X_i^\perp \,\big|\, C_i\right) = 0$ by construction. The pooled estimand is now defined as
This generalizes the definition in equation (ref) to account for controls. It is the population regression coefficient from regressing $Y_i$ on the residualized regressor $X_i^\perp$. A key observation is that $\beta^*$ can be expressed as a weighted average of the conditional effects $B(c)$. We have\footnote{By the law of iterated expectations we have $\mathbb{E}\left(Y_i X_i^\perp\right) = \mathbb{E}\left[\mathbb{E}\left(Y_i X_i^\perp \,\big|\, C_i\right)\right] = \mathbb{E}\left[{\rm Cov}\left(Y_i,\, X_i \,\big|\, C_i\right)\right]$ and $\mathbb{E}\left[\left(X_i^\perp\right)^2\right] = \mathbb{E}\left[\mathbb{E}\left(\left(X_i^\perp\right)^2 \,\big|\, C_i\right)\right] = \mathbb{E}\left[{\rm Var}\left(X_i \,\big|\, C_i\right)\right]$, where we used that $\mathbb{E}\left(X_i^\perp \,\big|\, C_i\right) = 0$.}
where the weights are given by
These weights are non-negative by construction and satisfy $\mathbb{E}\left[W^*(C_i)\right] = 1$. Thus $\beta^*$ is a convex combination of the conditional effects $B(c)$. Observations receive larger weight when the conditional variance of $X_i$ given $C_i$ is larger. This is the variance-weighted average treatment effect interpretation of regression estimands emphasized by Angrist1998 and subsequent literature. See also angrist2008mostly, abadie2018econometric, sloczynski2020interpreting, and 10.1257/aer.20221116.
An important feature of $\beta^*$ is that it can be consistently estimated by OLS without ever explicitly having to estimate $B(c)$. Suppose that either of the following two conditions holds\footnote{Note that condition (ref) (a) and (b) are both automatically satisfied when $C_i$ consists of a saturated set of dummy variables. This observation is particularly relevant for the panel data setting in the next subsection, where the unobserved controls $U_i$ can be thought of as unit-specific dummies.}
Then the OLS estimator $\widehat{\beta}$ on $X_i$ from regressing $Y_i$ on $(1, X_i, C_i')'$ consistently estimates $\beta^*$ (here we assume existence of second moments throughout). To see this, let $\widehat{X_i^\perp}$ denote the OLS residual from regressing $X_i$ on $(1, C_i')'$. Under condition (ref)(a), $\widehat{X_i^\perp}$ is a consistent estimator of $X_i^\perp$. The Frisch-Waugh-Lovell theorem then implies that
as $n \to \infty$.
Equally, under condition (ref)(b) and with $\beta^*$ defined in (ref), we have $Y_i - \beta^* X_i - \rho_1 - \rho_2' C_i$ uncorrelated with $(1, X_i, C_i')'$, so $\beta^*$ is the population OLS coefficient on $X_i$.\footnote{A more standard way to express condition (ref)(b) is as follows. Assume that $Y_i = \rho_1 + \beta X_i + \rho_2' C_i + \varepsilon_i$ for some $\rho_1$, $\beta$, $\rho_2$ such that $\mathbb{E}(\varepsilon_i \,|\, C_i) = 0$ and $\mathbb{E}(\varepsilon_i X_i) = 0$. Then $\beta = \beta^*$ (assuming $\beta^*$ is well-defined), and the OLS coefficient $\widehat{\beta}$ on $X_i$ from regressing $Y_i$ on $(1, X_i, C_i')'$ consistently estimates $\beta^*$.} Admittedly, condition (ref)(b) may seem redundant here, since both (ref)(a) and (ref)(b) justify the same OLS estimator. However, the dual justification will be important in Section (ref), where the generalization of each condition corresponds to a different estimation approach.
More generally, one may wish to estimate weighted averages of $B(c)$ with user-specified weights,
for some weighting function $w(\cdot)$. When $C_i$ takes discrete values, this can be done by first estimating $B(c)$ separately for each value $C_{i}=c$ and then forming the weighted average given the weights $w(c)$. However, establishing consistency for such estimators is generally more difficult than for $\widehat{\beta}$ above. This is particularly the case when the overlap condition ${\rm Var}\left(X_i \,\big|\, C_i = c\right) > 0$ fails for some values of $c$. We return to this issue in Section (ref).
We now extend the framework to panel data where the controls are unobserved. Consider a sample $(Y_{it}, X_{it})$ for $i = 1, \ldots, n$ and $t = 1, \ldots, T$. Here $Y_{it} \in \mathbb{R}$ is an outcome and $X_{it} \in \mathbb{R}$ is a scalar regressor. We assume that for each unit $i$ there exists an unobserved confounder $U_i$ (possibly multi-dimensional) that may affect both $Y_{it}$ and $X_{it}$. Formally, we write\footnote{ We do allow for a DGP of the form $Y_{it} = h(X_{it}, U_i, \varepsilon_{it})$, which when combined with $X_{it} = g_X(U_i, \varepsilon_{it})$ gives (ref). However, the assumption does rule out dynamic regressors such as $X_{it} = Y_{i,t-1}$. The variance-weighted representation of the within estimator carries over to the dynamic case, but the interpretation is subtle, and we do not pursue this here.}
for some unknown measurable function $g(\cdot,\cdot)$ and idiosyncratic shocks $\varepsilon_{it}$ (possibly multi-dimensional). The structure parallels the cross-sectional case, with $U_i$ playing the role of the control variable $C_i$. For each unit $i$, we define the conditional linear projection coefficient\footnote{When $X_{it}$ is binary, $\beta_i$ reduces to the conditional average treatment effect $\beta_i = \mathbb{E}\left(Y_{it} \,\big|\, X_{it} = 1,\, U_i\right) - \mathbb{E}\left(Y_{it} \,\big|\, X_{it} = 0,\, U_i\right)$.}
We assume covariance stationarity here, which implies that $\beta_i$ does not depend on $t$. As before, we define the orthogonal decomposition
The pooled estimand is
where the weights are given by
Thus $\beta^*$ is a variance-weighted average of the unit-specific effects $\beta_i$, where as before the weights are non-negative and sum to one in expectation.
At first sight, it may seem difficult to estimate $\beta^*$ since the confounder $U_i$ is unobserved. However, the panel structure allows us to estimate $X_i^\parallel$ nonparametrically. Under covariance stationarity and mild mixing conditions on the serial dependence of $X_{it}$, we have
as $T \rightarrow \infty$.\footnote{ The $T \to \infty$ framing in this subsection is used for expositional simplicity and for continuity with the two-way case in Section (ref), where it is genuinely needed. It is not essential for the variance-weighted interpretation of the within estimator in the one-way setting: for any fixed $T \geq 2$, the within demean gives $\mathbb{E}(Y_{it}\,\widehat{X_{it}^\perp})/\mathbb{E}((\widehat{X_{it}^\perp})^2) = \beta^*$ exactly, since the factor $(T-1)/T$ arising from demeaning cancels in the numerator and denominator.} This suggests a natural estimator for $X_{it}^\perp$, namely
The resulting estimator for $\beta^*$ is
This is the classical within group (WG) estimator, also known as the fixed effects (FE) estimator. Under appropriate regularity conditions, $\widehat{\beta} \to_P \beta^*$ as $n, T \to \infty$, see galvao2014estimation and Juodis2020 for details. Unlike the cross-sectional setting, no linearity condition analogous to (ref) is needed here. Covariance stationarity ensures that $X_i^\parallel$ is constant over $t$, so time-averaging removes it nonparametrically regardless of the functional form of $\mathbb{E}(X_{it} \,\big|\, U_i)$.
More generally, one may wish to estimate weighted averages of $\beta_i$ with user-specified weights. For each unit $i$, define the unit-specific estimator\footnote{Here and in (ref), we assume that $T^{-1} \sum_{t=1}^T \left(\widehat{X_{it}^\perp}\right)^2 \geq c > 0$ for all $i$. This rules out irregular identification issues that can arise for fixed $T$. See https://doi.org/10.3982/ECTA8220 for discussion.}
The Mean Group estimator of Pesaran199579 uses equal weights,
For general weights $w_i$, one can form $\widehat{\beta}_w = \sum_{i=1}^n w_i \widehat{\beta}_i$. However, establishing consistency for such estimators (including other functionals of $\widehat{\beta}_i$) in the general model-free setup is more involved than for the pooled estimator $\widehat{\beta}$. See e.g. galvao2014estimation, OkuiYanagi2015, and jochmans2024inference.
We now consider the setting where unobserved heterogeneity is present in both the cross-sectional and time-series dimensions. Consider a sample $(Y_{it}, X_{it})$ for $i = 1, \ldots, n$ and $t = 1, \ldots, T$. We assume that the observed variables $Z_{it} = (Y_{it}, X_{it})'$ are generated as\footnote{ As in the one-way case, we allow for a DGP of the form $Y_{it} = h(X_{it}, U_i, V_t, \varepsilon_{it})$, which when combined with $X_{it} = g_X(U_i, V_t, \varepsilon_{it})$ gives (ref). However, the assumption does rule out dynamic regressors such as $X_{it} = Y_{i,t-1}$.}
where $g(\cdot,\cdot,\cdot)$ is a general measurable function, $U_i$ and $V_t$ are unobserved unit-specific and time-specific heterogeneity, and $\varepsilon_{it}$ are idiosyncratic shocks (all possibly multi-dimensional). This decomposition is closely related to the Aldous-Hoover representation for exchangeable arrays. See 2019Menzel for background and chiang2022standard for an application to panel data.
The structure parallels Section (ref), but now we condition on both $U_i$ and $V_t$. For each $(i,t)$, we define the conditional linear projection coefficient
Due to the presence of time-specific heterogeneity $V_t$, this quantity now varies across both $i$ and $t$.\footnote{When $X_{it}$ is binary, $\beta_{it}$ reduces to the conditional average treatment effect $\beta_{it} = \mathbb{E}\left(Y_{it} \,\big|\, X_{it} = 1,\, U_i,\, V_t\right) - \mathbb{E}\left(Y_{it} \,\big|\, X_{it} = 0,\, U_i,\, V_t\right)$.} As before, we define the orthogonal decomposition
and analogously for $Y_{it}$. Under identical distribution across $(i,t)$, the pooled estimand is
where the weights are given by
Thus $\beta^*$ is again a variance-weighted average of the conditional effects $\beta_{it}$, with non-negative weights that sum to one in expectation. Since $\mathbb{E}(Y_{it}^\parallel X_{it}^\perp) = 0$ by iterated expectations, we also have
This equivalent representation will be useful for understanding the IFE estimator in Section (ref). As in Section (ref), the natural plug-in estimator for $\beta^*$ takes the form
where $\widehat{X_{it}^\perp}$ is an estimator of $X_{it}^\perp$. The key challenge is constructing $\widehat{X_{it}^\perp}$, which requires estimating $X_{it}^\parallel = \mathbb{E}\left(X_{it} \,\big|\, U_i,\, V_t\right)$.
This task is fundamentally different from the one-way setting of Section (ref). In that setting, covariance stationarity implied $X_{it}^\parallel = X_i^\parallel$, which could be estimated nonparametrically by time-averaging. Here, $X_{it}^\parallel$ varies in both dimensions. Without further restrictions, it is impossible to distinguish $X_{it}^\parallel$ from $X_{it}^\perp$ using $X_{it}$ alone.
One approach to estimating $X_{it}^\perp$ is to impose an additive structure on $X_{it}^\parallel$,\footnote{ Analogous to condition (ref)(b) in the cross-sectional setting, we could alternatively assume an additive structure on $Y_{it}^\parallel - \beta X_{it}^\parallel$. This would justify the same TWFE estimator under different assumptions.}
for some measurable functions $g_1(\cdot)$ and $g_2(\cdot)$. Under this assumption, we can estimate $X_{it}^\perp$ by the two-way demeaning
The resulting estimator $\widehat{\beta}$ is the two-way fixed effects (TWFE) estimator. However, the additive structure is restrictive. When $X_{it}^\parallel$ contains terms that interact $U_i$ and $V_t$ non-additively, the two-way demeaning leaves behind a residual that is generally correlated with $X_{it}^\perp$, and TWFE fails to consistently estimate $\beta^*$.
The common correlated effects (CCE) approach of pesaran2006estimation\footnote{See e.g.\ JUODIS2026106120 for a critical review of the recent literature on this methodology.} as well as generalizations of this approach (see e.g.\ JuodisEtal2017 and LU2026106183) generally do not fully resolve the shortcomings of the TWFE transformation. In particular, the CCE approach builds upon the assumption that
where now $g_1(\cdot)$ and $g_2(\cdot)$ are $R$-dimensional vectors. $R$ is generally assumed to be no larger than the number of regressors, e.g.\ $R = 1$ is the maximum permitted value for the single regressor case that we consider here.\footnote{This limitation can be circumvented by using additional proxy variables as in KarUrbWest2014 and JuodisSarafidis2019, but generally the dimension $R$ remains finite as $n, T \to \infty$.} Hence, while the CCE approach does not impose additive structure on $X_{it}^\parallel$, it keeps the restriction that $X_{it}^\parallel$ needs to admit exact low-rank structure, unlike the approaches that we describe below. Moreover, the CCE method imposes additional restrictions on the first conditional moments of $X_{it}^\parallel$ in the form of the so-called rank condition. In Appendix (ref) we provide a simple DGP that illustrates the potential failure of the TWFE and the CCE methods.
Next, we show that consistent estimation of $\beta^*$ is possible under weaker conditions. The key insight is that $X_{it}^\parallel = \mathbb{E}(X_{it} \,|\, U_i, V_t)$ can often be well-approximated by a low-rank structure, namely the leading principal components of $X_{it}$. This approximation is accurate when $X_{it}^\parallel$ is a sufficiently smooth function of $(U_i, V_t)$ and the dimensionality of $U_i$ and $V_t$ is not too high.
We now describe the two principal components based estimators that are the focus of this paper. Under suitable regularity conditions, both estimators consistently estimate the variance-weighted average effect $\beta^*$ defined in (ref). Throughout, we write $Y$, $X$, $X^\parallel$, and $X^\perp$ for the $n \times T$ matrices with $(i,t)$ entries $Y_{it}$, $X_{it}$, $X_{it}^\parallel$, and $X_{it}^\perp$ respectively.
The first estimator is the Principal Components (PC) estimator studied in GreenawayMcGrevy201248. This is a two-step procedure. In the first step, we estimate $X_{it}^\parallel$ by the rank-$R$ approximation to $X$,
This minimization has a closed-form solution given by the truncated singular value decomposition of $X$.\footnote{Let $X = \sum_{r=1}^{\min(n,T)} \sigma_r u_r v_r'$ be the singular value decomposition of $X$, with singular values $\sigma_1 \geq \sigma_2 \geq \ldots \geq 0$. By the Eckart-Young-Mirsky theorem, the solution to (ref) is $\widehat{X^\parallel} = \sum_{r=1}^{R} \sigma_r u_r v_r'$.} Let $\widehat{X_{it}^\perp} := X_{it} - \widehat{X_{it}^\parallel}$ denote the residuals. In the second step, the PC estimator is obtained by
Our second estimator is the Interactive Fixed Effects (IFE) estimator of bai2009panel. This estimator estimates the coefficient $\beta$ and the factor structure jointly by solving
Here, the first line uses notation analogous to (ref), while the second line replaces the rank restriction on $G$ by the explicit low-rank representation $G_{it} = \lambda_i' f_t$, where $\lambda_i = (\lambda_{i1}, \ldots, \lambda_{iR})' \in \mathbb{R}^R$ are unit-specific factor loadings and $f_t = (f_{t1}, \ldots, f_{tR})' \in \mathbb{R}^R$ are time-specific factors. The second line also shows that $\widehat{\beta}_{\rm IFE}$ is simply obtained as the joint least squares estimator over $(\beta, \lambda, f)$.
The relationship between PC and IFE mirrors the two justifications for OLS in the cross-sectional setting. Recall from Section (ref) that the OLS estimator consistently estimates $\beta^*$ under either condition (ref)(a), which restricts the conditional mean of $X_i$, or condition (ref)(b), which restricts the conditional mean of $Y_i - \beta^* X_i$. By the Frisch-Waugh-Lovell theorem, the same OLS estimator is consistent for $\beta^*$ under either condition. In the two-way panel setting, the analogous conditions are:
The PC estimator $\widehat{\beta}_{\rm PC}$ corresponds to condition (a): it first removes the low-rank structure from $X$, then regresses $Y$ on the residuals. The IFE estimator $\widehat{\beta}_{\rm IFE}$ corresponds to condition (b): it jointly estimates $\beta$ and a low-rank approximation to $Y - X\beta$. Unlike the cross-sectional case, the two approaches are not numerically equivalent. However, under appropriate regularity conditions, both estimators consistently estimate $\beta^*$, as we show below.
Both estimators depend on a tuning parameter $R$, which controls the number of principal components or factors used. Our consistency results require $R$ to grow with $n$ and $T$.
We now provide results for consistency of the PC and IFE estimators under high-level conditions. In the following subsection, we provide lower-level sufficient conditions that are easier to interpret and verify. We begin with the PC estimator. Recall that this estimator is defined by (ref) and (ref) as a function of $Y$, $X$, and $R$. We now write $R = R_{nT}$, and allow it to depend on the sample size.
The theorem is stated in a self-contained way: it assumes the existence of a constant $\beta^*$, a matrix $X^\parallel$, and derived quantities $E = Y - \beta^* X$ and $X^\perp = X - X^\parallel$ satisfying conditions (i)--(iv), and concludes that $\widehat{\beta}_{\rm PC}$ consistently estimates $\beta^*$. No specific interpretation of these quantities is assumed. Of course, the theorem becomes useful when we set $X_{it}^\parallel := \mathbb{E}(X_{it} \,|\, U_i, V_t)$ and define $\beta^*$ as the variance-weighted average in (ref). This is what we do in the sufficient conditions below.
We now briefly discuss each condition. Condition (i) requires that $X^\perp$ and $E$ are asymptotically uncorrelated, and that $X$ and $Y$ have bounded second moments. Condition (ii) is a non-degeneracy condition ensuring that $X_{it}^\perp$ has positive variance in the limit. Condition (iii) requires that the spectral norm of $X^\perp$ does not grow too fast relative to $nT/R_{nT}$. Condition (iv) is the key low-rank assumption: it requires that $X^\parallel$ can be well-approximated by a rank-$R_{nT}$ matrix. This corresponds to condition (ref)(a) in the previous subsection. Note that (iii) requires $R_{nT}$ to not be too large, while (iv) requires $R_{nT}$ to not be too small. The sufficient conditions below assume that $R_{nT}$ grows with $n$ and $T$, but slowly.
We now state the analogous result for the IFE estimator.
As with Theorem (ref), this result is stated in a self-contained way: it assumes the existence of a constant $\beta^*$, a matrix $E^\parallel$, and derived quantities $E = Y - \beta^* X$ and $E^\perp = E - E^\parallel$ satisfying conditions (i)--(iv). The theorem becomes useful when we set $E_{it}^\parallel := \mathbb{E}(E_{it} \,|\, U_i, V_t)$ and define $\beta^*$ as the variance-weighted average in (ref).
Compared to Theorem (ref), the roles of $X$ and $E$ are reversed, reflecting the different structure of the IFE estimator. Condition (i) requires that $X$ is asymptotically uncorrelated with the idiosyncratic component $E^\perp$, and that $X$ has bounded second moments. Condition (ii) is a non-collinearity condition requiring that $X$ cannot be well-approximated by a low-rank matrix.\footnote{Here, similarly to FreemanWeidner2021, we need to assume that ${\rm rank}(G) \leq 2R_{nT}$ and not simply ${\rm rank}(G) \leq R_{nT}$.} This ensures that the coefficient $\beta$ can be identified separately from the factor structure. Condition (iii) requires that the spectral norm of $E^\perp$ does not grow too fast relative to $nT/R_{nT}$. Condition (iv) is the key low-rank assumption: it requires that $E^\parallel$ can be well-approximated by a rank-$R_{nT}$ matrix. This corresponds to condition (ref)(b) in the previous subsection. As for the PC estimator, (iii) requires $R_{nT}$ to not be too large, while (iv) requires $R_{nT}$ to not be too small.
The proofs of both theorems are given in Appendix (ref). The proof strategy builds on the approach of FreemanWeidner2021, adapted to our setting.
We now provide primitive sufficient conditions on the data generating process (ref) that imply the high-level conditions in Theorems (ref) and (ref). We emphasize that these are only sufficient conditions: the high-level conditions may also be satisfied under alternative assumptions not covered by the theorem below.
The proof of Theorem (ref) is given in Appendix (ref). We now provide an informal discussion of why these conditions imply the high-level assumptions of Theorems (ref) and (ref). The formal proof proceeds through an intermediate set of conditions stated in Lemma (ref) in the appendix.
Conditions (i) and (ii) specify the dependence structure of the data generating process. Units are drawn independently, while time periods may exhibit serial dependence through both the common shocks $V_t$ and the idiosyncratic shocks $\varepsilon_{it}$. The $\alpha$-mixing assumption with summable coefficients ensures that this serial dependence decays sufficiently fast. These conditions are standard in the panel data literature with common factors; see fernandez_weidner_2014 and SuEtAL2014 for related discussions. Together with condition (iii), they imply laws of large numbers and bounds on spectral norms that are needed to verify the high-level conditions. In particular, the independence of $\varepsilon_{it}$ from $(U, V)$ ensures that $X^\perp$ and $E^\perp$ are asymptotically uncorrelated, which is essential for conditions (i) in both Theorems (ref) and (ref). The mixing conditions also yield bounds on the spectral norms $\|X^\perp\|$ and $\|E^\perp\|$ that, combined with condition (v), imply conditions (iii) in both theorems. See wang2022lowrankpanelquantileregression for related spectral norm bounds.
Condition (iii) provides moment bounds and a non-degeneracy requirement. The uniform fourth moment bound ensures that sample averages converge to their population counterparts. The requirement ${\rm Var}(X_{it} \,|\, U_i, V_t) > 0$ ensures that $X_{it}^\perp$ has positive variance, which is needed for condition (ii) in Theorem (ref) and condition (ii) in Theorem (ref).
Condition (iv) is the key smoothness assumption. It requires that the conditional mean functions $h_X$ and $h_E$ are sufficiently smooth relative to the underlying dimensionality of the unobserved heterogeneity. Classical results in approximation theory show that for an $s$-smooth function on $\mathbb{R}^d$, the singular values of the associated integral operator decay as $\sigma_r = \operatorname{\mathcal{O}}(r^{-s/d})$. When $s > d/2$, this decay is fast enough to ensure that $\sum_{r > R} \sigma_r^2 \to 0$ as $R \to \infty$. Applied to our setting with $d = \max(d_U, d_V)$, this implies that $X^\parallel$ and $E^\parallel$ can be well-approximated by low-rank matrices, which is precisely what conditions (iv) in Theorems (ref) and (ref) require.
Condition (v) restricts the growth rate of $R_{nT}$. The requirement $R_{nT} \to \infty$ ensures that the low-rank approximation error vanishes asymptotically. The upper bound $R_{nT} = o(\min(\sqrt{T}, \sqrt{n}/(\log T)^2))$ ensures that $R_{nT}$ does not grow so fast as to violate the spectral norm conditions (iii) in the high-level theorems.
Theorem (ref) is the main theoretical contribution of this paper. It establishes that the PC and IFE estimators have well-defined probability limits under fully nonparametric assumptions on the data generating process. Existing results for these estimators, such as bai2009panel, typically assume a parametric model of the form $Y_{it} = X_{it} \beta + \lambda_i' f_t + \varepsilon_{it}$. In contrast, our conditions do not impose any parametric structure. The representation $Y_{it} = X_{it}\beta^* + G_{it} + \xi_{it}$ is not assumed but derived from the definition of $\beta^*$ as a variance-weighted average. The low-rank condition (iv) restricts only the conditional mean functions $h_X$ and $h_E$, not the full distribution of $(Y_{it}, X_{it})$ given $(U_i, V_t)$. The key requirement for consistency, beyond regularity conditions, is that the number of factors $R_{nT}$ grows with $n$ and $T$, but slowly. This allows the low-rank approximation to capture increasingly complex patterns in the conditional means while ensuring that estimation error remains controlled.
The results in the previous subsections define the estimand $\beta^*$ as a population quantity that remains fixed as $n, T \to \infty$. This requires distributional assumptions on $(U_i)$ and $(V_t)$, such as the i.i.d.\ and stationarity conditions in Theorem (ref). An alternative approach, more in line with the standard interactive fixed effects literature, is to condition on $(U, V) = ((U_i)_{i=1}^n, (V_t)_{t=1}^T)$ throughout. This allows for arbitrary sequences $(U_i)$ and $(V_t)$, at the cost of working with an estimand that depends on $n$ and $T$. For fixed sequences $(U_1, \ldots, U_n)$ and $(V_1, \ldots, V_T)$, we define the conditional (or finite-population) estimand
where $\beta_{it}$ is defined in (ref) and the weights are
Thus $\beta^*_{nT}$ is a variance-weighted average of the conditional effects $\beta_{it}$, with non-negative weights satisfying $\sum_{i=1}^n \sum_{t=1}^T w_{it}^{nT} = 1$. Unlike $\beta^*$ in (ref), the estimand $\beta^*_{nT}$ depends on the realized sequences $(U_i)$ and $(V_t)$ and may change with $n$ and $T$, but its interpretation is unchanged.
Theorems (ref) and (ref) apply unchanged when $\beta^*$ is replaced by $\beta^*_{nT}$ and all statements are interpreted conditionally on $(U, V)$. The proofs require no modification: the key orthogonality condition $\sum_{i=1}^n \sum_{t=1}^T \mathbb{E}[X_{it}^\perp E_{it} \,|\, U_i, V_t] = 0$ holds by construction of $\beta^*_{nT}$, and all other conditions concern matrices and spectral norms that are well-defined conditional on $(U, V)$.
The following theorem provides sufficient conditions for consistency in the conditional setting. It parallels Theorem (ref), but replaces the distributional assumptions on $(U_i)$ and $(V_t)$ with direct conditions on the low-rank approximability of the conditional mean matrices.
The proof of Theorem (ref) is given in Appendix (ref). Condition (iii) of Theorem (ref) assumes directly that the conditional mean matrices admit good low-rank approximations. In Theorem (ref), this condition was derived from smoothness of the conditional mean functions $h_X$ and $h_E$, combined with the i.i.d.\ and stationarity assumptions on $(U_i)$ and $(V_t)$. In the conditional setting, smoothness alone does not suffice: we also require that the realized sequences $(U_i)$ and $(V_t)$ are sufficiently spread out, in the sense that the singular functions of $h_X$ and $h_E$ do not concentrate on these sequences. In Appendix (ref) we discuss a sufficient condition for Condition (iii) of Theorem (ref).
This section discusses limitations of the results presented so far as well as natural extensions of the framework. None of the discussions below involve new formal results, our goal here is just to clarify the scope of the paper and to point to directions for future work.
The consistency results in Theorems (ref) and (ref) are derived under minimal restrictions on the underlying DGP and are silent on the distribution of the two estimators. More progress can be achieved if additional restrictions on the DGP and on the approximation rate of $R_{nT}$ are imposed. In Appendix (ref) we follow FreemanWeidner2021 and derive rate results for the IFE estimator. As in that paper, the rate of convergence is driven by the approximation bias associated with $R_{nT} \to \infty$. The fact that the asymptotic distribution is dominated by this bias term is further confirmed by our Monte Carlo study in Section (ref). Since no general nonparametric method for correcting this bias is available, the PC and IFE estimators cannot be directly used for inference on $\beta^*$.\footnote{Even in settings where $E_{it}^\parallel$ or $X_{it}^\parallel$ have an exact finite factor structure, the estimators we consider suffer from the incidental parameters problem (see e.g.\ westerlund2015cross and MoonWeidnerET) and require bias correction.}
Some progress is possible if further restrictions on the DGP are imposed. Following https://doi.org/10.3982/ECTA15238, FreemanWeidner2021, and beyhum2024inferencediscretizingunobservedheterogeneity, one approach is to discretize the unobserved heterogeneity. We take the suggestion of beyhum2024inferencediscretizingunobservedheterogeneity as a leading example. The procedure assumes that there exist injective functions $a_i$ and $b_t$ of $U_i$ and $V_t$ respectively, with finite-sample counterparts $\widehat{a}_i$ and $\widehat{b}_t$, and proceeds in three steps:
Following beyhum2024inferencediscretizingunobservedheterogeneity, we refer to this as the two-way grouped fixed effects (TWGFE) estimator. Their Monte Carlo results indicate that its finite-sample distribution is well approximated by a normal distribution that ignores estimation uncertainty from the grouping step, provided the dimension of $U_i$ and $V_t$ is small (mirroring assumption (iv) of Theorem (ref)) and K-means clustering is used.
The TWGFE approach is not yet theoretically justified in our setting, due to the estimation of high-dimensional nuisance parameters indexed by group membership $g_i$ and $c_t$.\footnote{We conjecture that consistency of TWGFE for $\beta^*$ can be established under the low-level conditions of Theorem (ref) combined with the injectivity conditions in beyhum2024inferencediscretizingunobservedheterogeneity.} The theoretical results in beyhum2024inferencediscretizingunobservedheterogeneity are instead established via sample splitting, as it is common in the double machine learning literature 10.1111/ectj.12097. Such sample splitting could be justified in our setting only if the projection errors $E_{it}^\perp$ and $X_{it}^\perp$ were i.i.d.\ over both dimensions. Since in our setting these errors have no structural meaning beyond population linear projection residuals, imposing such restrictions is not appropriate.
A further difficulty is that the convergence rate of the infeasible pooled OLS estimator (even for $E^\parallel$ or $X^\parallel$ known) can range from $\sqrt{\min(n,T)}$ to $\sqrt{nT}$, depending on the DGP, without any intermediate rate being ruled out. As a result, uniform inference on $\beta^*$ is generally not possible in our setting, as discussed in 2019Menzel and Juodis2020.
The pooled estimand $\beta^*$ in (ref) uses variance-driven weights $w_{it}^*$ that are determined by the data generating process rather than the researcher. In many applications one may prefer a user-specified weighting scheme. A natural target is $\beta_w := \mathbb{E}(w_{it}\,\beta_{it})$ for some chosen weights $w_{it}$, with the equal-weighted case $w_{it} = 1$ being the panel analogue of the mean group estimator of Pesaran199579. Consistently estimating $\beta_w$ requires a way to estimate the unit-time specific coefficients $\beta_{it}$ themselves. The standard starting point is to write
and impose structure on how $\beta_{it}$ varies with $(i,t)$. Three specifications proposed in the literature are:
All three approaches reduce estimation of $\beta_{it}$ to a finite-dimensional problem by imposing separability (either additive or multiplicative) in the $(U_i, V_t)$ arguments. Estimation in each case further requires that $E_{it}^\parallel$ has an exact low-rank structure. These restrictions are convenient but potentially strong.\footnote{For example, the simple specification $\beta_{it} = \beta(U_i, V_t) = (g_{11}(U_i) + g_{21}(V_t))/(g_{12}(U_i) + g_{22}(V_t))$ satisfies neither additive nor multiplicative separability, and is not covered by any of the three approaches above.}
An alternative route for estimating $\beta_w = \mathbb{E}(w_{it}\,\beta_{it})$ is as follows. Let $Z_{it} = (Y_{it}, X_{it})'$ and define $Z_{it}^\perp = Z_{it} - \mathbb{E}(Z_{it} \,|\, U_i, V_t)$ as before. Then $\beta_{it}$ is a ratio of elements of the $2 \times 2$ conditional variance matrix
namely $\beta_{it} = \Sigma_{it,YX}/\Sigma_{it,XX}$, where $\Sigma_{it,YX}$ and $\Sigma_{it,XX}$ denote the $(Y,X)$ and $(X,X)$ elements respectively. Now observe that
so each element of $Z_{it}^\perp(Z_{it}^\perp)'$ has conditional mean equal to the corresponding element of $\Sigma_{it}$, with a mean-zero noise term $\Xi_{it}$. This is the same approximate low-rank structure exploited in the main results, now applied to the $2\times 2$ outer product matrix rather than to $X$ or $E$ alone. Using the identity
one can apply a low-rank approximation separately to each of the $n\times T$ matrices with $(i,t)$ entries $Y_{it}X_{it}$, $X_{it}^2$, $Y_{it}$, and $X_{it}$, obtaining estimators of the corresponding conditional means, and then form $\widehat{\Sigma}_{it,YX}$ and $\widehat{\Sigma}_{it,XX}$ by plugging in. Under smoothness conditions analogous to those of Theorem (ref), this yields consistent estimators of the elements of $\Sigma_{it}$. This leads to the estimator
We stress that (ref) is intended as a conceptual proposal rather than a final estimator. In particular, the denominator $\widehat{\Sigma}_{it,XX}$ estimates a conditional variance and may be close to zero for some $(i,t)$, so some form of regularization (such as thresholding) is needed to keep the ratio well-behaved. We leave a formal treatment of these questions for future work.
The estimator in (ref) can be seen as a natural extension of standard mean group methods to the two-way setting. In the one-way case, $\widehat{\Sigma}_{it,YX}$ and $\widehat{\Sigma}_{it,XX}$ are only allowed to vary along a single dimension (either $i$ or $t$), so that the conditional expectations can be estimated by simple within-group averages. Here the conditioning set is bivariate $(U_i, V_t)$, and the low-rank approximation plays the role that within-group averaging plays in the one-way case, handling both dimensions jointly. As an alternative to principal components, the discretization approach of beyhum2024inferencediscretizingunobservedheterogeneity can also be used: after clustering units into groups $g_i$ and time periods into groups $c_t$ as described in Section (ref), one can estimate $\widehat{\beta}_{g_i c_t}$ for every cell $(g_{i},c_{t})$ of the resulting partition and then form the weighted average $\widehat{\beta}_w$ directly.
Finally, our discussion so far addressed consistent estimation of $\beta_w$ but was silent on inference. Inference for this class of estimands is generally more involved than for the pooled estimand $\beta^*$, and typically requires either additional non-degeneracy conditions on the distribution of $\beta_{it}$ (see e.g.\ KeaneNeal2020 and LU2023694) or sample splitting and regularization schemes (see e.g.\ chernozhukov2019inference).
So far we have taken $X_{it}$ and $\beta$ to be scalars. We now consider the extension to $X_{it} \in \mathbb{R}^K$ and $\beta \in \mathbb{R}^K$, with $Y_{it} \in \mathbb{R}$ unchanged. Define $X_{it}^\parallel := \mathbb{E}(X_{it} \,|\, U_i, V_t) \in \mathbb{R}^K$ and $X_{it}^\perp := X_{it} - X_{it}^\parallel \in \mathbb{R}^K$ as before. The conditional projection coefficient (ref) then also becomes a $K$-vector,
The pooled estimand (ref) generalises to
where ${\rm Var}(X_{it} \,|\, U_i, V_t) = \mathbb{E}\bigl(X_{it}^\perp (X_{it}^\perp)' \,\big|\, U_i, V_t\bigr)$, and the second equality follows from the law of iterated expectations, exactly as in the scalar case. The PC and IFE estimators extend to
where the residuals $\widehat{X_{it}^\perp} \in \mathbb{R}^K$ are obtained by applying the truncated SVD (ref) to each of the $K$ regressor matrices $X^{(k)} \in \mathbb{R}^{n \times T}$ separately. With those straightforward modifications, Theorems (ref) and (ref) continue to hold, that is, one can still show\footnote{ The high-level conditions of Theorems (ref) and (ref) extend to this vector setting with minimal notational replacements, e.g.\ the non-degeneracy condition (ii) becomes $\frac{1}{nT}\sum_{i,t} X_{it}^\perp (X_{it}^\perp)' \to_P C$ for a positive definite matrix $C$.}
as $n,T \rightarrow \infty$.
The multivariate estimand (ref) merits careful interpretation. Consider first $\beta_{it}$ in (ref). This is simply the vector of conditional linear projection coefficients of $Y_{it}$ on $X_{it}$ given $(U_i, V_t)$: the $k$-th component $\beta_{it,k}$ measures the partial effect of $X_{it,k}$ on $Y_{it}$ holding the remaining $K-1$ treatments fixed, conditional on the unobserved heterogeneity $(U_i, V_t)$. This is a meaningful object: it generalises the scalar conditional projection coefficient (ref) in the natural way, and for binary treatments reduces to the conditional average treatment effect of the $k$-th treatment holding the other treatments fixed. Importantly, the conditioning on $(U_i, V_t)$ is fully nonparametric here, which is more demanding than the partially linear framework of 10.1257/aer.20221116, who condition on observed controls linearly.
The interpretation of $\beta^*$ in (ref) is more subtle. Writing out the $k$-th component,
it is clear that $\beta^*_k$ is generally a function of $\beta_{it,j}$ for all $j$, not only $j = k$. Specifically, $\beta^*_k$ can be expressed as
where the weights $\lambda_{kj,it}$ are determined by the conditional variance matrix ${\rm Var}(X_{it} \,|\, U_i, V_t)$ and its expectation. The own-treatment weights satisfy $\mathbb{E}(\lambda_{kk,it}) = 1$, while the contamination weights satisfy $\mathbb{E}(\lambda_{kj,it}) = 0$ for $j \neq k$. Since the contamination weights average to zero, they must be negative for some $(i,t)$ unless they are identically zero. Hence, $\beta^*_k$ is not a convex combination of the $\beta_{it,k}$: it is contaminated by the effects of the other $K-1$ treatments, with weights that need not be non-negative. This is precisely the contamination bias phenomenon identified by 10.1257/aer.20221116 in cross-sectional regressions with multiple treatments and flexible controls, and by de2023two in two-way fixed effects regressions with several treatments. Our results show that the same phenomenon arises in the factor model panel setting, and that it is not an artifact of any particular estimation approach: it is a property of the estimand $\beta^*$ itself.
The contamination bias disappears in two special cases that parallel those identified in 10.1257/aer.20221116. First, if the $K$ treatments are conditionally uncorrelated given $(U_i, V_t)$, so that ${\rm Var}(X_{it} \,|\, U_i, V_t)$ is diagonal for all $(i,t)$, then $\mathbb{E}[{\rm Var}(X_{it} \,|\, U_i, V_t)]$ is also diagonal and the contamination weights $\lambda_{kj,it}$ are identically zero. Second, contamination bias is zero if the effects $\beta_{it,j}$ are homogeneous across $(i,t)$ for all $j \neq k$, since the contamination weights average to zero. Outside these special cases, the vector $\beta^*$ retains a well-defined probability limit that is consistently estimated by the PC and IFE estimators, but its components do not admit the clean variance-weighted average interpretation that makes the scalar estimand (ref) attractive.
In this section we consider a simple setting with a single outcome variable $Y_{it}$ and a covariate $X_{it}$. Motivated by the framework developed in the previous sections, we construct the DGP in the form of a panel location-scale nonlinear factor model:\footnote{Alternatively, it can be seen as an extension of the linear panel quantile model in https://doi.org/10.3982/ECTA15746.}
where $\epsilon_{it,x}$ and $u_{it}$ are mutually independent random variables for all $(i,t)$. For the location functions $l_y(\lambda_i, f_t)$ and $l_x(\lambda_{i,x}, f_{t,x})$, we consider two setups.
The NLFM setup is inspired by the constant elasticity of substitution specification for unobserved heterogeneity proposed in https://doi.org/10.3982/ECTA15238, see also beyhum2024inferencediscretizingunobservedheterogeneity, while the LFM follows the standard framework of pesaran2006estimation and bai2009panel.
We restrict our attention to the setting where $\gamma_{it}$ takes the form
Note that $\gamma_{it}$ has a causal interpretation as the structural effect of $X_{it}$ on $Y_{it}$. We make this choice such that the implied value of $\beta_{it}$ for the overall DGP is given by
This allows us to consider setups where $\beta_{it}$ has either a causal interpretation (when $\rho = 0$) or a non-causal one (when $\rho \neq 0$), while keeping the overall structure of the estimand constant. In particular, the parameter $\beta^*$ estimated by the IFE and PC methods is given by
which depends on $\kappa$ and $\rho$ only through their sum $\kappa + \rho$. Throughout this section we use the linear two-way specification for the scale functions
such that the resulting $\beta_{it}$ violates the settings of chernozhukov2019inference and LU2023694.
The setup involves four unit-specific random variables $U_i = (\lambda_i, \lambda_{i,x}, \nu_i, \nu_{i,x})'$ and four time-specific random variables $V_t = (f_t, f_{t,x}, g_t, g_{t,x})'$. To keep the results tractable we impose
and
Here $\lambda_i^+$ and $\lambda_{i,x}$ are mutually independent cross-sectionally i.i.d.\ $\Gamma(1,1)$ sequences, while $f_t^+$ and $f_{t,x}$ are mutually independent AR(1) sequences with innovations drawn from a Gamma distribution with shape parameter $(1-\alpha)^2/(1-\alpha^2)$ and scale parameter $(1-\alpha^2)/(1-\alpha)$.\footnote{With this parametrization $\mathbb{E}(f_t^+) = \mathbb{E}(f_{t,x}) = 1$ and $\mathrm{Var}(f_t^+) = \mathrm{Var}(f_{t,x}) = 1$. } We set $\beta_0 = 0$ and $\alpha = 0.5$, and vary $\pi \in \{0.0, 0.5\}$.\footnote{Preliminary Monte Carlo results indicate that the choice of $\alpha$ is mostly irrelevant as long as $|\alpha| < 1$. The value $\beta_0 = 0$ is without loss of generality as long as non-zero values of $\rho$ and/or $\kappa$ are included in the analysis.} The specification for random variables follows the framework of beyhum2024inferencediscretizingunobservedheterogeneity. Given the properties of the Gamma distribution it is straightforward to establish that
We consider the following four special cases.
In the first two designs $\beta_{it}$ has no causal (NC) interpretation, while in the last two it does, hence the label C.\footnote{The CCE estimator of pesaran2006estimation is consistent for $\beta^*$ and asymptotically normal only under (DGP.1).}
Two remarks are in order. First, $\beta^*$ is a function of $l_y(\lambda_i, f_t)$ and $l_x(\lambda_{i,x}, f_{t,x})$, but not of $s_y(\nu_i, g_t)$ and $s_x(\nu_{i,x}, g_{t,x})$. As a result, $E_{it}^\parallel = Y_{it}^\parallel - \beta^* X_{it}^\parallel$ is not affected by the dimensionality of the unobservables in the scaling functions. Second, the dimensions of unobservables in $E_{it}^\parallel$ are larger than those in $X_{it}^\parallel$ (two-dimensional for both the cross-sectional and time series dimensions in $E_{it}^\parallel$, versus one-dimensional for $X_{it}^\parallel$). In this regard the PC approach is advantageous relative to the IFE method.
We set $n = T$ and vary $n \in \{25, 50, 75, 100, 200\}$. The number of factors $R_{nT}$ is set following FreemanWeidner2021 to $R_{nT} = \lfloor 3n^{3/8} \rfloor$. The number of Monte Carlo replications is $M = 10000$, and all results are reported for the normalized estimator $\sqrt{\min(n,T)}(\widehat{\beta} - \beta^*)$.
The results in Tables (ref)--(ref) are broadly similar across all four DGPs, for all estimators considered. This is perhaps surprising for the CCE estimator, which is consistent for $\beta^*$ only under (DGP.1). The bias appears to be of order $\operatorname{\mathcal{O}}(1)$ while the variance diminishes as $n, T \to \infty$, consistent with the theoretical finding that the asymptotic distribution is dominated by approximation bias. The value of $\pi$, which determines $\beta^*$, has a large effect on the bias. The IFE, CCE, and PC(X) estimators are positively biased for larger values of $\pi$, while the PC(YX) estimator has negative bias when $\pi = 0$. An intuitive explanation for the behaviour of PC(X) and PC(YX) can be related to the bias expression in westerlund2015cross and the non-invariance of PC estimators to $\beta^*$ and the correlations between the factor loadings and the errors. Finally, the finite-sample distributions of the estimators appear close to normal, as illustrated in Figure (ref).
This paper shows that two widely used large panel estimators, the PC estimator of GreenawayMcGrevy201248 and the IFE estimator of bai2009panel, have well-defined and interpretable probability limits under fully nonparametric assumptions on the data generating process. Specifically, both estimators converge to the same variance-weighted average of unit-time-specific treatment effects $\beta_{it}$, where the weights are proportional to the conditional variance of the regressor given the unobserved heterogeneity. This result requires no parametric structure on the relationship between the outcome, the regressor, and the unobserved heterogeneity. The key requirement is that the number of estimated factors $R_{nT}$ grows with the sample size, but slowly enough that estimation error remains controlled. This is a meaningful departure from the existing literature, which typically requires $R_{nT}$ to be fixed and the factor structure to be correctly specified. Thus, even if the underlying linear panel regression model with interactive fixed effects is misspecified, one can still attach a clear causal or descriptive meaning to the resulting estimates.
The main limitation of our results is that they apply to the case of a single regressor. When multiple regressors are included, the variance-weighted average interpretation breaks down due to contamination bias, as we discuss in Section (ref). A second limitation concerns inference. Our consistency results are silent on the distribution of the estimators, and we argue in Section (ref) that the asymptotic distribution is dominated by approximation bias, making inference on $\beta^*$ generally impossible without additional parametric assumptions.
This paper benefited from the use of generative AI tools to assist with proofreading, language editing, and \LaTeX formatting. All output was carefully reviewed by the authors. All substantive content, results, and any remaining errors are the authors' responsibility.
The authors declare none.
{2pt}