Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,792 characters · 8 sections · 10 citation commands
\lineskip=1ex \baselineskip 5ex \pagestyle{empty}
\vskip 3em
\noindentAbstract: This paper studies the interactive fixed effects (IFE) estimator in a panel-data setting with heterogeneous treatment effects. We show that, if the treatment-effect heterogeneity admits a linear factor structure, the IFE estimator could fail to recover the average treatment effect on the treated units. The problem arises because the interactive fixed effects absorb the heterogeneity in the treatment effect, creating a bad-control problem. With time-invariant factors or unit-invariant loadings in the treatment effect heterogeneity, identification may further break down due to multicollinearity. These problems are not present in alternative estimation methods that exclude treated units in post-treatment periods from the factor estimation.
\noindentJEL classification: C18, C19, C23, C38, C55\\ \noindentKeywords: bad control; causality; latent factor; potential outcomes; synthetic control
{\noindentAcknowledgments: The authors acknowledge financial support from CAPES, FAPESP (2023/01728-0) and CNPq (302153/2022-5). The usual disclaimer applies.}
\pagestyle{plain}
There are many panel-regression methods that allow for time-varying unobserved heterogeneity in the estimation of average treatment effects on the treated units (ATT). A prominent class of such methods replaces parallel-trends assumptions with the restriction that potential outcomes follow a linear factor structure in the absence of treatment arkhangelsky2024causalmodelslongitudinalpanel. Among the many alternatives in the literature, a natural one is the interactive fixed effects (IFE) estimator that \cn{bai2009panel} introduces in his seminal paper. However, it is not a priori clear whether the IFE estimator indeed recovers the ATT under heterogeneous treatment effects.
There is a vast literature that investigates heterogeneous treatment effects in difference-in-differences (DiD) settings under parallel-trends assumptions. It shows that two-way fixed effects (TWFE) estimators could well recover nonconvex or negatively weighted averages of treatment effects in the presence of treatment-timing variation de2020two,callaway2021difference,goodman2021difference,sun2021estimating,borusyak2024revisiting. In contrast, we know very little about the behavior of the estimators that rely on a factor structure for the potential outcomes when untreated under heterogeneous treatment effects.
We fill this gap by studying the behavior of the IFE estimator in the presence of heterogeneous treatment effects. We show that, in empirically relevant settings, heterogeneity in treatment effects can induce a fundamental identification problem, so that the IFE estimator fails to recover the ATT or any interpretable average of treatment effects. The problem arises if the heterogeneous treatment effects admit a linear factor structure, as in the case that treatment effects depend on common latent shocks that affect all treated units with different intensity levels. For instance, the effect of a job-search assistance program may depend on aggregate unemployment conditions common to all individuals, while personal characteristics determine the magnitude of the response.
In this setting, the interaction between the treatment indicator and heterogeneous effects admits a linear factor structure. As a result, when using post-treatment outcomes of treated units to estimate factors as in the IFE estimator, part of the treatment effect might end up in the factor estimates. If the IFE estimator partials out variation due to the treatment, it would then fail to identify the ATT. There is a bad control problem in that the IFE estimator cannot distinguish between latent factors governing potential outcomes when untreated (that it should control) and latent factors stemming from the heterogeneity in treatment effects (that it should not). As such, it absorbs both components, precluding their individual identification.
The issue becomes more severe if the heterogeneous treatment effect contains either time-invariant factors or unit-invariant loadings. In this case, the interaction between treatment status and the constant factor or loading is perfectly collinear with the treatment indicator, violating the rank conditions for identification in \cn{bai2009panel}. As a result, the IFE estimator is not necessarily consistent for the ATT and, in addition, it could even fail to converge to a unique probability limit.
\cn{xu2017generalized} argues that treatment-effect heterogeneity could well generate bias in the IFE estimator. However, he does not formally characterize when and why heterogeneity causes problems in such a setting. In contrast, we fully establish the conditions under which heterogeneity in treatment effects leads to inconsistency, thereby offering some intuition for the particular forms of heterogeneity that drive these failures. Moreover, we revisit the data-generating process that \cn{xu2017generalized} considers in his Monte Carlo simulations to illustrate concerns about heterogeneous treatment effects. We show that the IFE estimator is inconsistent under a static specification, but consistent under a dynamic specification. We also demonstrate that we recover consistency when the proportion of post-treatment periods or the proportion of treated units becomes negligible. This illustrates that heterogeneity in treatment effects does not necessarily imply inconsistent ATT estimates. Rather, whether inconsistency arises depends on several features, including the form of treatment-effect heterogeneity, the specification used for estimation, and the proportion of treated units and post-treatment periods.
The problems we identify are specific to estimation methods that use post-treatment outcomes of treated units to estimate the factor structure, such as the \pc{bai2009panel} IFE estimator. Alternative approaches, such as synthetic-control and imputation-based methods, do not suffer from this bad-control problem, precisely because they exclude treated units in post-treatment periods from factor estimation. See, among others, \cn{abadie2010synthetic}, \cn{Doudchenko}, \cn{gobillon2016regional}, \cn{xu2017generalized}, \cn{amjad2018robust}, \cn{arkhangelsky2021synthetic}, \cn{athey2021matrix}, \cn{ben2021augmented}, \cn{ferman2021synthetic}, and \cn{porreca2022synthetic}.
Interestingly, the issue we raise here differs markedly from those affecting the TWFE estimator under heterogeneous treatment effects. In the latter case, the problem is that the estimator implicitly relies on comparisons between units at different stages of treatment exposure, including comparisons between units that are newly treated and units that remain treated over the same period. If treatment effects are heterogeneous, such comparisons could induce negative weighting of treatment effects in the presence of variation in treatment timing. In stark contrast, our point about the IFE estimator derives from a bad-control issue. By using post-treatment outcomes of treated units to estimate the factor structure, the IFE estimator could eventually partial out components of the treatment effect itself even in settings without staggered treatment adoption.
We organize the rest of this paper as follows. Section (ref) introduces the framework. Section (ref) derives the main theoretical results. In particular, we first review the IFE estimator under homogeneous treatment effects in Section (ref) and then examine the case of heterogeneous treatment effects and the resulting bad-control and collinearity issues in Section (ref). Section (ref) reports some Monte Carlo simulations to illustrate our point, whereas Section (ref) provides an empirical illustration. Section (ref) concludes.
Consider a panel-data setting in which we observe $i\in\{1,\ldots,N\}$ units at $t\in\{1,\ldots,T\}$ periods. We want to estimate the effect of a policy $d_{i,t}=d_id_t$, where $d_i=\bs{1}(i\in\mathcal{T})$ with $\mathcal{T}$ indicating the set of treated units and $d_t=\bs{1}(t>T_0)$ denoting an irreversible treatment that starts immediately after period $T_0<T$. In addition, we denote the number of post-treatment periods by $T_1=T-T_0<T$. We then define the potential outcome (PO) depending on whether unit $i$ is treated or not at time $t$: $y_{i,t}(1)$ and $y_{i,t}(0)$, respectively. In particular, we assume that the potential outcomes for untreated units admit a linear factor representation:
where $\bs{F}_{0,t}$ and $\bs{\lambda}_{0,i}$ are $k_0\times 1$ vectors of deterministic factors and their corresponding loadings, $e_{i,t}$ is the idiosyncratic error for unit $i$ at time $t$, and $\alpha_{i,t}$ is the (possibly heterogeneous) treatment effect for unit $i$ at time $t$.
We consider a setting conditional on treatment assignment, so that we can treat the latter as fixed. We denote by $N_1$ and $N_0$ the number of treated and control units, respectively. We entertain only deterministic heterogeneous treatment effects, in that $\alpha_{i,t}$ is a fixed parameter, even if it potentially changes across units and over time alvarez2025inference. The parameter of interest is the average treatment effect on treated units (ATT):
The main challenge in the estimation of $\bar\alpha$ is that we observe realized (rather than potential) outcomes: namely, $y_{i,t}=(1-d_{i,t})\,y_{i,t}(0)+d_{i,t}\,y_{i,t}(1)$. We next summarize the sampling scheme and exogeneity conditions in Assumptions (ref) and (ref).
While we consider a setting in which treatment allocation and factor structure are fixed, we could alternatively rewrite the framework as conditional on them. In this case, we would capture the characteristics of the treatment assignment mechanism by the distribution of potential outcomes conditional on treatment allocation. For example, if the units select into treatment based on unobservable (to the econometrician) factors, then the conditional distribution of these unobservables given treatment allocation depends on treatment status \ca{Ferman_JASA}{see discussion in}. In addition, we would have to formulate Assumption (ref) in terms of conditional moments to ensure that treatment assignment is as good as random once we control for the factor structure. Finally, it is worth saying that we require iid errors just for convenience \ca{bai2009panel}{see, for instance, discussion in}.
The linear factor structure for the potential outcomes when untreated implies that the DiD estimator is not necessarily unbiased for $\bar\alpha$, except in special cases such as if $\bs{\lambda}_{0,i}'\bs{F}_{0,t}=\theta_i+\gamma_t$. The two-way fixed effects (TWFE) allows for confounders $\theta_i$ that are unit-specific but constant over time and time-varying confounders $\gamma_t$ that affect all units in the same way. This specification rules out the possibility of time-varying unobserved confounders that differ across units, though.
In this section, we discuss the estimation of $\bar\alpha$ using the IFE estimator, as well as some alternative estimators that assume a linear factor model for the potential outcomes.
We first examine the identification and estimation of treatment effects using the IFE framework under the assumption that the treatment effect is homogeneous across units and periods: $\alpha_{i,t}=\alpha_0$ for all $i$ and $t$. Let now $\bs{Y}_i=(y_{i,1},\ldots,y_{i,T})'$, $\bs{D}_i=(d_{i,1},\ldots,d_{i,T})'$, $\bs{F}_0=(\bs{F}_{0,1}',\ldots,\bs{F}_{0,T}')'$, $\bs{\Lambda}_0=(\bs\lambda_{0,1}',\ldots,\bs\lambda_{0,N}')'$, and $\bs{e}_i=(e_{i,1},\ldots,e_{i,T})'$. We can then represent the PO model in (ref) as
\cn{bai2009panel} defines the IFE estimator as
where $\bs{M}_{\bs{F}}=\bs{I}_{T}-\bs{F}(\bs{F}'\bs{F})^{-1}\bs{F}'=\bs{I}_{T}-\bs{F}\bs{F}'/T$. It differs from the least-squares estimator only because it must also retrieve the latent factor structure. Implementation rests on an interactive procedure that estimates both the parameter of interest and the factor structure. \cn{bai2009panel} establishes consistency of $\hat\alpha$ for $\alpha_0$ as $N,T\rightarrow\infty$, under Assumptions (ref) and (ref), and the regularity conditions (ref) in Appendix (ref). Although his result treats the dimension of the latent factors as known, \cn{moon2015linear} demonstrate that estimating the factor dimension does not affect the asymptotic behavior of $\hat\alpha$ as long as we do not underestimate the number of factors.
There are alternative methods to estimate causal effects in the event that potential outcomes follow a linear factor structure. See, for instance, the synthetic-controls (SC) approach by \cn{abadie2010synthetic}, the generalized synthetic-controls (GSC) framework by \cn{xu2017generalized}, the demeaned synthetic-controls (DSC) estimator by \cn{Doudchenko} and \cn{ferman2021synthetic}, and the synthetic difference-in-differences (SDiD) approach by \cn{arkhangelsky2021synthetic}. In contrast to the IFE estimator, they do not rely on treated units in the post-treatment period to estimate the factor structure.
We next relax the homogeneity assumption on the treatment effects by assuming they admit a linear factor representation: namely, $\alpha_{i,t}=\gamma+\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}$, where $\gamma$ represents the common treatment effect, and $\bs{F}_{\alpha,t}$ and $\bs\lambda_{\alpha,i}$ denote the $k_\alpha-$vectors of latent factors and loadings. The main problem arises from the fact that the PO model under heterogeneous treatment effects admits a linear factor representation that combines both $\bs{F}_{0,t}$ and $\bs{F}_{\alpha,t}$:
This happens because $d_{i,t}\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}+\bs\lambda_{0,i}'\bs{F}_{0,t}=(d_i\bs\lambda_{\alpha,i})' (d_t\bs{F}_{\alpha,t})+\bs\lambda_{0,i}'\bs{F}_{0,t}$, giving way to the extended factor representation in (ref), with $\bb F_t=(d_t\bs{F}_{\alpha,t}^\prime,\bs{F}_{0,t}^\prime)^\prime$ and $\bb\lambda_i=(d_{i}\bs\lambda_{\alpha,i}^\prime,\bs\lambda_{0,i}^\prime)^\prime$ such that $\bb\Lambda'\bb\Lambda$ is diagonal. We next impose some high-level conditions on this extended linear factor structure with factors $\bb F_t$ and loadings $\bb\lambda_i$.
Although we restrict attention to a deterministic factor structure, assuming a stochastic idiosyncratic component to the heterogeneity in treatment effects would not affect the results. It is just for simplicity that we do not entertain this case. Assumption (ref)$(i)$ is just a normalization, given that we can always find $\bb F$ and $\bb\Lambda$ that satisfy these conditions bai2009panel. Note that $\breve{k}\le k_0+k_\alpha$ because the dimension of the linear subspace spanned by $\bb F$ is at most the sum of the dimensions of the linear subspaces spanned by $\bs F_0$ and $\bs F_\alpha$. For instance, if $Y(0)$ contains factors that are active only for treated units in post-treatment periods, some components of $(d_i\bs{\lambda}_{\alpha,i})'(d_t\bs{F}_{\alpha,t})$ could then lie in the factor space of $Y(0)$, thereby violating Assumption (ref)$(ii)$. We would have to drop the redundant factors so that the resulting extended factor structure $(\bb F,\bb\Lambda)$ would have a dimension smaller than $k_0+k_\alpha$. This does not affect the proof or the conclusions of Proposition (ref), though.
Assumption (ref)$(ii)$ implies that factors and loadings in the extended factor representation asymptotically generate variation in the outcomes. This assumption would not hold, for example, if the proportion of treated units and/or periods becomes negligible as the sample size grows (see Remark (ref)). Finally, Assumption (ref)$(iii)$ ensures that the extended factor structure is not collinear with the treatment dummy. We discuss in Remark (ref) two important cases in which this assumption could fail.
We next document that the minimization problem in (ref) does not result in a consistent estimator for $\bar\alpha$.
Intuitively, the objective function increases with the heterogeneity among treated units, so the minimizer forces the factor structure to absorb it entirely. A direct implication of Proposition (ref) is that the IFE estimator $\hat\alpha$ is not a consistent estimator of the average treatment effect on the treated (ATT). A useful way to interpret this result is that heterogeneous treatment effects generate a factor structure that appears only for treated units in the post-treatment period. Because the IFE estimator searches for latent factors that explain systematic variation in outcomes, it attributes this pattern to additional factors rather than to treatment effects. As such, the factor structure we estimate absorbs part of the heterogeneous treatment effects, effectively acting as a bad control that removes this variation from the treatment effects. This explains why the IFE estimator converges to $\gamma$ rather than to the ATT parameter.
\pc{xu2017generalized} simulations illustrate potential problems with the IFE estimator under heterogeneous treatment effects of the form $\alpha_{i,t}=\psi_t+a_{i,t}$, where $a_{i,t}$ is purely idiosyncratic. An interesting implication of Remark (ref) is that, while the static IFE estimator is indeed inconsistent under this data-generating process (DGP), the IFE estimator remains consistent in the dynamic specification. This highlights that whether heterogeneous treatment effects lead to inconsistency depends not only on the form of the heterogeneity, but also on the specification used for estimation.
To illustrate how factors affect the estimation of ATT under heterogeneous treatment effects, we carry out a Monte Carlo study using the PO model in (ref), with (ref) as the parameter of interest. We entertain two data generating processes. The first considers a homogeneous treatment effect in which $\alpha_{i,t}=2$ for all units and time periods, implying $\bar \alpha = 2$. The second features heterogeneous treatment effects $\alpha_{i,t}=\gamma+\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}$, with $\gamma=1$ and $\frac{1}{N_1T_1}\sum_{i\in\mathcal{T}} \sum_{t=T_0+1}^T\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}=1$, so that $\bar\alpha=2$. For more details on the data generating processes and the factor structure of the heterogeneous treatment effects, see the notes of Table 1 as well as Appendix (ref).
We consider five ATT estimators: IFE, SC, DSC, GSC, and SDiD. They are all consistent for the DGP with homogeneous treatment effects. Under heterogeneous treatment effects, however, the IFE estimator captures only the common treatment-effect component $\gamma$, while the PO factor structure soaks up $\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}$ as a bad control (see Section (ref)). In contrast, this problem does not arise for the other estimators, since they do not use treated units in post-treatment periods to estimate the factor structure (see Remark (ref)).
Table (ref) displays the simulation results. For the DGP with homogeneous treatment effects, all estimators are, on average, close to the ATT $\bar\alpha=2$. The IFE estimator performs slightly better in terms of variance relative to the other alternatives. This should not come as a surprise, since it uses more information: namely, the treated units in the post-treatment periods—to estimate the factor structure. Importantly, the assumption that treatment effects are homogeneous is crucial for the IFE estimator to recover only the factor structure of the untreated potential outcomes, rather than inadvertently controlling for part of the treatment effect.
Once we turn to heterogeneous treatment effects, the picture changes dramatically. The IFE estimator completely fails to recover the ATT parameter $\bar\alpha=2$, retrieving instead $\gamma=1$. Following the intuition from Section (ref), the treatment-effect component $\bs\lambda_{\alpha,i}'\bs{F}_{\alpha,t}$ is now absorbed into the PO factor structure rather than into the treatment effect as it should. In contrast, the alternative estimators remain close to the true ATT value of $\bar\alpha=2$.
To illustrate empirically the bad-control issue in the IFE estimation, we revisit the study by \cn{gobillon2016regional} about the effect of the Enterprise Zone Program on the unemployment rate. In particular, we assess how the Enterprise Zone Program affects the exit rates from unemployment, either to a job or for unknown reasons, and the entry rate into unemployment (all on a logarithmic scale). The analysis is for the Paris region, with filters to match the sample in \cn{gobillon2016regional}, resulting in 135 control units and 13 treated units for 7 pre-treatment and 13 post-treatment periods.
For each outcome, we compare the IFE estimator of $\bar\alpha$ with the alternative SC, DSC, GSC and SDiD estimators. Although the latter approaches yield robust estimators to heterogeneous treatment effects, they rely on different assumptions for consistency. As such, they could well converge to different probability limits in the event some of these assumptions fail. For instance, the SC, DSC, and SDiD approaches impose convex weighting constraints on the factor loadings. In particular, the standard SC estimator is the most restrictive, calling for treated units' factor loadings lying in the convex hull of the donors’ loadings for both time-variant and time-invariant factors. Both DSC and SDiD approaches relax this assumption by previously removing time-invariant fixed effects, so that treated units must lie in the convex hull of the factor loadings, excluding those corresponding to time-invariant factors. In view that the GSC estimator does not require convex-hull assumptions, we take it as the main benchmark to assess whether the IFE estimator is indeed recovering the ATT parameter.
\newcolumntype{P}[1]{>{\arraybackslash}p{#1}}
Table (ref) summarizes our empirical findings. We conduct inference by employing a permutation-based placebo test under the sharp null hypothesis of no treatment effect. This means that, at each permutation, we randomly select $N_1$ units as pseudo-treated before computing the ATT estimators and their corresponding test statistics. We denote the latter by $\theta_r^k$ for $k\in\{\text{IFE},\text{SC},\text{DSC}, \text{GSC},\text{SDiD}\}$ and $r\in\{1,\ldots,R\}$. We then compute the p-value of $\theta_r^k$ using the empirical distribution across $R=10{,}000$ permutations, i.e., $\frac{1}{R}\sum_{r=1}^R\mathbb{I}(\theta_r^k>\theta_{\bar{\alpha}}^k)$, where $\theta_{\bar{\alpha}}^k$ is the sample value of the test statistic using the observed data. For the SC and DSC estimators, we use the mean squared prediction error ratio as the test statistic to adjust for imperfect pre-treatment fit abadie2010synthetic,abadie2015comparative,firpo2018synthetic. In turn, we consider the absolute treatment effect $|\hat{\alpha}_r^{k}|$ as the test statistic for the other estimators.
For the exit rate to a job (Column 1), the IFE estimator departs from all other estimators, indicating a significant positive ATT in contrast to the negative signs of the other ATT estimates. In particular, it differs sharply from the GSC, which suggests the presence of treatment effect heterogeneity. It is also interesting to observe that the GSC approach yields a much more prominent ATT than the SC, DSC and SDiD estimators, perhaps suggesting that the convex-hull conditions do not hold.
The latter does not seem a problem for the estimation of the ATT of the Enterprise Zone Program on the exit rate for unknown reasons in that most estimates are close to 5% and mostly significant (Column 2). The only exception is the IFE approach, which yields an ATT estimate very close to zero. This pattern strongly suggests that the convex hull restrictions for the validity of the SC, DSC, and SDiD estimators hold, whereas the heterogeneous treatment effects prevent the IFE estimator from recovering the ATT. Finally, IFE and GSC estimators for entry rate are close to zero and not statistically significant at standard significance levels, which might suggest that treatment effects are homogeneous and equal to zero for this outcome (column 3).
We examine the behavior of \pc{bai2009panel} IFE estimator under heterogeneous treatment effects. We find that, if the heterogeneity in treatment effects admits a linear factor representation, the IFE estimator fails to recover the average treatment effect on the treated. This happens because the interactive fixed effects in the PO model absorb the factors in the heterogeneity, thereby becoming a bad control in that it accounts for too much. We also show that, in the presence of time-invariant components in the heterogeneous treatment effects, identification breaks down due to multicollinearity. In contrast, these issues do not arise for alternative estimators that are robust to heterogeneous treatment effects, essentially because they exclude treated units in post-treatment periods from the factor estimation in the PO model.
Empirically, we revisit \pc{gobillon2016regional} analysis about how the Enterprise Zone Program affects the unemployment rate in the Paris region. The IFE estimator yields very different results relative to the generalized synthetic-control approach in two out of three outcomes analyzed, indicating that it is likely assigning the heterogeneity factors to the interactive fixed effects in the PO model, rather than in the treatment effect.
\lineskip=1ex \baselineskip 3.5ex
\baselineskip 4.5ex