Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,911 characters · 17 sections · 25 citation commands
Forecasted Treatment Effects with Short Panels
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Polynomial regressions; Forecast unbiasedness; Counterfactuals; Misspecification; Heterogeneous treatment effects
\spacingset{1.6}
Evaluating the effect of a policy or a treatment usually requires comparing treated and untreated units. Standard approaches such as difference-in-differences (DiD) or two-way fixed effects (TWFE) regressions rely on outcomes of untreated units to identify the counterfactual outcomes of treated units. These methods rely on assumptions such as unconfoundedness or parallel trends (e.g., GoodmanBacon, CallawaySantAnna, SunAbraham). They cannot be applied when all units receive the treatment.
This situation is common. Examples include national tax reforms, health programs, and environmental regulations that affect all individuals or regions at the same time. Even when adoption is staggered, the last periods after full treatment do not contain untreated units, so conventional DiD or event-study designs cannot recover treatment effects for those periods. In such cases, researchers often turn to structural models or to simple before–after comparisons. These settings call for methods that make the required assumptions explicit and allow researchers to evaluate their plausibility.
This paper studies how to estimate treatment effects when no valid control group exists. We propose a transparent and practical approach based on forecasting each treated unit’s counterfactual outcome from its own pre-treatment data using basis function regressions. The key idea is that in short panels, the property required for identification is forecast unbiasedness rather than forecast accuracy. Forecasting methods in time-series and machine learning aim to minimize prediction error and often allow some bias to reduce variance. In policy evaluation, this objective is not appropriate. An unbiased forecast of the counterfactual outcome guarantees an unbiased estimate of the treatment effect, while a forecast that is accurate but biased does not. Of course, our approach relies on its own identifying assumptions, which we make explicit, but they differ from those used in designs that compare treated and untreated units.
Our approach, the Forecasted Average Treatment effect (FAT) estimator, computes the average treatment effect on the treated as the average difference between the observed post-treatment outcome and an unbiased forecast of the outcome in the absence of treatment. What matters is that the forecasting rule produces unbiased predictions on average across individuals.
The paper makes three contributions. First, we develop a simple and general framework for estimating treatment effects when no untreated units are available. Under appropriate assumptions, the FAT estimator is consistent and asymptotically normal even when the time dimension is short, the panel is unbalanced, and treatment effects are heterogeneous.
Second, we characterize a broad class of data generating processes (DGPs) under which the required forecast unbiasedness condition holds. When individual counterfactual outcomes can be generally expressed as the sum of (at most) three unobserved components - a stationary process, a unit root process, and a deterministic trend - unbiased forecasts can be obtained by regressing pre-treatment outcomes on basis functions of time, such as low-order polynomials. This result allows for DGPs with fixed effects, heterogeneous autoregressive parameters, and non-stationary trends, without requiring a detailed specification of the stochastic component. The framework also accommodates heterogeneous dynamic processes, such as unit-specific autoregressive parameters, which are difficult to handle with standard short-panel methods. We call our proposed approach “Unobserved-Components FAT”.
Third, we compare our general “Unobserved-Components FAT” approach with what we refer to as “Parametric-Model FAT”, where the researcher specifies and estimates a particular parametric model such as an AR(1) to forecast counterfactuals. This approach relies on stronger assumptions and can perform poorly in short panels because of misspecification or estimation bias. In contrast, our regression-based approach that focuses directly on forecast unbiasedness is robust to specification errors and is not affected by the incidental parameter problem.
The FAT estimator can also be used when untreated units exist but are imperfect comparators. In such cases, FAT can provide a robustness check for studies that rely on, e.g., parallel-trends assumptions. The framework complements recent work on DiD and event-study designs (e.g., deChaisemartinX2020, CallawaySantAnna, SunAbraham) by providing an alternative benchmark based on explicit assumptions about the individual-specific counterfactual process rather than on cross-group comparability.
Crucially, the approach based on FAT is advantageous when outcomes exhibit dynamics. Recent work MarxTamerTang, Klosin2024, BotosaruLiu2025, Cornwall2025 has shown that traditional TWFE and dynamic panel estimators often yield biased or negatively weighted estimates of the treatment effect in the presence of serial correlation. The solutions proposed in this literature typically rely on the homogeneity of the persistence parameters across units. By construction, FAT avoids conflating structural treatment effects with outcome persistence and is robust to feedback mechanisms Bonhomme2025, while allowing for a much richer class of models, which include heterogeneous forms of serial correlation, where the degree of outcome persistence can vary arbitrarily across individuals.
This paper relates to several strands of literature. It is thematically related to work on causal inference in settings where contemporaneous untreated comparison groups are not the primary source of identification, and where time-series structure plays a central role in constructing counterfactual outcomes from pre-treatment data. These include Bayesian approaches to causal forecasting (e.g., Brodersen2015); (comparative) interrupted time-series methods, in which untreated potential outcomes are modeled using finite-dimensional parametric mean functions in time—typically linear or piecewise-linear trends with level and/or slope changes at the intervention date (e.g., Schochet2022,Bernaletal2017,BrownWarner1985); and methods based on data-driven aggregation to purge unobserved aggregate confounders in designs with aggregate shocks (e.g., ArkhangelskyKorovkin2023).
The FAT framework differs from these strands along three dimensions. First, relative to Bayesian causal forecasting, FAT is designed to deliver forecast-unbiased counterfactuals under high-level restrictions on the untreated outcome process and does not require the researcher to commit to a fully specified likelihood or prior. Second, relative to interrupted time-series methods, FAT strictly nests the canonical segmented-regression mean specifications while allowing for richer latent evolution, including stochastic trend components and heterogeneous unit-specific dynamics. Third, relative to the aggregate-shock/aggregation literature, FAT does not take identification from exogenous aggregate shocks or instruments, nor does it rely on cross-sectional weighting to eliminate unobserved aggregate confounding; instead, it is a short-panel forecasting approach in which identification is driven by the internal time-series information in treated units’ pre-intervention histories.
In addition, the paper contributes to the panel-data literature on heterogeneous dynamic processes and random coefficients (e.g., Chamberlain1992,ArellanoBonhomme2012,GrahamPowell2012). For example, the projection estimator in ArellanoBonhomme2012 can be interpreted as a special case of FAT under specific restrictions on the latent outcome process. In particular, for a unit-root process without drift, the Arellano–Bonhomme forecast corresponds to an equal-weighted average of pre-treatment outcomes, whereas the FAT polynomial forecast assigns greater weight to more recent observations. When deterministic trends are present, the projection approach in ArellanoBonhomme2012 requires correct specification of the trend component, while FAT only requires the choice of a sufficiently rich family of basis functions, such as low-order polynomials. Therefore, FAT delivers unbiased counterfactual forecasts under weaker restrictions on the trend component, while retaining the projection-based intuition underlying existing random-coefficient panel estimators.
Our analysis also relates to work on forecast unbiasedness in time-series models (e.g., FullerHasza, Dufour1984) and to the dynamic-panel literature on bias correction and consistent estimation with short time dimensions (e.g., AngristPischke,BlundellBond1998). Our approach differs by showing that consistent treatment effect estimation in short panels can be achieved through unbiased forecasts rather than through the specification of a full stochastic model.
The rest of the paper is organized as follows. Section (ref) presents the framework and the FAT estimator. Section (ref) extends the method to include covariates and discusses forecasting based on a parametric model. Section (ref) reports simulation results. Section (ref) provides an empirical illustration. Section (ref) concludes.
In this section we develop our baseline “Unobserved-Components FAT” estimator for the average treatment effect on the treated (ATT) under universal treatment. We use individual pre‑treatment time series to forecast each unit’s post‑treatment counterfactual outcome and then average, across units, the difference between the observed outcome and its forecast. We first establish consistency and asymptotic normality of this estimator under a high‑level forecast‑unbiasedness condition, and then show that this condition holds for a broad class of unobserved‑components DGPs when counterfactuals are forecast using basis‑function regressions.
Consider a treatment or a policy that is implemented at a time $\tau$.\footnote{In Section (ref), we discuss individual-specific timing.} Here, the treatment affects all individuals in the population at the same time, so that the treatment indicator of individual $i$ at time $t$ is given by $$d_{it}:=\ 1(t>\tau) \text{ for all } i=1,\dots, n.$$ We adopt the potential outcomes framework with each individual $i$ having two potential outcomes at each time $t$: $y_{it}(1)$ if the individual is exposed to the treatment and $y_{it}(0)$ if the individual is not exposed to the treatment. Due to the absence of an untreated group in our setting, we will henceforth simply refer to $y_{it}(0)$ as the “counterfactual".
Under the stable unit treatment value assumption (SUTVA), the observed outcome of individual $i$ at $t$ is:
When all individuals are treated after time $\tau$, we have
We follow the literature on heterogeneous treatment effects in defining the ATT $h\geq 1$ periods after $\tau$ as:
where we used that $y_{i\tau+h}(1)=y_{i\tau+h}$ for $h\geq1$.\footnote{Note that with identically and independently distributed data across $i$, the right hand side of (ref) reduces to the conventional $\mathbb{E}\left[y_{i\tau+h}-y_{i\tau+h}\left(0\right)\right]$.}
The challenge in identifying and estimating ${\rm ATT}_h$ is that the counterfactual $y_{i\tau+h}\left(0\right)$ is not observed for $h \geq 1$. The conventional approach in the presence of an untreated group is to impose sufficient assumptions that identify the parameter of interest from the observed post-treatment outcomes of the untreated group. In the absence of an untreated group, we exploit pre-treatment individual time series to obtain a forecast for $y_{i\tau+h}\left(0\right)$. We denote this forecast by $\widehat{y}_{i\tau+h}(0)$.
We call our proposed estimator for ${\rm ATT}_{h}$ the Forecasted Average Treatment effect estimator (FAT), defined as:
where $\widehat{y}_{i\tau+h}(0)$ is a measurable function of past outcomes $\left\{y_{it}\right\}_{t\leq \tau}$.\footnote{In the baseline case, the individual forecast depends only on the past outcomes of the treated, in particular, there are no covariates in the information set.} We explain how to obtain the forecast $\widehat{y}_{i\tau+h}(0)$ below. For now, we note that $\widehat{y}_{i\tau+h}(0)$ uses individual-specific pre-treatment outcomes, which naturally accommodates unbalanced panels and heterogeneous treatment effects.
We make the following high-level assumptions.
Note that Assumption (ref) guarantees that $\mathbb{E}(\widehat{{\rm FAT}}_{h}) = {\rm ATT}_{h}$ so that ${\rm ATT}_{h}$ is identified.\footnote{The result follows trivially by writing ${\rm ATT}_{h} =\frac{1}{n}\sum_i\mathbb{E}\left[y_{i\tau+h}-\widehat{y}_{i\tau+h}(0) + \widehat{y}_{i\tau+h}(0) - y_{i\tau+h}\left(0\right)\right].$}
Let $u_{i\tau+h}:=y_{i\tau+h}-\widehat{y}_{i\tau+h}(0)$ be the forecasted individual treatment effect at $\tau+h, \; h\geq 1$.
For example, when $\{u_{i\tau+h}\}$ is a sequence of cross-sectionally independent (but not identically distributed) random variables, Theorem 5.11 in White gives an asymptotic normality result.
What Assumptions 1 and 2 fundamentally rule out is shocks that affect all individuals after the treatment and that are unforecastable.\footnote{In principle, the assumptions allow for some common shocks to be captured by the method used to forecast the counterfactuals. For example, if the common shocks are deterministic and can be modeled as polynomials, our baseline method will be able to account for it.} This is an example of the assumptions that any method for treatment effect evaluation must inevitably make in the absence of a valid control group. The assumptions are more likely to be valid the higher the frequency at which the data is measured and the closer to the treatment date one evaluates the effect (i.e., they more plausibly hold for monthly data and for effects evaluated one month after the treatment than for yearly data and effects evaluated several years after the treatment). Below we provide some discussion of how the assumption of no common shocks after the treatment could be weakened. As one would expect, the assumption can be relaxed when there is a group of untreated individuals that are subject to the same unforecastable shock, as we discuss in the Online Appendix (Section (ref)). Perhaps less obviously, the assumption could also be relaxed in the absence of an untreated group, as long as we observe another variable for the treated units that is subject to the same unforecastable shock but not subject to the treatment (see the discussion in Section (ref)).
In the remainder of the paper, we provide low-level sufficient assumptions, including a full description of the class of DGPs for the counterfactuals $y_{it}(0)$ and the forecast methods which satisfy Assumption (ref).
In this section, we characterize the class of DGPs for the counterfactuals. The need to discuss the DGP for counterfactuals arises because of the lack of an untreated group, which means that we must rely on forecasting counterfactuals from pre-treatment observations.
In this section, we consider a class of DGPs such that Assumption (ref) is satisfied generally, namely, by any forecast that can be written as a weighted average of pre-treatment outcomes with weights summing to 1. The DGPs in this class express the counterfactual as the sum of potentially two unobserved stochastic components. This includes a variety of processes, such as stationary and non-stationary (unit root) ARMA processes with individual-specific parameters.
A key insight of this paper is that a correctly specified parametric model (e.g., a specific ARIMA model) is not necessary to obtain unbiased forecasts of the counterfactuals. In fact, as the next result shows, any forecast expressed as a weighted average of pre-treatment observations satisfies the unbiasedness condition.
There are many ways to obtain forecasts that are weighted averages of pre-treatment data. In this paper we focus on a general class of forecasts obtained via basis function regressions, such as polynomial time trends regressions.
This definition makes it clear that for any type of basis function the choice $q_i=0$ yields the sample mean of pre-treatment outcomes as the forecast. The following result shows that forecasts obtained via basis function regressions satisfy the weighted average requirement of Theorem (ref).
Our framework connects to the literature on random coefficient panel models, particularly Chamberlain1992, GrahamPowell2012, and ArellanoBonhomme2012. Linear panel models with unit-specific intercepts and trend coefficients, similar to specifications in this literature, belong to the class of DGPs in Assumption (ref) or Assumption (ref). However, our goals differ: that literature studies identification of the distribution of heterogeneous coefficients, whereas we focus on forecasting untreated potential outcomes to estimate average treatment effects.
In this section, we consider an expanded class of DGPs that is appropriate for applications where: 1) there is more than one pre-treatment outcome; 2) it makes sense to model the outcomes as trending over time; 3) the trend is deterministic rather than (or in addition to) stochastic. We show that the basis function regression considered in the previous section gives unbiased forecasts of the counterfactuals, under certain conditions.
The expanded class of DGPs always includes a deterministic trend component, possibly in addition to (either or both) the stochastic components considered in Assumption (ref).
Theorem (ref) below clarifies when forecasts obtained via basis function regressions satisfy the unbiasedness assumption (Assumption (ref)) in the presence of deterministic trends.
Our proposed method for forecasting counterfactuals in Definition (ref) requires choosing: 1) the number of pre-treatment periods used for the estimation $R_i$; 2) the type of basis functions $b_k(t)$; and 3) the number of basis functions $q_i$. We discuss how these choices affect inference and offer some practical recommendations for empirical researchers.
The choice of estimation window, $R_i$, involves a trade-off. In our class of DGPs, a larger $R_i$ generally yields an estimator with smaller variance, but a shorter $R_i$ can guard against violations of our assumptions stemming from parameter instability in pre-treatment data. While plotting pre-treatment time series offers informal guidance on such instability, applying our method to long-T settings reveals a heightened sensitivity to the chosen estimation window. This is because the increased likelihood of structural changes and parameter instability over extended periods becomes a critical concern. This challenge resonates with discussions in the regression discontinuity literature regarding bandwidth selection (e.g., GelmanImbens2019).
Regarding the choice of basis functions, Theorems (ref) and (ref) imply that this choice only matters for unbiasedness when the DGP has a deterministic time trend, in which case the basis functions need to be correctly specified to ensure unbiasedness (up to the order). When the DGP is mean stationary or has a stochastic trend, the choice of basis functions does not matter for unbiasedness. Basis functions may be chosen based on the time series properties of pre-treatment outcomes. Polynomial time trends seem to be a natural choice of basis functions for DGPs with deterministic trends.\footnote{In applications using difference-in-differences (DiD) methods it is typical to assume the presence of time trends (mostly linear) that are common between control and treatment groups. Our results make it clear that a linear time trend can only be dealt with by either using an untreated group or, when an untreated group is not available, by using (at least two) pre-treatment time periods to model the trend (which leads to our polynomial regression).} Other basis functions could be used, e.g., periodicity could be captured by Fourier basis functions. Our practical recommendation, and what we focus on henceforth, is to consider polynomial basis functions by letting $b_k(t)=t^k$ in Definition (ref).\footnote{An advantage of polynomial basis functions is that it is in principle possible to relax the assumption that the basis functions are correctly specified by instead assuming that $y_{it}^{(3)}(0)$ in Assumption (ref) is a continuous (but otherwise unspecified) function of time. Since time in our setting is defined on a compact interval and the deterministic trend is a continuous function, $y_{it}^{(3)}(0)$ can be approximated arbitrarily well by a polynomial in time. In fact, by the Weierstrass Approximation Theorem, the approximation error approaches zero as the order of the polynomial goes to infinity. Under additional smoothness assumptions on the deterministic trend, an approximation theorem could then be used (e.g. the Polynomial Approximation Error Theorem) to derive a bound on the approximation error of $y_{it}^{(3)}(0)$ by the polynomial regression. The forecast of $y_{i\tau+h}(0)$ is biased, but we conjecture that it may be possible to do bias correction given an expression for the bias obtained via the approximation theorem.}
Regarding the choice of the order of the basis functions $q_i$ used for the estimation, this only matters for unbiasedness if the DGP has a deterministic trend component. For DGPs without such a trend, any $q_i$ ensures unbiasedness, with $q_i=0$ offering the lowest variance; however, larger $q_i$ can mitigate bias from non-stationary initial conditions. If a deterministic trend exists, $q_i$ cannot be smaller than the unknown true order, forcing a trade-off where larger $q_i$ for unbiasedness risks higher variance. This choice is inherently constrained by the number of pre-treatment periods ($T$): larger $T$ allows for more flexible polynomial forecasts and extensive placebo testing, yet estimation consumes periods from this budget; conversely, smaller $T$ drastically limits both model flexibility and the scope for placebo tests.
Standard cross-validation methods typically target predictive accuracy, not the unbiasedness crucial for our context. While bias-targeting cross-validation is a promising avenue for selecting $q_i$, its implementation faces challenges in our short-T panel context. It inherently requires long pre-treatment series for data splitting, which is often not feasible. Moreover, even with long-T data, this can heighten concerns about parameter instability (as stationarity becomes less defensible over extended periods). Therefore, our practical recommendation is to report results for a small range of $q_i$ values (e.g., $0,1,2,3$), with pre-treatment time series plots providing informal guidance.
The treatment timing $\tau_i$ can be individual-specific. While staggered adoption designs typically allow for identification via comparisons with not-yet-treated units, such strategies are inherently limited to the window before the last unit becomes treated. In settings where all units eventually adopt the treatment, such as the universal rollout of a national policy, standard event-study designs cannot recover treatment effects for the periods following full adoption, as no valid comparison group remains.
Our approach overcomes this limitation. Because FAT constructs counterfactuals using only the treated unit's own pre-treatment history, it does not rely on the existence of a contemporaneous control group. Consequently, it allows researchers to estimate treatment effects over longer horizons, extending into periods after all units have been treated.
We require that the treatment timing be exogenous with respect to the counterfactual outcome path $y_{it}(0)$, conditional on the unit's history used for forecasting. The basis functions $\{b_k(t)\}$ in Definition (ref) are then functions of the time to adoption $t - \tau_i$. Additionally, it is possible to allow for treatment anticipation, as long as it is limited. In this case, one can simply modify the pre-treatment estimation window ${\cal T}_{i}$ in Definition (ref) to include observations only up to the time $\tau_i-\delta_i$ at which it is still reasonable to assume that there was no treatment anticipation, that is, ${\cal T}_{i} \equiv \left\{ \tau_i-\delta_i-R_{i}+1,\ldots,\tau_i-\delta_i\right\}$ (and $h$ is adjusted accordingly).
For balanced panels, our proposed estimator is algebraically equivalent to two simpler pooled estimation strategies. These alternative approaches are also suitable for repeated cross-section data or when cohort of birth plays the role of time. However, in unbalanced panels, these simpler strategies may yield inconsistent estimates for ATT. Therefore, our exposition primarily focuses on individual-level estimators, given their ability to accommodate unbalanced panels; these same pooled estimators, however, are how the FAT is obtained in balanced panels.
Assume that $R=R_{i}$ and $q=q_{i}$ are constant across $i$ and focus on polynomial basis functions in Definition (ref). The first alternative way to obtain our proposed estimator is to consider the cross-sectional averages $\overline{y}_{t}=\frac{1}{n}\sum_{i=1}^{n}y_{it}$ of the observed outcomes in time period $t$. Due to linearity of the forecasting procedure, we can rewrite:
where ${\cal T}=\{\tau-R+1,\ldots,\tau\}$, and we suppress the dependence on $\tau$, $q$, $R$. Here, the cross-sectional averages for $t\leq\tau$ are used to obtain a forecast of the average counterfactual for $t=\tau+h$, which is then subtracted from the cross-sectional average observed at that time.
The second alternative is to consider a pooled regression estimator, $\widehat{{\rm FAT}}_{h}=\widehat{\beta}_{h}$, where
which is the OLS estimator obtained from regressing $y_{it}$ on a set of time dummies $1(t=\tau+k)$, for $k\in\{1,\ldots,h\}$, and individual-specific time trends.\footnote{It actually does not matter for $\widehat{\beta}$ whether the coefficients $\alpha$ on the time trend are individual-specific.}
In this section, we consider an alternative approach that explicitly incorporates covariates, including lagged outcomes, in the estimation of FAT. Compared to the baseline case, this involves specifying a parametric model for the counterfactuals and imposing additional assumptions.
Consider the following parametric model for the counterfactual $y_{it}(0)$:
where $x_{it}\in\mathbb{R}^{\dim x_{it}}$ is a vector of covariates, $\beta\in\mathbb{R}^{\dim x_{it}}$ is restricted to be homogeneous across individuals, $c_{ik} \in \mathbb{R}$ are unknown parameters and $\varepsilon_{it}\in\mathbb{R}$ is such that
Assuming that we have estimates $\widehat{\beta}$ for the common parameter that are consistent as $n\rightarrow\infty$ under correct model specification\footnote{For example, when $q_{i}=0$ and $x_{it}=(y_{it-1},z'_{it})'$, a consistent estimator for $\beta=(\rho,\theta')'$ can be obtained by applying an IV regression to the first-differenced model
using, for example, $y_{it-2}$ and $z_{it-1}$ as instruments. In the Monte Carlo simulations, we further extend this case to $q_{i}=1$.} leads to the parametric-model FAT, defined as:
where $\widehat{y}_h(\widehat{\beta},y_i,x_i)$ is the forecast obtained as:
where $c_{i}=(c_{i,0},\ldots,c_{i,q_i})$ is a $q_i+1$ vector, and ${\cal T}_i=\{\tau-R_i+1,\ldots,\tau\}$ is the set of the $R_i$ time periods directly preceding the treatment date. The parameters $q_i$ and $R_i\in\{q_i+1,\ldots,\tau\}$ are chosen by the researcher.
When the covariate vector includes lagged outcomes (e.g., $x_{it} = y_{i,t-1}$), the forecast in (ref) requires these lagged values to be observed. For $h=1$, this poses no problem since $y_{i\tau}$ is observed. For $h \geq 2$, however, the required intermediate counterfactuals $y_{i,\tau+1}(0), \ldots, y_{i,\tau+h-1}(0)$ are not observed under treatment, so iterated forecasting is not directly feasible. One solution is to adopt a local-projection approach that directly models the $h$-step-ahead relationship between $y_{i,\tau+h}(0)$ and $x_{i\tau}$, analogous to Jorda2005. Alternatively, one may use the Unobserved-Components FAT from Section (ref), which does not require dynamic covariates and remains feasible at all horizons.
If the model for the counterfactuals is an AR(p) with heterogeneous parameters, the time series literature (e.g., FullerHasza,Dufour1984) has derived conditions under which forecasts from an individual AR(p) model are unbiased. The assumptions are stationarity of the initial condition and symmetry of the error term.
A second example is strictly exogenous covariates with heterogeneous coefficients. For example, suppose that the $h=1$ period forecast of $y_{i\tau+1}\left(0\right)$ is given by
The forecast (ref) is unbiased provided that a Vandermonde matrix which includes functions of $\left(\tau-R_{i}+s\right)^{j},j=0,\dots,q_{i},s=1,2,\dots,R_{i}$ and the covariates $x_{i\tau-R_{i}+s},s=1,2,\dots,R_{i},$ is invertible. This condition constrains how the covariates can change over time.
In this section, we study the finite-sample performance of our estimators. Although we focus on a universal-treatment setting, the insights from the simulations carry over unchanged to staggered adoption, as linearity of the forecasting operator ensures unbiased individual forecasts aggregate to unbiased cohort forecasts.
In Section (ref), we compare the Unobserved-Components FAT (UC FAT) to the Parametric-Model FAT (PM FAT) across forecast horizons, sample sizes, autoregressive persistence, and alternative initial-condition regimes (stationary vs.\ nonstationary). PM FAT can inherit small-sample biases from first-stage estimation of the AR parameter, whereas UC FAT avoids these by construction; the contrast is most pronounced under high persistence and small $n$. When initial conditions are stationary, finite-sample biases are generally small across forecast horizons.
Section (ref) focuses on UC FAT and evaluates how tuning parameter choices - polynomial order $q$ and pre-treatment window $R$ - trade off bias and variance under DGPs satisfying Assumptions (ref) and (ref). In the small-$T$ setting we consider, the simulations support the recommendation to use the lowest polynomial order that accommodates any deterministic trend in the outcome and, conditional on that choice, select the largest feasible pre-treatment window. When the trend order is underfit, bias emerges; in that case, shortening the pre-treatment window can mitigate misspecification. Once the polynomial order is rich enough, bias is largely insensitive to $(q,R)$, and precision improves with lower order and longer windows. These patterns are robust across DGPs with heterogeneous autoregressive parameters and individual-specific trends, and they provide clear guidance for empirical implementation.
We consider the following DGP for the counterfactual outcome: \[ y_{it}(0) \;=\; y^{(1)}_{it}(0) + y^{(3)}_{it}(0), \qquad t=1,\dots,T, \] where the autoregressive component has a homogeneous AR parameter \[ y^{(1)}_{it}(0) = \mu_i + \rho\, y^{(1)}_{i,t-1}(0) + u_{it}, \quad \mu_i \sim \mathcal U[-1,1],\quad u_{it}\sim\mathcal N(0,1),\quad \rho\in\{0.2,0.9\}, \] and the deterministic trend is homogeneous linear \[ y^{(3)}_{it}(0) = \delta_i t,\qquad \delta_i=1. \] We consider two regimes for $y^{(1)}_{i0}(0)$:
We work with a balanced panel, $T=8$ and $\tau=5$ (five pre-treatment periods), forecast horizons $h\in\{1,2,3\}$, and sample sizes $n\in\{50,1000\}$. Treatment effects at $\tau+h$ are zero by construction.
\paragraph{Estimators.} For $b_k(t)=t^k$, the UC forecast for unit $i$ at $\tau+h$ is \[ \widehat y^{\text{UC}}_{i,\tau+h}(0)=\sum_{k=0}^{q} \widehat c_{ik}(\tau+h)^k,\quad \widehat c_i=\arg\min_{c\in\mathbb R^{q+1}}\sum_{t=1}^{\tau}\Big(y_{it}-\sum_{k=0}^q c_k t^k\Big)^2, \] and our UC FAT estimator is
We compute (ref) using the fatEstimator package in R, with $R=q+1$ and $q\in\{0,1,2,3\}$.
The PM forecast first estimates $\rho$ by Anderson–Hsiao using $y_{i,t-2}$ as an instrument for $\Delta y_{i,t-1}$ in $\Delta y_{it}=\rho\,\Delta y_{i,t-1}+\Delta u_{it}$, then fits a polynomial to the residuals $y_{it}-\widehat\rho\,y_{i,t-1}$, so that: \[ \widehat y^{\text{PM}}_{i,\tau+h}(0)=\widehat\rho\,y_{i,\tau}+\sum_{k=0}^{q}\widehat c_{ik}(\tau+h)^k,\quad \widehat c_i=\arg\min_{c}\sum_{t=1}^{\tau}\Big(y_{it}-\widehat\rho\,y_{i,t-1}-\sum_{k=0}^{q} c_k t^k\Big)^2, \] with the PM FAT estimator given by:
We estimate $\rho$ using the ivreg package in R and then use the fatEstimator package on the residuals $y_{it}-\widehat\rho\,y_{i,t-1}$. The specification fits polynomials of degree $q\in\{0,1,2,3\}$ to $R=q+1$ pre-treatment periods.
\paragraph{Simulation scenarios.} We consider four scenarios, varying the cross-sectional size and the autoregressive parameter: \[ (n,\rho) \in \{(50,0.2),\ (50,0.9),\ (1000,0.2),\ (1000,0.9)\}. \]
We report bias and RMSE for UC and PM FAT across polynomial orders $q\in\{0,1,2,3\}$, forecast horizons $h\in\{1,2,3\}$, AR parameter magnitude $\rho\in\{0.2,0.9\}$, and sample size $n\in\{50,1000\}$, based on 500 Monte Carlo replications.
Tables (ref) and (ref) report bias and RMSE for the UC FAT for a DGP with initial condition drawn from the nonstationary distribution (Table (ref)) and from the stationary distribution (Table (ref)). With a nonstationary initial condition (Table (ref)), the $q=0$ specification - which omits the linear trend - exhibits large upward bias that grows roughly linearly with the forecast horizon $h$. Including a linear term ($q=1$) effectively removes this bias: the remaining deviations from zero are small across all horizons and values of $\rho$. Higher-order specifications ($q=2,3$) do not reduce asymptotic bias in this DGP (the true trend is linear), but they affect small-sample behavior: they can attenuate some initialization effects at the cost of substantially larger RMSE because of polynomial extrapolation. In all cases, RMSE increases with the forecast horizon.
With a stationary initial condition (Table (ref)), the $q=0$ specification still omits the linear time trend and exhibits a systematic upward bias. Once a linear term is included ($q=1$), the forecast specification matches the deterministic component of the DGP, and bias is essentially zero across horizons and values of $\rho$; the small nonzero entries in the table reflect finite-sample variability. Higher-order specifications ($q=2,3$) remain asymptotically unbiased but exhibit larger RMSE due to higher-order polynomial extrapolation. As expected, RMSE increases with $h$, while bias patterns are stable across $\rho$.
Tables (ref) and (ref) report bias and RMSE of PM FAT for a DGP with initial condition drawn from the nonstationary distribution (Table (ref)) and from the stationary distribution (Table (ref)). With a nonstationary initial condition (Table (ref)), the $q=0$ specification again omits the linear trend and generates a large upward bias that increases with the forecast horizon $h$. Introducing a linear term ($q=1$) removes this trend-misspecification bias at the population level, so that any remaining bias is driven by first-stage estimation of the AR parameter $\rho$. In finite samples, PM FAT inherits a small Nickell-type bias through $\hat{\rho}$, which propagates into the $h$-step forecast and yields modest biases even when $q=1$. Higher-order specifications ($q=2,3$) do not further reduce this asymptotic bias but amplify RMSE because polynomial extrapolation increases forecast variance, especially at longer horizons. As with UC FAT, RMSE rises with $h$, while the qualitative bias pattern is similar across values of $\rho$.
With a stationary initial condition (Table (ref)), the $q=0$ model still omits the linear trend and exhibits a systematic upward bias, though somewhat attenuated relative to the nonstationary case. Once a linear term is included ($q=1$), the forecasting specification matches the deterministic component of the DGP, and PM FAT is asymptotically unbiased; the small residual biases reflect the Nickell-type distortion from estimating $\rho$ in the first step. As before, higher-order specifications ($q=2,3$) remain asymptotically unbiased but exhibit larger RMSE due to variance inflation from higher-order polynomials. Compared to the nonstationary case, the stationary initial condition reduces the magnitude of the finite-sample biases, while RMSE continues to increase with $h$ and bias patterns remain stable across $\rho$.
In this section, we compare the finite-sample performance of the Unobserved-Components FAT estimator across different tuning parameters: the polynomial order, $q$, and the estimation window, $R$, for different specifications of the DGP. All specifications satisfy Assumption (ref), with the counterfactual process specified as the sum of up to three different components.
Let $I_j, j=1,2,3$ be an indicator that equals to one whenever the associated components is present in the specification of the counterfactual process. For each $i=1,\dots,n$:
Note that the initial observation ${y}_{i0}^{(1)}$ is drawn from the stationary distribution and the time trend is linear and homogeneous, so it can be interpreted as a common shock.
Table (ref) shows results for the bias and standard error of $\widehat{\rm FAT}^{\rm UC}_1$ across different tuning parameters. The table shows that, when the DGP is mean stationary (first panel) or when it is the sum of a mean stationary and a random walk (second panel), the bias does not vary much across different values of the tuning parameters, while the standard errors are smaller for smaller $q$ and larger $R$. When the DGP contains a time trend component, we observe bias when the polynomial-order $q$ is less than the true order of the time trend, as the theory predicts. In this case, however, a smaller estimation window $R$ gives a smaller bias. When $q\geq 1$, the performance of the estimator in terms of bias is again robust to the choice of tuning parameters, with smaller standard errors for smaller $q$ and larger $R$.\footnote{We observe the same behavior with a sample size of $N=30$, see Table (ref) in the Online Appendix.}
\spacingset{1.9}
We consider an empirical exercise showing that FAT, despite not requiring an untreated control group, yields conclusions consistent with established methods that rely on a control group. This positions FAT as a valuable complementary robustness check when a control group is available, rather than a replacement for existing methods such as TWFE or DiD.\footnote{Additional replications are included in the Online Appendix (ref) and (ref).}
We consider the application in GoodmanBacon on the effect of no-fault laws on female suicide. The data are on U.S. states that adopted no-fault divorce laws from 1969 to 1985. The outcome is the age-adjusted average suicide mortality rate per million women (ASMR) in the state. We work with a balanced panel of $n=36$ treated states and 10 time periods, 5 of which are pre-treatment.
We assume the same tuning parameters $q$ and $R$ for FAT across states. To help select the basis functions and the tuning parameters, Figure (ref) plots the outcome variable averaged across states, as a function of time to adoption. The visible trending behaviour (plausibly reflecting a deterministic component) in the pre-treatment data suggests using a polynomial basis function and choosing orders $q>0$. There is no strong indication of structural instability in the pre-treatment data, so we use all available data for estimation, that is, we set $R=5$. We compute the forecast $\widehat{y}^{UC}_{i\tau+h}$ in (ref) for a number of forecast horizons $h$ and orders $q$. The Unobserved-Components FAT estimates at $\tau+h$ are shown in the top panel of Figure (ref) as a function of $h\in\{1,\dots,5\}$ and for $q\in\{1,2\}$. The variance of the estimator is computed as explained in Section (ref). The results show a statistically insignificant decrease in the suicide rate after adoption of the no-fault divorce laws across different forecast horizons. This replicates the findings in GoodmanBacon, ClarkeTapiaSchythe, but without using a control group. The standard error of our estimator is comparable to that obtained in the above-mentioned studies that employ event-study design approaches and, as expected, our $95\%$ confidence intervals are greater for longer forecast horizons (see, e.g., the results in Figure 2 in ClarkeTapiaSchythe).\footnote{To see how our estimator performs in terms of standard error with a small sample size, as in this application, see Table (ref) in the Online Appendix for simulation results with $N=30$.}
To validate the procedure, in Figure (ref) we report placebo FAT at lag $j$, $j\in\{0,\dots,3\}$, calculated as if the law was adopted $j$ years earlier than its actual adoption date (for each state). The forecast horizon is $h=1$. The figure shows that the placebo FATs for both polynomial orders are close to zero. This offers suggestive evidence that FAT estimates may be interpreted as ATT estimates.
This paper proposes an estimator of the average treatment effects in the absence of an untreated or control group, based on forecasting individual counterfactuals using basis function regressions over a (short) time series of pre-treatment data. Forecast unbiasedness is a key requirement that is satisfied by our approach under a broad class of DGPs that express the individual counterfactuals as the sum of up to three unobserved components: a stationary process, a stochastic trend, and a deterministic trend. The approach is robust, general, and flexible, allowing for unbalanced panels, heterogeneous treatment effects, and staggered treatment timing. Forecasting counterfactuals using a parametric model instead requires stronger assumptions and can perform poorly due to misspecification and estimation bias in small samples (e.g., incidental parameter problem).
\FloatBarrier
\spacingset{1.9}