EconBase
← Back to paper

What Do We Get from Two-Way Fixed Effects Regressions? Implications from Numerical Equivalence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,537 characters · 16 sections · 34 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

What Do We Get from Two-Way Fixed Effects Regressions? Implications from Numerical Equivalence

abstractThis paper develops numerical and causal interpretations of two-way fixed effects (TWFE) regressions in settings with nonbinary, nonstaggered treatments and time-varying covariates. Using the equivalence between TWFE and pooled first-difference regressions, I express the TWFE coefficient as a weighted average of first-difference coefficients across all horizons, clarifying how short- and long-run changes contribute to the estimate. Causal interpretation relies on common-trends assumptions across all horizons and conditioning on covariate changes rather than levels. I propose diagnostic procedures to assess these assumptions across horizons and illustrate them by reexamining TWFE estimates of minimum-wage effects on employment. \global\long

Introduction

Linear regression methods are widely used in empirical economics for their simplicity, but their ability to deliver clear causal insights is debated when treatment effects are heterogeneous angrist2010credibility,heckman2010comparing. Two-way fixed effects (TWFE) regressions, increasingly common in panel data settings, illustrate this tension. They build on the linear regression framework with unit and time fixed effects, drawing conceptual motivation from the canonical two-period difference-in-differences (DID) design. Yet in settings with multiple periods and nonbinary treatments, the link between the data and the TWFE coefficients becomes less transparent. Recent methodological work has proposed alternative estimators, but few simultaneously accommodate nonbinary treatments, time-varying covariates, and complex treatment paths. TWFE regressions therefore remain widely used, highlighting the need to understand their behavior in general settings.

This paper examines TWFE regressions from both numerical and causal perspectives, without assuming that their motivating linear equation fully reflects the true causal relationships. My analysis builds on a simple yet powerful insight: TWFE regressions can be understood through their equivalence to pooled first-difference (FD) regressions. This equivalence holds in panel data with units $i=1,\ldots,N$ and periods $t=1,\ldots,T$, in which TWFE regressions are based on the equation

equation[equation omitted — 108 chars of source]

where $Y_{it}$ is a scalar outcome, $\alpha_{i}$ is a unit-specific effect, $X_{it}$ is a vector of explanatory variables, $C_{i}$ is a vector of time-invariant covariates, $\mu_{t}$ is a period-specific effect, and $\varepsilon_{it}$ is a residual. Consider the corresponding time-series differences of the TWFE equation:

equation[equation omitted — 202 chars of source]

where $\Delta_{k}a_{t}\equiv a_{t+k}-a_{t}$ denotes a $k$-period difference of any time series $\{a_{t}\}^{T}_{t=1}$. Remarkably, a least-squares estimate of the TWFE equation ((ref)) and a pooled least-squares estimate of the FD equation ((ref)) across all $k=1,\ldots,T-1$ yield algebraically identical estimates of the coefficient $\beta$.\footnote{In standard econometric terminology, \textquotedbl first\textquotedbl in FD refers to the order of differencing (differencing once, as opposed to twice in second differences) rather than the time gap. Thus, this paper uses FD to denote $\Delta_{k}$ operations regardless of the value of $k$.} Building on this equivalence, the TWFE coefficient can be decomposed into a weighted average of FD coefficients across all difference lengths $k$. This decomposition reveals how TWFE regressions aggregate short- and long-run changes, providing a diagnostic tool for assessing its identifying assumptions.

This equivalence generalizes the well-known TWFE--FD identity from two-period panels $(T=2$) to multiperiod settings ($T\ge2$). It follows from a basic U-statistics identity (see Section (ref)) and has been applied to bias correction of the within estimator under correctly specified models HanLee2017,han2022bias. However, its implications for understanding TWFE estimators without assuming the linear model's causal validity have received little attention. My contribution lies in applying this perspective to general settings---across binary, discrete, and continuous regressors, and regardless of treatment timing or the inclusion of covariates---including settings where related decomposition results exist goodman2018difference,strezhnev2018semiparametric and settings where they do not.

Building on this insight, I clarify how TWFE coefficients admit causal interpretation within a potential outcome framework when $X_{it}$ consists of a scalar treatment ($D_{it}$) and time-varying covariates ($W_{it}$). A key condition is the conditional common trends assumption: \[ E\left[\Delta_{k}Y_{it}(d)|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d)|\Delta_{k}W_{it},C_{i}\right]\thinspace\thinspace\thinspace\text{for }k=1,\ldots,T-1, \] imposing conditional mean independence of the potential outcome change $\Delta_{k}Y_{it}(d)$, evaluated at treatment level $d$, from the treatment change $\Delta_{k}D_{it}$. Crucially, this assumption builds on $k$-period changes and conditions only on concurrent changes, not full histories as in strict exogeneity, reflecting that TWFE regressions rely on $k$-period changes. This assumption highlights potential challenges in supporting causal interpretation of the TWFE coefficient, including the difficulty of justifying common trends over increasingly long horizons, and the reliance on conditioning on covariate changes rather than predetermined levels, both of which contrast with the canonical two-period DID framework.

The core issue with TWFE lies in its aggregation across all possible difference lengths. I show that individual $k$-period FD regressions admit causal interpretation under the same identifying conditions, which need only hold for the specific $k$ rather than across all possible $k$. Traditional panel data econometrics justifies TWFE's aggregation on efficiency grounds, which rely on three joint conditions: common trends for all difference lengths $k$, static treatment effects, and homogeneous treatment effects. When at least one of these conditions fails, the TWFE aggregation loses its advantage, and systematic variation in FD coefficients across difference lengths $k$ provides a signal of such failures.

In response to concerns about TWFE's reliance on homogeneity, recent work has developed heterogeneity-robust alternatives (e.g., Chaisemartin2020,de2020difference; callaway2020difference,wooldridge2021two,de2022difference,Borusyak2018,callaway2021difference). Yet these methods do not readily extend to applications involving nonbinary, nonstaggered treatments and time-varying covariates. In such cases, practitioners continue to rely on TWFE estimators. Moreover, some of these alternatives still rely on common trends assumptions that span the entire panel duration, inheriting the fundamental aspect of TWFE's identification structure. Diagnosing when these conditions fail remains practically important, both for evaluating TWFE and for assessing alternative estimators that share similar identifying assumptions.

To this end, I develop diagnostic procedures that signal violations of the common trends assumption for increasing $k$. Unlike the familiar pre-trend tests designed for staggered-adoption designs, these diagnostics apply more broadly and highlight the often-overlooked question of how far across horizons one is willing to assume parallel trends. When diagnostics support common trends across all horizons, researchers may proceed with TWFE or---if the setting permits---heterogeneity-robust alternatives that also rely on long-horizon common trends. When diagnostics reveal violations at longer horizons, $k$-period FD regressions with appropriately chosen $k$ provide a more defensible approach, as do heterogeneity-robust alternatives that do not rely on common trends across the entire panel.

This paper contributes to the recent literature on TWFE regressions, which has primarily focused on settings with binary or staggered treatments. A growing number of studies investigate numerical and causal properties of TWFE estimators in such settings, diagnose their issues, and propose alternative estimators. The literature covers binary and staggered treatment cases athey2018design,callaway2020difference,goodman2018difference,wooldridge2021two, binary and nonstaggered cases Chaisemartin2020, event-study settings Borusyak2018,lin2022interpreting,schmidheiny2020event,sun2020estimating, cases with multiple binary treatment variables de2022several, and continuous and staggered treatment settings callaway2021difference.\footnote{wooldridge2021two applies to general settings numerically but focuses on binary, staggered cases for causal interpretation. Chaisemartin2020 cover nonbinary treatments in an appendix; the relationship to the present paper is discussed in Section (ref).} While these existing studies offer valuable insights into specific well-defined settings, this paper provides general results for TWFE regressions that hold across a broader range of specifications, including binary, discrete, or continuous regressors, with or without covariates, and under any treatment paths.

The rest of this paper is organized as follows. Section (ref) investigates the numerical properties of TWFE regressions, and Section (ref) offers causal interpretation of TWFE coefficients. Section (ref) discusses implications of these findings and develops diagnostic procedures for assessing common trends violations across different time horizons. Section (ref) illustrates these insights by examining the TWFE estimates of the minimum wage effect on employment outcomes. Section (ref) concludes. All proofs are in Appendix (ref).

Numerical Properties

This section develops a numerical interpretation of the TWFE estimator, revealing how it aggregates information across time horizons. Unlike existing approaches that decompose the TWFE estimator into between-group comparisons, the approach here isolates the contribution of each time horizon without presuming well-defined treatment and comparison groups. Applicable to arbitrary treatment patterns, this perspective enables both causal interpretation under horizon-specific assumptions (Section (ref)) and diagnostics for assumption violations (Section (ref)).

I focus on a balanced panel for simplicity; Appendix (ref) discusses how the results in the main paper extend to or differ in an unbalanced panel. For any time series $\{a_{t}\}^{T}_{t=1}$, I use $\Delta_{k}a_{t}\equiv a_{t+k}-a_{t}$ to denote a $k$-period difference and $\overline{a}\equiv\frac{1}{T}\sum^{T}_{t=1}a_{t}$ to denote the time-series average.

Equivalence of Least-Square Objectives

In a two-period panel, it is well known that TWFE and FD regressions yield the same coefficient estimates. This equivalence extends to multiperiod panels. The following theorem shows that the TWFE estimator can be obtained from a least-squares problem that pools across all $k$-period first differences. This result holds for both univariate and multivariate regressors, whether they are binary, discrete, or continuous.

thmLet $\widehat{\beta}_{\text{FE}}$ be the coefficient on $X_{it}$ from a least-squares problem: \begin{equation} \underset{\beta,\{\alpha_{i}\}^{N}_{i=1},\{\gamma_{t},\mu_{t}\}^{T}_{t=1}}{\min}\sum^{N}_{i=1}\sum^{T}_{t=1}\left(Y_{it}-\alpha_{i}-X_{it}'\beta-C_{i}'\gamma_{t}-\mu_{t}\right)^{2}. \end{equation} Then $\beta=\widehat{\beta}_{\text{FE}}$ is also solves: \begin{align} \underset{\beta,\{\gamma_{t},\mu_{t}\}^{T}_{t=1}}{\min} & \sum^{N}_{i=1}\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}\left(\Delta_{k}Y_{it}-\Delta_{k}X_{it}'\beta-C_{i}'\Delta_{k}\gamma_{t}-\Delta_{k}\mu_{t}\right)^{2}. \end{align}

The equivalence follows from a well-known U-statistics identity: the sample covariance between $\{a_{t}\}^{T}_{t=1}$ and $\{b_{t}\}^{T}_{t=1}$ equals the average of ($a_{s}-a_{t})(b_{s}-b_{t})$ over all pairs with $s>t$, up to normalization. HanLee2017,han2022bias leverage this identity to study bias correction under a correctly specified linear model, in a setting without period-specific parameters $\{\gamma_{t},\mu_{t}\}^{T}_{t=1}$. The same algebraic fact, however, enables a fundamentally different use: reinterpreting TWFE regressions without assuming a linear causal model, and doing so under arbitrary treatment patterns and covariate structures.

The pooled FD objective reveals that TWFE estimates reflect changes in the variables of interest across all possible time differences (from one period to $T-1$ periods), rather than their levels. This shift in perspective opens a door to interpreting TWFE estimators in more general settings than prior work has addressed. For binary treatment settings, goodman2018difference and strezhnev2018semiparametric show that TWFE estimators can be expressed as averages of two-unit, two-period DID comparisons. While their analyses, tailored to these settings, do not explicitly invoke the U-statistics identity, my approach reveals that their results stem from the same numerical structure. Appendix (ref) elaborates on this connection.

Weighted-Average Relationship

A least-squares estimate of equation ((ref)) using only $k$-period differences is given by:

equation[equation omitted — 221 chars of source]

which yields $\widehat{\beta}_{\text{FE},k}$ as the coefficient on $\Delta_{k}X_{it}$ for each $k=1,\ldots,T-1$.\footnote{While this specification has more parameters than degrees of freedom due to $\{\gamma_{t}\}{}^{T}_{t=1}$ and $\{\mu_{t}\}{}^{T}_{t=1}$, the solution for $\beta$ remains unique and is equivalent to the solution given by a more natural specification:$\underset{\beta,\{\gamma^{*}_{t},\mu^{*}_{t}\}{}^{T-k}_{t=1}}{\min}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\left(\Delta_{k}Y_{it}-\Delta_{k}X_{it}'\beta-C_{i}'\gamma^{*}_{t}-\mu^{*}_{t}\right)^{2}.$} The least-squares problem ((ref)), which produces $\widehat{\beta}_{\text{FE}}$ according to Theorem (ref), pools the objective ((ref)) across all $k=1,\ldots,T-1$. The similarity of the two objectives indicates a tight numerical connection between the TWFE coefficient $\widehat{\beta}_{\text{FE}}$ and the FD coefficients $\left\{ \widehat{\beta}_{\text{FD},k}\right\} ^{T-1}_{k=1}$.

In fact, this connection admits an exact matrix-weighted-average interpretation: the TWFE coefficient can be written as a matrix-weighted average of the FD coefficients, with weight matrices that sum to the identity. This result is purely algebraic and follows from an application of the Frisch--Waugh--Lovell theorem. For completeness, the formal statement and proof are provided in Appendix (ref).

The interpretation of this matrix-weighted representation depends on whether $X_{it}$ is univariate or multivariate. When $X_{it}$ is univariate, the matrix weights reduce to scalars and the TWFE coefficient is simply a convex combination of the $k$-period FD coefficients. When $X_{it}$ is multivariate, the aggregation generally involves matrix-valued weights, which are harder to interpret directly. In practice, however, empirical applications typically focus on the coefficient on a single treatment variable, motivating a decomposition that isolates that coefficient even in multivariate specifications.

I therefore consider a common empirical setting in which $X_{it}$ consists of a scalar treatment variable $D_{it}$ and a vector of time-varying covariates $W_{it}$, and focus on the interpretation of the coefficient on $D_{it}$. Using $D_{it}$ as the first element of $X_{it}$, I write: \[ X_{it}=\left[

array[array omitted — 31 chars of source]

\right],\thinspace\thinspace\thinspace\widehat{\beta}_{FE}=\left[

array[array omitted — 81 chars of source]

\right],\thinspace\thinspace\thinspaceand\,\thinspace\thinspace\widehat{\beta}_{FD,k}=\left[

array[array omitted — 85 chars of source]

\right]. \] I also define

equation[equation omitted — 322 chars of source]

to be a residual from a regression of $Y_{it}$ on $C_{i}$ independently performed for each $t$.

The following theorem establishes that $\widehat{\beta}^{D}_{\text{FE}}$ can be expressed as a weighted average of $\widehat{\beta}^{D}_{\text{FD},k}$ after an adjustment for the influence of time-varying covariates.

thmThe TWFE coefficient on $D_{it}$ is given by: \begin{equation} \widehat{\beta}^{D}_{FE}=\sum^{T-1}_{k=1}\widehat{w}_{k}\widehat{\beta}^{D}_{FD,k}+\frac{\sum^{T-1}_{k=1}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\left(\widehat{\delta}^{W}_{FD,k}-\widehat{\delta}^{W}_{FE}\right)'\left(\Delta_{k}\widetilde{W}_{it}\Delta_{k}\widetilde{W}_{it}'\right)\left(\widehat{\beta}^{W}_{FD,k}-\widehat{\beta}^{W}_{FE}\right)}{\sum^{T-1}_{k=1}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\Delta_{k}\widetilde{D}_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)}, \end{equation} where $\widehat{\delta}^{W}_{\text{FE}}$ and $\widehat{\delta}^{W}_{\text{FD},k}$ are the coefficients from the TWFE and $k$-period FD regressions of $D_{it}$ on $W_{it}$, and the weights are defined as: \begin{equation} \widehat{w}_{k}\equiv\frac{\sum^{N}_{i=1}\sum^{T-k}_{t=1}\Delta_{k}\widetilde{D}_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)}{\sum^{T-1}_{\ell=1}\sum^{N}_{i=1}\sum^{T-\ell}_{t=1}\Delta_{\ell}\widetilde{D}_{it}\left(\Delta_{\ell}\widetilde{D}_{it}-\Delta_{\ell}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)}. \end{equation}

This theorem shows that $\widehat{\beta}^{D}_{\text{FE}}$ can be decomposed into two components. The first is a weighted average of the FD coefficients. The weights, which reflect variation in treatment changes not explained by covariates or time effects, sum to one but may be negative in exceptional cases.\footnote{The denominator in ((ref)) equals the sum of $(\Delta_{\ell}\widetilde{D}_{it}-\Delta_{\ell}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})^{2}$, while the numerator equals the sum of $(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FD},k})(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})$ and is not identical to the sum of $(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})^{2}$ unless $\widehat{\delta}^{W}_{\text{FE}}=\widehat{\delta}^{W}_{\text{FD},k}$ or $\widehat{\delta}^{W}_{\text{FE}}=0$. As such, some of the weights do not remain positive in highly unusual cases where unexplained treatment changes given by the TWFE and $k$-period FD regressions are negatively correlated.} The second is an adjustment term, which arises from discrepancies between the TWFE and FD regressions in how they account for covariate effects. While a TWFE regression applies the same coefficients $(\widehat{\beta}^{W}_{\text{FE}}$ and $\widehat{\delta}^{W}_{\text{FE}}$) to all $k$-period covariate changes, FD regressions allow the coefficients on covariate changes ($\widehat{\beta}^{W}_{\text{FD},k}$ and $\widehat{\delta}^{W}_{\text{FD},k}$) to vary across $k=1,\ldots,T-1$, leading to a deviation from a clean weighted-average interpretation. While the adjustment term complicates this interpretation, it has implications for causal interpretation of the population TWFE coefficient, as discussed in Section (ref). When there are no time-varying covariates $W_{it}$, the adjustment term vanishes and the weights are nonnegative.

This decomposition isolates how different time horizons contribute to the TWFE estimator. Unlike decompositions by between-group comparisons (e.g., goodman2018difference), which reveal what units are being compared, this decomposition reveals over what time spans comparisons are made. Group-based decompositions presume well-defined treatment and comparison groups and are most naturally interpreted when common trends assumptions hold across all time horizons; the present decomposition applies to arbitrary treatment patterns and facilitates assessment of whether these assumptions hold uniformly across different horizons. Systematic variation in $\widehat{\beta}^{D}_{\text{FD},k}$ across different $k$ signals potential violations of key assumptions underlying TWFE's causal interpretation. Section (ref) formalizes this perspective by establishing that each $\widehat{\beta}^{D}_{\text{FD},k}$ admits causal interpretation under different assumption sets, while Section (ref) develops diagnostic tools to assess which assumptions are violated.

Causal Interpretation

TWFE Interpretation

This section clarifies the conditions under which the population TWFE coefficient can be interpreted causally, focusing on the case in which $X_{it}$ consists of a scalar treatment ($D_{it}$) and a vector of time-varying covariates ($W_{it}$). I consider a large--$N$ fixed--$T$ setting, with $\left(Y_{it},D_{it},W_{it}\right)^{T}_{t=1}$ and $C_{i}$ being independent and identically distributed across $i=1,\ldots,N$. The cross-sectional mean is represented by $E\left[Y_{it}\right]$, while the time-series mean is denoted by $\overline{Y}_{i}$. In addition, I define \[ \widetilde{Y}_{it}\equiv Y_{it}-\left(

array[array omitted — 25 chars of source]

\right)'E\left[\left(

array[array omitted — 25 chars of source]

\right)\left(

array[array omitted — 25 chars of source]

\right)'\right]^{-1}E\left[\left(

array[array omitted — 25 chars of source]

\right)Y_{it}\right] \] to be a residual from the population projection of $Y_{it}$ on $(1,C_{i})$. This parallels the sample definition in ((ref)); for simplicity of notation I reuse the same symbol.

The equivalence result in Section (ref) reduces a TWFE regression to a regression of $\Delta_{k}\widetilde{Y}_{it}$ on $\Delta_{k}\widetilde{D}_{it}$ controlling for $\Delta_{k}\widetilde{W}_{it}$. The population TWFE coefficient on $D_{it}$ is given by

equation[equation omitted — 375 chars of source]

where

equation[equation omitted — 298 chars of source]

represents the population version of $\widehat{\delta}^{W}_{\text{FE}}$. With incidental unit-specific effect parameters being eliminated, this structure becomes analogous to a standard cross-sectional regression, making the problem tractable with standard causal inference tools.

The following assumptions characterize the conditions under which $\beta^{D}_{\text{FE}}$ admits causal interpretation. The goal is not to justify these assumptions but to make explicit what the TWFE approach relies on, enabling researchers to evaluate when it is appropriate.

assumptionx(Potential Outcome, Static) For each $t=1,\ldots,T$, $\{Y_{it}(d):d\in(\underline{d},\overline{d})\}$ is a stochastic process that defines a potential outcome associated with each possible treatment level $d\in(\underline{d},\overline{d})$, where $-\infty\le\underline{d}<\overline{d}\le\infty$. The observed outcome is given by $Y_{it}=Y_{it}(D_{it})$.

Assumption (ref) restricts potential outcomes to depend only on current treatment status, ruling out dynamic treatment effects. In staggered adoption settings with binary treatments, TWFE regressions can be analyzed while allowing dynamics callaway2020difference,goodman2018difference. In the general setting considered here with continuous treatments and arbitrary treatment paths, allowing both heterogeneity and dynamics would make potential outcomes intractably high-dimensional. Appendix (ref) considers dynamics under a homogeneous treatment effect restriction, demonstrating that the TWFE coefficient remains difficult to interpret even under this strong assumption.

assumptionx(Conditional Common Trends) There exists $d^{0}\in(\underline{d},\overline{d})$ such that \[ E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right] \] for any $(k,t)$ with $1\le t<t+k\le T$.

Assumption (ref) requires that trends in potential outcomes $\Delta_{k}Y_{it}(d^{0})$ at a baseline treatment level $d^{0}$ be mean independent from treatment changes $\Delta_{k}D_{it}$, conditional on time-invariant covariates $C_{i}$ and the concurrent changes $\Delta_{k}W_{it}$ in time-varying covariates. It requires mean independence from the concurrent treatment changes $\Delta_{k}D_{it}$, rather than from the entire treatment path $(D_{i1},\ldots,D_{iT})$. Therefore, when considered for any individual $k$, the condition is weaker than the strict exogeneity assumption standard in panel data and commonly invoked in DID settings.\footnote{In staggered adoption designs, the standard assumption imposes exogeneity with respect to treatment group (defined by adoption timing), which is by construction equivalent to exogeneity with respect to the entire treatment path.} However, requiring this weaker condition to hold at all difference lengths $k=1,\ldots,T-1$ simultaneously does impose collective restrictions on treatment dynamics, as detailed in the remarks at the end of this section.

Assumption (ref) presumes the existence of a natural baseline treatment level with economic meaning, such as \textquotedbl no treatment\textquotedbl ($d^{0}=0$) in a binary treatment setting. The baseline level $d^{0}$ serves as the counterfactual reference point for causal interpretation; the TWFE coefficient measures treatment effects relative to this baseline.\footnote{One could select any arbitrary $d^{0}$ and assert that the assumption holds. However, this makes the causal interpretation meaningless, since effects would be measured relative to an economically irrelevant reference point. Moreover, the assumption itself becomes difficult to defend on economic grounds.} The baseline level $d^{0}$ need not be experienced in the data as long as it represents an economically meaningful reference point and the parallel trends assumption at $d^{0}$ can be economically justified.\footnote{For instance, a tariff rate of zero in trade policy analysis may provide a natural baseline even if no country implements complete free trade, and common trends in the absence of trade barriers may be plausible on economic grounds.} When no natural baseline exists, such as with minimum wage policy where all levels represent active intervention, it can be problematic to assume common trends hold at one arbitrarily chosen level while potentially failing at all others. Appendix (ref) addresses this issue by invoking the common trends assumption for every treatment level.

assumptionx(Variation) $\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[Var\left(\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right)\right]>0$.

Assumption (ref) requires that $\Delta_{k}D_{it}$ have some variation conditional on $\left(\Delta_{k}W_{it},C_{i}\right)$.

assumptionx(Linearity and Time Homogeneity) For any $(k,t)$ with $1\le t<t+k\le T$, the conditional expectation of $\Delta_{k}D_{it}$ given $\left(\Delta_{k}W_{it},C_{i}\right)$ can be expressed as \[ E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]=\Delta_{k}W_{it}'\delta^{W}+C_{i}'\delta^{C}_{k,t}+\delta^{I}_{k,t}, \] where the coefficient $\delta^{W}$ does not vary across $(k,t)$.

Assumption (ref) imposes two restrictions: linearity of the conditional mean function, and time homogeneity of the coefficient $\delta^{W}$ across ($k$,$t$). The linearity component can be made less restrictive through flexible covariate specifications (angrist1999empirical), while the time homogeneity component imposes a fundamental restriction that cannot be relaxed within the TWFE framework.

Assumption (ref) is closely related to an imperfect weighted-average formula in Theorem (ref). The following lemma clarifies the relationship.

lemUnder Assumption (ref), $\delta^{W}_{\text{FE}}=\delta^{W}=\delta^{W}_{\text{FD},k}$ for all $k=1,\ldots,T-1$, where $\delta^{W}_{\text{FE}}$ is as defined in ((ref)) and $\delta^{W}_{\text{FD},k}$ is a coefficient from a regression of $\Delta_{k}\widetilde{D}_{it}$ on $\Delta_{k}\widetilde{W}_{it}$.

Lemma (ref) suggests that Assumption (ref) is a sufficient condition for $\widehat{\delta}^{W}_{\text{FE}}$ and $\widehat{\delta}^{W}_{\text{FD},k}$ to be equivalent in the population. Therefore, under Assumption (ref) the adjustment term in Theorem (ref) vanishes in the population.

The following theorem characterizes $\beta^{D}_{\text{FE}}$ under the full set of assumptions. Appendix (ref) shows that Assumption (ref) and both linearity and time homogeneity components of Assumption (ref) are necessary: relaxing any introduces bias terms that preclude causal interpretation.

thmUnder Assumptions (ref), (ref), (ref), and (ref), \[ \beta^{D}_{\text{FE}}=\sum^{T}_{t=1}E\left[\tau_{it}\omega_{it}\right], \] where \[ \tau_{it}\equiv\begin{cases} \frac{Y_{it}(D_{it})-Y_{it}(d^{0})}{D_{it}-d^{0}} & D_{it}\ne d^{0}\\ 0 & D_{it}=d^{0} \end{cases} \] is a per-unit effect of a deviation of the treatment $D_{it}$ from the baseline level $d^{0}$. The weights satisfy $\sum^{T}_{t=1}E\left[\omega_{it}\right]=1$ and \[ \omega_{it}\propto\left(D_{it}-d^{0}\right)\left\{ \left(\widetilde{D}_{it}-\overline{\widetilde{D}}_{i}\right)-\left(\widetilde{W}_{it}-\overline{\widetilde{W}}_{i}\right)'\delta^{W}_{\text{FE}}\right\} . \]

Theorem (ref) interprets $\beta^{D}_{\text{FE}}$ as a weighted average of per-unit treatment effects, with weights summing to one but possibly negative. The weight function $\omega_{it}$ interacts a deviation of $D_{it}$ from the baseline level $d^{0}$ with a TWFE residual of $D_{it}$.

However, this weighted-average interpretation faces multiple challenges. Three issues concern the assumptions needed for the weighted-average interpretation, while a fourth issue arises even when this interpretation holds.

First, the common trends assumption (Assumption (ref)) must hold across all difference lengths $k$. This requirement extends far beyond the canonical two-period DID setting, and becomes increasingly difficult to justify as the number of time periods grows. Section (ref) explores practical implications from this requirement in detail.

Second, the role of time-varying covariates in a TWFE regression differs entirely from a two-period DID setting that allows for covariates. When a TWFE regression includes time-varying covariates, a common trends assumption must be made conditional on changes in covariates, not their initial levels, as illustrated by Assumption (ref). By contrast, a common trends assumption is typically made conditional on pretreatment observables in a two-period DID setting heckman1997matching,heckman1998matching,abadie2005semiparametric. As formalized by caetano2022difference, a common trends assumption that uses covariate changes is less theoretically favored than the one that uses predetermined covariate levels. In particular, if the treatment change $\Delta_{k}D_{it}$ influences the covariate change $\Delta_{k}W_{it}$, then variation in $\Delta_{k}D_{it}$ with $\Delta_{k}W_{it}$ being fixed does not correspond to ceteris paribus variation that defines the causal effect of the treatment.\footnote{This type of issue is often called a “bad control” problem, a term popularized by angrist2008mostly. They describe occupation as a “bad control” when examining the effects of a college degree on earnings. Indeed, occupation is not held constant when we typically define the ceteris paribus effect of a college degree on earnings.}

Third, the linearity and time homogeneity assumption (Assumption (ref)) can be restrictive unless all covariates are discrete and time-invariant. This assumption demands that the relationship between treatment changes and covariate changes be not only linear but also constant across all time periods and difference lengths. Despite being essential for causal interpretation, these requirements are typically assumed implicitly rather than explicitly defended, unlike common trends assumptions.

Fourth, even when the weighted-average interpretation holds, the weights $\omega_{it}$ can be negative, as similarly highlighted in a binary treatment case Borusyak2018,Chaisemartin2020,goodman2018difference. Whether this poses a problem depends on the correlation between weights and treatment effects. If some weights are negative and weights are strongly correlated with per-unit treatment effects, then $\beta^{D}_{\text{FE}}$ can even lie outside the support of per-unit treatment effects $\tau_{it}$. Nevertheless, negative weights alone are not necessarily problematic if treatment effects are uncorrelated with weights, as noted by Chaisemartin2020 in their Corollary 2. Appendix (ref) shows that the weight function $\omega_{it}$ contains redundant variation that is orthogonal to treatment effects under a more strict common trends assumption, which may alleviate the concern about negative weights in some cases.

Overall, these issues highlight that causal interpretation of TWFE estimators involves more restrictive assumptions than commonly recognized, particularly regarding the length over which common trends can be justified, the role of time-varying covariates, and the need for linearity and time homogeneity conditions, all of which are rarely scrutinized in practice.

Remarks

Chaisemartin2020 provide a weighted-average interpretation of the TWFE coefficient that parallels Theorem (ref) in their online appendix. They separately consider cases with nonbinary treatment without covariates and binary treatment with time-varying covariates. However, their proof can be extended to allow for nonbinary treatment with both time-invariant and time-varying covariates. In the current setup, their counterpart to Assumption (ref) can be expressed as

equation[equation omitted — 181 chars of source]

They also assume linearity and time homogeneity of $E\left[\Delta_{1}Y_{it}(d^{0})|\Delta_{1}W_{it},C_{i}\right]$. This assumption corresponds to Assumption (ref), which imposes similar conditions on $E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]$ for $k=1,\ldots,T-1$.

Their assumption ((ref)) conditions on entire treatment and covariate histories. Thus, this single-period ($k=1$) assumption implies the corresponding assumption for all difference lengths $k=1,\ldots,T-1$ when combined with their linearity and time homogeneity assumption. By contrast, Assumption (ref) conditions only on concurrent changes $(\Delta_{k}D_{it},\Delta_{k}W_{it})$, and the assumption for $k=1$ does not imply the corresponding ones for longer differences. For example, $\Delta_{2}Y_{it}(d^{0})$ may correlate with $\Delta_{2}D_{it}$ even when $\Delta_{1}$ terms show no correlation, due to feedback between outcome changes in one period and treatment changes in another.

The contribution of Theorem (ref) relative to their result lies in showing that Assumption (ref), which relies solely on concurrent changes rather than full histories, is central to causal interpretation of the TWFE coefficient. The importance of covariate changes rather than levels for TWFE interpretation is well-recognized in canonical two-period settings, stemming from the known equivalence between TWFE and FD. Theorem (ref) shows that concurrent changes remain the fundamental basis for causal interpretation even in multiperiod settings. caetano2024difference also emphasize the role of covariate changes for TWFE interpretation but focus on the two-period case; they condition on full covariate histories in their multiperiod extension. By contrast, Theorem (ref) demonstrates the sufficiency of conditioning only on concurrent changes.

While this emphasis on concurrent changes clarifies the role of covariates, the implications for treatment dynamics warrant closer examination. Exogeneity of an entire treatment series $\left(D_{i1},\ldots,D_{iT}\right)$ as in ((ref)), often called a strict exogeneity condition, is a common assumption in panel data models. This assumption rules out the possibility that a shock to the current outcome influences future treatment status or is correlated with past treatment status. Although constraints that Assumption (ref) imposes on treatment dynamics appear weaker than strict exogeneity when considered for each $k=1,\ldots,T-1$ individually, similar restrictions arise when considered collectively. Specifically, the common trends assumption for periods $t$ to $t+2k$ inherently makes it unlikely for potential outcome changes $\Delta_{k}Y_{it}(d^{0})$ in the first $k$ periods to influence treatment changes $\Delta_{k}D_{i,t+k}$ in the subsequent $k$ periods. It also makes correlations between prior treatment changes $\Delta_{k}D_{it}$ and subsequent potential outcome changes $\Delta_{k}Y_{i,t+k}(d^{0})$ practically implausible. Without excluding such scenarios, correlations between $\Delta_{2k}Y_{it}(d^{0})$ and $\Delta_{2k}D_{it}$ would arise, even in the absence of correlations between $\Delta_{k}Y_{it}(d^{0})$ and $\Delta_{k}D_{it}$ or between $\Delta_{k}Y_{i,t+k}(d^{0})$ and $\Delta_{k}D_{i,t+k}$.

Yet a critical distinction of Assumption (ref) from strict exogeneity is that these restrictions manifest as testable departures from common trends at specific horizons rather than as a failure of a monolithic condition that must be maintained or abandoned entirely. This horizon-specific structure enables the diagnostic approach developed in Section (ref).

FD Interpretation

The causal interpretation of the $k$-period FD coefficient follows directly from the TWFE framework, but invokes assumptions only for the specific difference length $k$:

assumptionx(Conditional Common Trends, $k$-period FD) There exists $d^{0}\in(\underline{d},\overline{d})$ such that the following holds for any $t\in\{1,\ldots,T-k\}$: \[ E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]. \]
assumptionx(Variation, $k$-period FD) $\sum^{T-k}_{t=1}E\left[Var\left(\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right)\right]>0$.
assumptionx(Linearity and Time Homogeneity, $k$-period FD) For $t\in\{1,\ldots,T-k\}$, the conditional expectation of $\Delta_{k}D_{it}$ given $\left(\Delta_{k}W_{it},C_{i}\right)$ can be expressed as \[ E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]=\Delta_{k}W_{it}'\delta^{W}_{k}+C_{i}'\delta^{C}_{k,t}+\delta^{I}_{k,t}, \] where the coefficient $\delta^{W}_{k}$ does not vary across $t$.

These assumptions are identical to Assumptions (ref), (ref), and (ref), except that they need only hold for the specific $k$ rather than across all $k=1,\ldots,T-1$. As noted in the preceding remarks, Assumption (ref) for smaller $k$ does not imply it holds for larger $k$.

The population $k$-period FD coefficient $\beta^{D}_{\text{FD},k}$ is obtained from a regression of $\Delta_{k}Y_{it}$ on $\Delta_{k}D_{it}$ controlling for $\Delta_{k}W_{it}$ and $C_{i}$ interacted with period indicators. Under the stated assumptions, this coefficient has the following causal interpretation.

thmUnder Assumptions (ref), (ref), (ref), and (ref), \[ \beta^{D}_{\text{FD},k}=\sum^{T}_{t=1}E\left[\tau_{it}\omega^{(k)}_{it}\right], \] where $\tau_{it}$ is a per-unit effect defined in Theorem (ref). The weights satisfy $\sum^{T}_{t=1}E\left[\omega^{(k)}_{it}\right]=1$ and \[ \omega^{(k)}_{it}\propto(D_{it}-d^{0})\sum^{T}_{s=1}\mathbbm{1}_{|t-s|=k}\left(\widetilde{D}_{it}-\widetilde{D}_{is}-\left(\widetilde{W}_{it}-\widetilde{W}_{is}\right)'\delta^{W}_{k}\right). \]

Theorem (ref) suggests that the $k$-period FD coefficient admits a weighted-average interpretation of per-unit treatment effects, with weights that depend on the unpredictable component of $k$-period treatment changes.

The FD interpretation inherits several problems from the TWFE framework. The common trends assumption remains difficult to justify for large $k$, though researchers can mitigate this concern by choosing smaller values of $k$. The identification strategy still relies on covariate changes rather than predetermined covariate levels, creating the same theoretical concerns about ceteris paribus variation. Covariate impacts should be homogeneous across time periods $t$, though homogeneity across difference lengths $k$ is no longer required. Finally, the possibility of negative weights persists, complicating the interpretation of the weighted average.

Among these inherited problems, those associated with covariates can be readily addressed within the FD framework. Rather than controlling for covariate changes $\Delta_{k}W_{it}$, FD regressions can be modified to control for the initial covariate levels $W_{it}$ or other predetermined variables. Moreover, the homogeneity restriction on covariate impacts across time periods can be relaxed by allowing period-specific coefficients.

The FD framework thus offers greater flexibility than the TWFE approach in addressing covariate-related concerns while requiring weaker assumptions overall. Given that a TWFE regression is merely a pooled version of FD regressions as shown in Section (ref), this raises the question of whether such aggregation is desirable when an individual FD coefficient may admit causal interpretation under more plausible conditions.

Practical Implications and Diagnostics

Discussion: When Should We Use TWFE?

Section (ref) shows that the TWFE estimator is a weighted average of FD estimators; Section (ref) establishes that both of them admit causal interpretations. However, the FD interpretation invokes assumptions only for the specific difference length $k$, whereas the TWFE interpretation demands these assumptions hold simultaneously across all possible difference lengths. This asymmetry raises a natural question: why aggregate across all available FD estimators to form a TWFE estimator, when each individual FD estimator may provide valid causal identification under weaker conditions?

In traditional panel data econometrics, the answer centers on efficiency gains. Within the textbook fixed effects model, TWFE and all $k$-period FD estimators identify the same population parameter, and the TWFE estimator achieves greater efficiency under additional distributional assumptions about the error term. Indeed, adding a constant effects assumption to the assumptions in Section (ref) yields this equivalence:

assumptionx(Constant Effects) The per-unit treatment effect defined in Theorem (ref) satisfies $\tau_{it}=\tau$ for all $(i,t)$.
corIf Assumptions (ref), (ref), (ref), (ref), and (ref) hold, then $\beta^{D}_{FE}=\tau=\beta^{D}_{\text{FD},k}$ for all $k=1,\ldots,T-1$.

However, this efficiency justification loses its force when the estimators identify different parameters. The possibility of heterogeneous treatment effects alone undermines this foundation, even when all other identifying assumptions hold. Moreover, potential bias from violated common trends assumptions at longer horizons can be too substantial to justify the gains from reduced standard errors. Long differences may be less reliable for causal identification if unobserved heterogeneity in trends grows over time.\footnote{millimet2025mis formalize a related mechanism within the linear model: when unit-specific heterogeneity drifts over time, longer differences introduce greater bias than shorter ones.} In addition, as noted in Section (ref), the common trends assumption essentially rules out feedback from past outcome changes to future treatment changes within its time frame. Over longer horizons, this assumption becomes harder to justify, as the likelihood of such feedback increases.

One might alternatively justify TWFE aggregation by appealing to interest in long-run effects, since TWFE regressions exploit long-run differences as well as short-run differences. Identifying effects over $k$-period horizons inevitably requires assuming common trends over at least $k$ periods, making the assumption a necessary compromise on this ground. However, even $k$-period FD coefficients have little connection to $k$-period treatment effects when dynamic effects are present, as formalized in Appendix (ref) under an assumption of homogeneous treatment effects. The TWFE coefficient would therefore represent a difficult-to-interpret weighted average of these already complex effect measures, even under the strong homogeneous effect assumption.

These considerations suggest that TWFE maintains clear advantages only under three joint conditions: (1) common trends assumptions hold for all difference lengths $k$, (2) treatment effects are static rather than dynamic, and (3) treatment effects are linear and homogeneous across units and time periods. When any of these conditions fails, the case for TWFE becomes weaker. Researchers may then focus on FD regressions with specific $k$, which inherit some limitations of TWFE but allow more targeted identifying assumptions. Alternatively, they may employ heterogeneity-robust estimators tailored to their specific setting (e.g., Chaisemartin2020,de2020difference; callaway2020difference,wooldridge2021two,de2022difference,Borusyak2018,callaway2021difference), though such methods may be more challenging to apply in broader settings with nonbinary, nonstaggered treatments and time-varying covariates, and some of them still rely on common trends assumptions across the entire panel duration.

The relationship between TWFE and FD estimators provides valuable diagnostic information about assumption violations. If TWFE and $k$-period FD estimates differ substantially, this signals violation of at least one of the three justifying conditions. Two of these---violations of common trends assumptions for specific difference lengths and the presence of dynamic effects---can be assessed through common trends diagnostics developed in Section (ref). For the third condition concerning treatment effect homogeneity, the weight diagnostics established in the recent literature (e.g., Chaisemartin2020) are useful for evaluating how severely violations affect the TWFE coefficient, though with important caveats that arise from assuming common trends at only one treatment level, as highlighted by fabre2022robustness in binary treatment settings and by Appendix (ref) in broader settings.

Diagnosing Violations of the Common Trends Assumption

The common trends assumption cannot be directly tested, since potential outcome changes $\Delta_{k}Y_{it}(d^{0})$ are unobserved. In staggered adoption settings, pre-trend tests provide suggestive evidence about the plausibility of common trends. By contrast, under general treatment paths there is no obvious way to scrutinize this assumption. Yet the plausibility of this assumption can be assessed through testable implications involving the relationship among nonconcurrent treatment and outcome changes. This section develops diagnostic procedures that exploit the temporal structure of panel data to flag violations of common trends across different time horizons.

Sufficient Conditions for Common Trends

The key insight underlying these diagnostics is that the common trends assumption becomes implausible when treatment changes are systematically related to nonconcurrent outcome changes. Suppose that for given $k\in\{2,\ldots,T-1\}$, potential outcome changes in the first $\ell$ periods are correlated with treatment changes in the subsequent $k-\ell$ periods, or treatment changes in the first $\ell$ periods are correlated with potential outcome changes in the subsequent $k-\ell$ periods. This would suggest intertemporal outcome-to-treatment feedback or unobserved factors that drive both treatment and outcome dynamics over the horizon of $k$ periods, indicating a potential violation of the common trends assumption.

More formally, for each $\ell\in\{1,\ldots,k-1\}$, the following two conditions jointly comprise a sufficient condition for Assumption (ref):

description$E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]$ for $t\in\{1,T-k\}$. • $E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{k}W_{it},C_{i}\right]$ for $t\in\{1,\ldots,T-k\}$.

Condition I requires that potential outcome changes over periods $t$ to $t+\ell$ be mean independent of future treatment changes over periods $t+\ell$ to $t+k$ as well as concurrent treatment changes, conditional on covariates. Condition II requires that potential outcome changes over periods $t+\ell$ to $t+k$ be mean independent of past treatment changes over periods $t$ to $t+\ell$ as well as concurrent treatment changes, conditional on covariates.

Testable Implications and Diagnostic Regressions

To overcome the challenge that potential outcome changes are unobserved, I consider the following decomposition of observed outcome changes, which follows from Assumption (ref) and the definition of per-unit treatment effects $\tau_{it}$. \[ \Delta_{k}Y_{it}=\Delta_{k}Y_{it}(d^{0})+\left(\frac{\tau_{i,t+k}+\tau_{it}}{2}\right)\Delta_{k}D_{it}+\Delta_{k}\tau_{it}\left(\frac{D_{i,t+k}+D_{it}}{2}-d^{0}\right). \] While heterogeneity of per-unit treatment effects $\tau_{it}$ complicates the analysis in general, combining Assumption (ref) and Condition I yields: \[ E\left[\Delta_{\ell}Y_{it}|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]+\tau\Delta_{\ell}D_{it}. \] Similarly, under Assumption (ref) and Condition II: \[ E\left[\Delta_{k-\ell}Y_{i,t+\ell}|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{k}W_{it},C_{i}\right]+\tau\Delta_{k-\ell}D_{i,t+\ell}. \] These expressions provide the foundation for empirical diagnostics: under the joint hypothesis of Conditions I and II (sufficient for common trends) and constant effects, regressions of outcome changes on concurrent treatment changes and lead/lag treatment changes controlling for covariates ($\Delta_{k}W_{it},C_{i}$) should yield zero coefficients on the lead/lag terms.

The following regression specifications implement the diagnostic procedures:

eqnarray[eqnarray omitted — 378 chars of source]

Under the joint hypothesis of Conditions I and II (sufficient for Assumption (ref)) and constant effects (Assumption (ref)), we should observe $\beta_{1}=0$ for each regression. These diagnostics can be implemented for increasing values of $k$, starting with $k=2$. Since Conditions I and II also demand exogeneity of concurrent treatment changes, violation of common trends for some $k^{*}$ implies violation of Conditions I and II for larger $k>k^{*}$, which in turn signals (though does not imply) violation of Assumption (ref).

These diagnostics resemble distributed-lag regressions but differ in a key respect: these diagnostics exploit treatment changes over nonoverlapping time periods, whereas distributed-lag regressions, being a form of multivariate TWFE regression, pool all possible FDs and therefore incorporate numerous overlapping treatment changes, obscuring the connection to specific common trends assumptions.

Limitations and Interpretation

These diagnostic procedures should not be interpreted as direct tests of the common trends assumption---a limitation inherent to all such diagnostics, including standard pre-trend tests. The link between diagnostic results and the common trends assumption involves multiple inferential steps, each with important caveats.

First, Conditions I and II are sufficient but not necessary for the common trends assumption (Assumption (ref)). Their violation therefore does not in itself imply that the assumption fails. Nevertheless, for the common trends assumption to hold despite violations of Conditions I and II, correlations in different subperiods must coincidentally cancel out when aggregated over the full $k$-period horizon. While such cancellation is theoretically possible, it would demand a precise offsetting of intertemporal feedback, which is difficult to justify on economic grounds.

Second, $\beta_{1}=0$ in the lead and lag diagnostics is not sufficient for Conditions I and II to hold, since the diagnostics are silent about the exogeneity of concurrent treatment changes---a key part of these conditions. This mirrors a well-known limitation of pre-trend tests in staggered adoption designs, where absence of differential pre-trends does not guarantee common trends in post-treatment periods.

Third, the purest interpretation of the diagnostics requires the constant effects assumption (Assumption (ref)). When treatment effects $\tau_{it}$ vary and depend on treatment paths, the coefficients $\beta_{1}$ can be contaminated by this heterogeneity rather than purely reflecting violations of the sufficient conditions. Yet, this limitation does not undermine diagnostic value: if $\beta_{1}\ne0$ is detected, it indicates either a violation of the sufficient conditions for common trends or the presence of treatment effect heterogeneity that complicates causal interpretation. Both scenarios suggest that standard TWFE estimates may be problematic, making the diagnostic informative regardless of which interpretation applies.

These caveats notwithstanding, the diagnostics remain informative for assessing the plausibility of common trends at different horizons, especially since intertemporal outcome--treatment correlation captures economically meaningful threats to identification, such as anticipation, policy responses to past outcomes, or persistent unobserved shocks. They also have the ancillary benefit of detecting dynamic treatment effects, since $\beta_{1}\neq0$ can arise when treatment effects depend on treatment history (Appendix (ref)). Accordingly, the diagnostics help evaluate whether TWFE estimates are likely to be interpretable as causal effects, and whether estimators that rely more heavily on short-run (or long-run) variation may be preferable.

Empirical Illustration

This section illustrates insights from the econometric results by studying the impact of minimum wages on employment using TWFE regressions. The TWFE regression analysis based on state-level panel data in the US has attracted considerable attention for decades. The econometric framework of this paper, applicable to a broad class of multiperiod panel settings, is particularly well-suited for this empirical context. The treatment variable, the state minimum wage, is continuous and changes many times in a given state. The existing insights from the TWFE regression literature, typically addressing binary treatments and staggered adoption designs, do not directly apply to this setting.

To present the estimates, I use a state--year panel of employment outcomes and minimum wage laws in 50 US states and the District of Columbia from 1979 to 2019. As a dependent variable in a TWFE regression, I use the log employment rate of teens (ages 16--19) from the Current Population Survey (CPS).\footnote{These sample and outcome variable choices follow manning2021elusive, who uses CPS data from 1979 to 2019, aggregated into quarterly state-level observations. However, my analysis is based on annual frequency to incorporate the BDS data.} Another dependent variable is the net job creation rate from the Business Dynamics Statistics (BDS), using four major sectors with high shares of workers without college education.\footnote{The four sectors are: Construction, Manufacturing, Retail Trade, and Accommodation and Food Services. Using the data from all sectors does not qualitatively change the estimates presented below, but the estimated effects become smaller in magnitude.} These employment measures, as stock and flow variables, capture different aspects of labor market adjustment. The focus on age groups or sectors more likely affected by the minimum wage is practical, since the majority of workers in the US are not directly affected by the minimum wage. The minimum wage data are from Vaghul2019, and I use the log of the annual average of the state minimum wage as the treatment variable in each regression. I use census region indicators as covariates.

I begin by illustrating the weighted-average characterization suggested in Section (ref). Figure (ref) presents the FD coefficients $\widehat{\beta}_{\text{FD},k}$ for $k=1,\ldots,40$ as dots accompanied by standard error bars, the TWFE coefficient $\widehat{\beta}_{\text{FE}}$ as a horizontal dashed line, and the weights given by Theorem (ref) as bars. Since all covariates are time-invariant, the adjustment term in Theorem (ref) can be ignored. While the inclusion of time-varying covariates would result in a deviation from the exact weighted-average interpretation, Appendix (ref) finds that the deviation is minimal in this setting even after including time-varying covariates such as log per-capita income and log teen population.

figure[figure omitted — 752 chars of source]

In Panel (a) of Figure (ref), the TWFE coefficient is $\widehat{\beta}_{\text{FE}}=-0.266$, whereas the FD coefficient is $\widehat{\beta}_{\text{FD},1}=-0.001$ with a one-year gap and becomes larger in magnitude as the gap increases. While the difference among “short” FD, “long” FD, and TWFE estimates itself has long been recognized in the minimum wage effects literature neumark1992employment, an explicit numerical relationship among these estimates has not been previously established. The theoretical results in Section (ref) demonstrate that a TWFE regression aggregates these heterogeneous FD coefficients into a single coefficient. The weighted average of all 40 FD coefficients is indeed $-0.266$. Panel (b) of Figure (ref) presents analogous plots for a regression of the net job creation rate. Even though the TWFE coefficient is $\widehat{\beta}_{\text{FE}}=-2.65$, one-year FD coefficient $\widehat{\beta}_{\text{FD},1}=-5.74$ is larger in magnitude and some long-run FD coefficients have different signs. The weighted average of the FD coefficients is again confirmed to be identical to the TWFE coefficient.

The weights depicted in Figure (ref) are identical across the two panels. This consistency arises because the weights depend only on the explanatory variables, not on the dependent variables. As observed in ((ref)), these weights are associated with the unexplained variation of treatment changes. Consequently, the weights exhibit a distinct pattern: they are small for both short- and long-run FD coefficients, albeit for different reasons. For the short-run FD coefficients, the weights are small because minimum wage adjustments tend to be incremental over short periods. By contrast, for the long-run FD coefficients, the weights are small due to the limited number of observations of long-run changes.

The substantial differences among FD coefficients documented in Figure (ref) raise fundamental questions about causal interpretation. Traditional panel data econometrics justifies TWFE's aggregation of these estimates on efficiency grounds, but this efficiency justification is valid only when all FD estimators identify the same population parameter. The patterns observed here suggest that different FD estimators likely identify distinct parameters rather than noisy estimates of a common effect. Indeed, the hypothesis that these coefficients are all identical can be rejected even at a significance level of 0.01%. More concerning, some of these estimates may not identify causal parameters at all if the common trends assumption is violated. With a 41-year panel spanning 1979--2019, it seems unlikely that this assumption holds uniformly across all $k=1,...,40$. The procedures developed in Section (ref) provide a systematic way to evaluate these concerns across different time horizons.

Table (ref) presents regressions based on equations ((ref)) and ((ref)). These diagnostics test whether past and future treatment changes have zero coefficients when regressing outcome changes on them, controlling for concurrent treatment changes and covariates. For practical presentation, I focus on even values of $k\le10$ with lead/lag length $\ell=k/2$.

table[table omitted — 3,017 chars of source]

For teen employment, the hypothesis that both past and future treatment change coefficients equal zero is rejected at the 5% significance level for $k=$ 2, 8, and 10. For $k=2$, the future change coefficient (0.101) and the past treatment change coefficient (--0.128) exceed the concurrent change coefficient in magnitude, suggesting short-term feedback mechanisms or dynamic treatment effects. For $k=$ 8 and 10, both past and future minimum wage changes have negative coefficients, which raises the possibility of unobserved long-term factors that simultaneously drive minimum wage policies and teen employment trends over extended periods.

The job creation rate estimates also exhibit patterns that are hard to reconcile with the common trends assumption. The hypothesis of zero coefficients is rejected at the 5% level for $k=$ 4, 6, 8, and 10. Most notably, the future treatment change coefficients are consistently positive and significant across these horizons, indicating that increases in net job creation systematically precede minimum wage increases. This pattern may reflect intertemporal outcome-to-treatment feedback, where policymakers respond to improving labor market conditions by raising minimum wages.

Overall, the diagnostics indicate that the common trends assumption is most credible only at short horizons. This suggests that estimators relying exclusively on short-run changes may be more appropriate than the TWFE estimator for identifying causal effects in this setting. Yet, because employment is a stock variable that may adjust only gradually, an exclusive focus on short-run variation may never recover the economically relevant long-run effect, as emphasized in the literature (e.g., meer2016effects). This tension underscores the difficulty of designing empirical strategies that are both credible and substantively informative, and highlights the value of diagnostics in making such trade-offs explicit.

Conclusion

Building on a long-overlooked numerical equivalence between TWFE and pooled FD regressions, this paper provides insights into how TWFE regressions translate data into the coefficient of interest. A key observation is that TWFE regressions reflect associations among changes in the variables of interest, pooling all possible short- and long-run comparisons. As a natural consequence of this property, causal interpretation of the TWFE coefficient relies on the common trends assumption holding across all time horizons. To address this challenge, I develop diagnostic procedures that exploit testable implications to help researchers assess whether the common trends assumptions are plausible across different time horizons.

This paper focuses on diagnosing problems of TWFE regressions rather than proposing solutions. While recent literature has developed various alternative estimators in specific contexts such as binary treatments or staggered adoption designs, extending robust causal inference methods to broader settings with general treatment paths and time-varying covariates remains an important challenge for future research.