EconBase
← Back to paper

What Do We Get from Two-Way Fixed Effects Regressions? Implications from Numerical Equivalence

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

73,524 characters

What Do We Get from Two-Way Fixed Effects Regressions? Implications from Numerical Equivalence


\title{What Do We Get from Two-Way Fixed Effects Regressions? Implications
from Numerical Equivalence}
\author{Shoya Ishimaru\footnote{Hitotsubashi University, Department of Economics (email: [email removed]). \newline I thank Mark Colas, Jonathan Meer, Nobuhiko Nakazawa, Masayuki Sawada, Christopher Taber, Kensuke Teshima, Takahide Yanagi, and many seminar participants for helpful comments and suggestions. Support from the Grant-in-Aid for Early-Career Scientists (No. 21K13310) from the Japan Society for Promotion of Science is gratefully acknowledged. All errors are mine.}}
\maketitle
\begin{abstract}
This paper develops numerical and causal interpretations of two-way
fixed effects (TWFE) regressions in settings with nonbinary, nonstaggered
treatments and time-varying covariates. Using the equivalence between
TWFE and pooled first-difference regressions, I express the TWFE coefficient
as a weighted average of first-difference coefficients across all
horizons, clarifying how short- and long-run changes contribute to
the estimate. Causal interpretation relies on common-trends assumptions
across all horizons and conditioning on covariate changes rather than
levels. I propose diagnostic procedures to assess these assumptions
across horizons and illustrate them by reexamining TWFE estimates
of minimum-wage effects on employment.\vfill
\global\long\end{abstract}

\section{Introduction}

Linear regression methods are widely used in empirical economics for
their simplicity, but their ability to deliver clear causal insights
is debated when treatment effects are heterogeneous \citep{angrist2010credibility,heckman2010comparing}.
Two-way fixed effects (TWFE) regressions, increasingly common in panel
data settings, illustrate this tension. They build on the linear regression
framework with unit and time fixed effects, drawing conceptual motivation
from the canonical two-period difference-in-differences (DID) design.
Yet in settings with multiple periods and nonbinary treatments, the
link between the data and the TWFE coefficients becomes less transparent.
Recent methodological work has proposed alternative estimators, but
few simultaneously accommodate nonbinary treatments, time-varying
covariates, and complex treatment paths. TWFE regressions therefore
remain widely used, highlighting the need to understand their behavior
in general settings.

This paper examines TWFE regressions from both numerical and causal
perspectives, without assuming that their motivating linear equation
fully reflects the true causal relationships. My analysis builds on
a simple yet powerful insight: TWFE regressions can be understood
through their equivalence to pooled first-difference (FD) regressions.
This equivalence holds in panel data with units $i=1,\ldots,N$ and
periods $t=1,\ldots,T$, in which TWFE regressions are based on the
equation
\begin{equation}
Y_{it}=\alpha_{i}+X_{it}'\beta+C_{i}'\gamma_{t}+\mu_{t}+\varepsilon_{it},\label{eq:FE_model}
\end{equation}
where $Y_{it}$ is a scalar outcome, $\alpha_{i}$ is a unit-specific
effect, $X_{it}$ is a vector of explanatory variables, $C_{i}$ is
a vector of time-invariant covariates, $\mu_{t}$ is a period-specific
effect, and $\varepsilon_{it}$ is a residual. Consider the corresponding
time-series differences of the TWFE equation:
\begin{equation}
\Delta_{k}Y_{it}=\Delta_{k}X_{it}'\beta+C_{i}'\Delta_{k}\gamma_{t}+\Delta_{k}\mu_{t}+\Delta_{k}\varepsilon_{it}\thinspace\thinspace\thinspace\text{for }k=1,\ldots,T-1,\label{eq:FD_model}
\end{equation}
where $\Delta_{k}a_{t}\equiv a_{t+k}-a_{t}$ denotes a $k$-period
difference of any time series $\{a_{t}\}^{T}_{t=1}$. Remarkably,
a least-squares estimate of the TWFE equation (\ref{eq:FE_model})
and a pooled least-squares estimate of the FD equation (\ref{eq:FD_model})
across all $k=1,\ldots,T-1$ yield algebraically identical estimates
of the coefficient $\beta$.\footnote{In standard econometric terminology, \textquotedbl first\textquotedbl{}
in FD refers to the order of differencing (differencing once, as opposed
to twice in second differences) rather than the time gap. Thus, this
paper uses FD to denote $\Delta_{k}$ operations regardless of the
value of $k$.} Building on this equivalence, the TWFE coefficient can be decomposed
into a weighted average of FD coefficients across all difference lengths
$k$. This decomposition reveals how TWFE regressions aggregate short-
and long-run changes, providing a diagnostic tool for assessing its
identifying assumptions.

This equivalence generalizes the well-known TWFE--FD identity from
two-period panels $(T=2$) to multiperiod settings ($T\ge2$). It
follows from a basic U-statistics identity (see Section \ref{sec:Numerical-Equivalence})
and has been applied to bias correction of the within estimator under
correctly specified models \citep{HanLee2017,han2022bias}. However,
its implications for understanding TWFE estimators without assuming
the linear model's causal validity have received little attention.
My contribution lies in applying this perspective to general settings---across
binary, discrete, and continuous regressors, and regardless of treatment
timing or the inclusion of covariates---including settings where
related decomposition results exist \citep{goodman2018difference,strezhnev2018semiparametric}
and settings where they do not.

Building on this insight, I clarify how TWFE coefficients admit causal
interpretation within a potential outcome framework when $X_{it}$
consists of a scalar treatment ($D_{it}$) and time-varying covariates
($W_{it}$). A key condition is the conditional common trends assumption:
\[
E\left[\Delta_{k}Y_{it}(d)|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d)|\Delta_{k}W_{it},C_{i}\right]\thinspace\thinspace\thinspace\text{for }k=1,\ldots,T-1,
\]
imposing conditional mean independence of the potential outcome change
$\Delta_{k}Y_{it}(d)$, evaluated at treatment level $d$, from the
treatment change $\Delta_{k}D_{it}$. Crucially, this assumption builds
on $k$-period changes and conditions only on concurrent changes,
not full histories as in strict exogeneity, reflecting that TWFE regressions
rely on $k$-period changes. This assumption highlights potential
challenges in supporting causal interpretation of the TWFE coefficient,
including the difficulty of justifying common trends over increasingly
long horizons, and the reliance on conditioning on covariate changes
rather than predetermined levels, both of which contrast with the
canonical two-period DID framework.

The core issue with TWFE lies in its aggregation across all possible
difference lengths. I show that individual $k$-period FD regressions
admit causal interpretation under the same identifying conditions,
which need only hold for the specific $k$ rather than across all
possible $k$. Traditional panel data econometrics justifies TWFE's
aggregation on efficiency grounds, which rely on three joint conditions:
common trends for all difference lengths $k$, static treatment effects,
and homogeneous treatment effects. When at least one of these conditions
fails, the TWFE aggregation loses its advantage, and systematic variation
in FD coefficients across difference lengths $k$ provides a signal
of such failures.

In response to concerns about TWFE's reliance on homogeneity, recent
work has developed heterogeneity-robust alternatives (e.g., \citealt{Chaisemartin2020,de2020difference};
\citealt{callaway2020difference,wooldridge2021two,de2022difference,Borusyak2018,callaway2021difference}).
Yet these methods do not readily extend to applications involving
nonbinary, nonstaggered treatments and time-varying covariates. In
such cases, practitioners continue to rely on TWFE estimators. Moreover,
some of these alternatives still rely on common trends assumptions
that span the entire panel duration, inheriting the fundamental aspect
of TWFE's identification structure. Diagnosing when these conditions
fail remains practically important, both for evaluating TWFE and for
assessing alternative estimators that share similar identifying assumptions.

To this end, I develop diagnostic procedures that signal violations
of the common trends assumption for increasing $k$. Unlike the familiar
pre-trend tests designed for staggered-adoption designs, these diagnostics
apply more broadly and highlight the often-overlooked question of
how far across horizons one is willing to assume parallel trends.
When diagnostics support common trends across all horizons, researchers
may proceed with TWFE or---if the setting permits---heterogeneity-robust
alternatives that also rely on long-horizon common trends. When diagnostics
reveal violations at longer horizons, $k$-period FD regressions with
appropriately chosen $k$ provide a more defensible approach, as do
heterogeneity-robust alternatives that do not rely on common trends
across the entire panel.

This paper contributes to the recent literature on TWFE regressions,
which has primarily focused on settings with binary or staggered treatments.
A growing number of studies investigate numerical and causal properties
of TWFE estimators in such settings, diagnose their issues, and propose
alternative estimators. The literature covers binary and staggered
treatment cases \citep[e.g.,][]{athey2018design,callaway2020difference,goodman2018difference,wooldridge2021two},
binary and nonstaggered cases \citep[e.g.,][]{Chaisemartin2020},
event-study settings \citep[e.g.,][]{Borusyak2018,lin2022interpreting,schmidheiny2020event,sun2020estimating},
cases with multiple binary treatment variables \citep[e.g.,][]{de2022several},
and continuous and staggered treatment settings \citep[e.g.,][]{callaway2021difference}.\footnote{\citet{wooldridge2021two} applies to general settings numerically
but focuses on binary, staggered cases for causal interpretation.
\citet{Chaisemartin2020} cover nonbinary treatments in an appendix;
the relationship to the present paper is discussed in Section \ref{subsec:TWFE-Interpretation}.} While these existing studies offer valuable insights into specific
well-defined settings, this paper provides general results for TWFE
regressions that hold across a broader range of specifications, including
binary, discrete, or continuous regressors, with or without covariates,
and under any treatment paths.

The rest of this paper is organized as follows. Section \ref{sec:Numerical-Equivalence}
investigates the numerical properties of TWFE regressions, and Section
\ref{sec:Causal-Interpretation} offers causal interpretation of TWFE
coefficients. Section \ref{sec:Practical-Implications-and} discusses
implications of these findings and develops diagnostic procedures
for assessing common trends violations across different time horizons.
Section \ref{sec:Empirical-Illustration} illustrates these insights
by examining the TWFE estimates of the minimum wage effect on employment
outcomes. Section \ref{sec:Conclusion} concludes. All proofs are
in Appendix \ref{sec:Proofs}.

\section{Numerical Properties\label{sec:Numerical-Equivalence}}

This section develops a numerical interpretation of the TWFE estimator,
revealing how it aggregates information across time horizons. Unlike
existing approaches that decompose the TWFE estimator into between-group
comparisons, the approach here isolates the contribution of each time
horizon without presuming well-defined treatment and comparison groups.
Applicable to arbitrary treatment patterns, this perspective enables
both causal interpretation under horizon-specific assumptions (Section
\ref{sec:Causal-Interpretation}) and diagnostics for assumption violations
(Section \ref{sec:Practical-Implications-and}).

I focus on a balanced panel for simplicity; Appendix \ref{Asec:Unbalanced}
discusses how the results in the main paper extend to or differ in
an unbalanced panel. For any time series $\{a_{t}\}^{T}_{t=1}$, I
use $\Delta_{k}a_{t}\equiv a_{t+k}-a_{t}$ to denote a $k$-period
difference and $\overline{a}\equiv\frac{1}{T}\sum^{T}_{t=1}a_{t}$
to denote the time-series average.

\subsection{Equivalence of Least-Square Objectives}

In a two-period panel, it is well known that TWFE and FD regressions
yield the same coefficient estimates. This equivalence extends to
multiperiod panels. The following theorem shows that the TWFE estimator
can be obtained from a least-squares problem that pools across all
$k$-period first differences. This result holds for both univariate
and multivariate regressors, whether they are binary, discrete, or
continuous.
\begin{thm}
\label{thm:FE_FD_DID}Let $\widehat{\beta}_{\text{FE}}$ be the coefficient
on $X_{it}$ from a least-squares problem:
\begin{equation}
\underset{\beta,\{\alpha_{i}\}^{N}_{i=1},\{\gamma_{t},\mu_{t}\}^{T}_{t=1}}{\min}\sum^{N}_{i=1}\sum^{T}_{t=1}\left(Y_{it}-\alpha_{i}-X_{it}'\beta-C_{i}'\gamma_{t}-\mu_{t}\right)^{2}.\label{eq:FE_reg}
\end{equation}
Then $\beta=\widehat{\beta}_{\text{FE}}$ is also solves:

\begin{align}
\underset{\beta,\{\gamma_{t},\mu_{t}\}^{T}_{t=1}}{\min} & \sum^{N}_{i=1}\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}\left(\Delta_{k}Y_{it}-\Delta_{k}X_{it}'\beta-C_{i}'\Delta_{k}\gamma_{t}-\Delta_{k}\mu_{t}\right)^{2}.\label{eq:PFD_reg}
\end{align}
\end{thm}
The equivalence follows from a well-known U-statistics identity: the
sample covariance between $\{a_{t}\}^{T}_{t=1}$ and $\{b_{t}\}^{T}_{t=1}$
equals the average of ($a_{s}-a_{t})(b_{s}-b_{t})$ over all pairs
with $s>t$, up to normalization. \citet{HanLee2017,han2022bias}
leverage this identity to study bias correction under a correctly
specified linear model, in a setting without period-specific parameters
$\{\gamma_{t},\mu_{t}\}^{T}_{t=1}$. The same algebraic fact, however,
enables a fundamentally different use: reinterpreting TWFE regressions
without assuming a linear causal model, and doing so under arbitrary
treatment patterns and covariate structures.

The pooled FD objective reveals that TWFE estimates reflect changes
in the variables of interest across all possible time differences
(from one period to $T-1$ periods), rather than their levels. This
shift in perspective opens a door to interpreting TWFE estimators
in more general settings than prior work has addressed. For binary
treatment settings, \citet{goodman2018difference} and \citet{strezhnev2018semiparametric}
show that TWFE estimators can be expressed as averages of two-unit,
two-period DID comparisons. While their analyses, tailored to these
settings, do not explicitly invoke the U-statistics identity, my approach
reveals that their results stem from the same numerical structure.
Appendix \ref{sec:Comparison-GB} elaborates on this connection.

\subsection{Weighted-Average Relationship\label{subsec:Weighted-Average-Relationship}}

A least-squares estimate of equation (\ref{eq:FD_model}) using only
$k$-period differences is given by:
\begin{equation}
\underset{\beta,\{\gamma_{t},\mu_{t}\}{}^{T}_{t=1}}{\min}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\left(\Delta_{k}Y_{it}-\Delta_{k}X_{it}'\beta-C_{i}'\Delta_{k}\gamma_{t}-\Delta_{k}\mu_{t}\right)^{2},\label{eq:FD_LS}
\end{equation}
which yields $\widehat{\beta}_{\text{FE},k}$ as the coefficient on
$\Delta_{k}X_{it}$ for each $k=1,\ldots,T-1$.\footnote{While this specification has more parameters than degrees of freedom
due to $\{\gamma_{t}\}{}^{T}_{t=1}$ and $\{\mu_{t}\}{}^{T}_{t=1}$,
the solution for $\beta$ remains unique and is equivalent to the
solution given by a more natural specification:$\underset{\beta,\{\gamma^{*}_{t},\mu^{*}_{t}\}{}^{T-k}_{t=1}}{\min}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\left(\Delta_{k}Y_{it}-\Delta_{k}X_{it}'\beta-C_{i}'\gamma^{*}_{t}-\mu^{*}_{t}\right)^{2}.$} The least-squares problem (\ref{eq:PFD_reg}), which produces $\widehat{\beta}_{\text{FE}}$
according to Theorem \ref{thm:FE_FD_DID}, pools the objective (\ref{eq:FD_LS})
across all $k=1,\ldots,T-1$. The similarity of the two objectives
indicates a tight numerical connection between the TWFE coefficient
$\widehat{\beta}_{\text{FE}}$ and the FD coefficients $\left\{ \widehat{\beta}_{\text{FD},k}\right\} ^{T-1}_{k=1}$.

In fact, this connection admits an exact matrix-weighted-average interpretation:
the TWFE coefficient can be written as a matrix-weighted average of
the FD coefficients, with weight matrices that sum to the identity.
This result is purely algebraic and follows from an application of
the Frisch--Waugh--Lovell theorem. For completeness, the formal
statement and proof are provided in Appendix \ref{subsec:FE_not_equal_FD}.

The interpretation of this matrix-weighted representation depends
on whether $X_{it}$ is univariate or multivariate. When $X_{it}$
is univariate, the matrix weights reduce to scalars and the TWFE coefficient
is simply a convex combination of the $k$-period FD coefficients.
When $X_{it}$ is multivariate, the aggregation generally involves
matrix-valued weights, which are harder to interpret directly. In
practice, however, empirical applications typically focus on the coefficient
on a single treatment variable, motivating a decomposition that isolates
that coefficient even in multivariate specifications.

I therefore consider a common empirical setting in which $X_{it}$
consists of a scalar treatment variable $D_{it}$ and a vector of
time-varying covariates $W_{it}$, and focus on the interpretation
of the coefficient on $D_{it}$. Using $D_{it}$ as the first element
of $X_{it}$, I write:
\[
X_{it}=\left[\begin{array}{c}
D_{it}\\
W_{it}
\end{array}\right],\thinspace\thinspace\thinspace\widehat{\beta}_{\text{FE}}=\left[\begin{array}{c}
\widehat{\beta}^{D}_{\text{FE}}\\
\widehat{\beta}^{W}_{\text{FE}}
\end{array}\right],\thinspace\thinspace\thinspace\text{and}\,\thinspace\thinspace\widehat{\beta}_{\text{FD},k}=\left[\begin{array}{c}
\widehat{\beta}^{D}_{\text{FD},k}\\
\widehat{\beta}^{W}_{\text{FD},k}
\end{array}\right].
\]
I also define
\begin{equation}
\widetilde{Y}_{it}\equiv Y_{it}-\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)'\left(\sum^{N}_{i=1}\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)'\right)^{-1}\sum^{N}_{i=1}\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)Y_{it}\label{eq:Resid}
\end{equation}
to be a residual from a regression of $Y_{it}$ on $C_{i}$ independently
performed for each $t$.

The following theorem establishes that $\widehat{\beta}^{D}_{\text{FE}}$
can be expressed as a weighted average of $\widehat{\beta}^{D}_{\text{FD},k}$
after an adjustment for the influence of time-varying covariates.
\begin{thm}
\label{theorem:FE_not_equal_FD}The TWFE coefficient on $D_{it}$
is given by:
\begin{equation}
\widehat{\beta}^{D}_{\text{FE}}=\sum^{T-1}_{k=1}\widehat{w}_{k}\widehat{\beta}^{D}_{\text{FD},k}+\frac{\sum^{T-1}_{k=1}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\left(\widehat{\delta}^{W}_{\text{FD},k}-\widehat{\delta}^{W}_{\text{FE}}\right)'\left(\Delta_{k}\widetilde{W}_{it}\Delta_{k}\widetilde{W}_{it}'\right)\left(\widehat{\beta}^{W}_{\text{FD},k}-\widehat{\beta}^{W}_{\text{FE}}\right)}{\sum^{T-1}_{k=1}\sum^{N}_{i=1}\sum^{T-k}_{t=1}\Delta_{k}\widetilde{D}_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)},\label{eq:cov_bias}
\end{equation}
where $\widehat{\delta}^{W}_{\text{FE}}$ and $\widehat{\delta}^{W}_{\text{FD},k}$
are the coefficients from the TWFE and $k$-period FD regressions
of $D_{it}$ on $W_{it}$, and the weights are defined as:
\begin{equation}
\widehat{w}_{k}\equiv\frac{\sum^{N}_{i=1}\sum^{T-k}_{t=1}\Delta_{k}\widetilde{D}_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)}{\sum^{T-1}_{\ell=1}\sum^{N}_{i=1}\sum^{T-\ell}_{t=1}\Delta_{\ell}\widetilde{D}_{it}\left(\Delta_{\ell}\widetilde{D}_{it}-\Delta_{\ell}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}}\right)}.\label{eq:weight_cov}
\end{equation}
\end{thm}
This theorem shows that $\widehat{\beta}^{D}_{\text{FE}}$ can be
decomposed into two components. The first is a weighted average of
the FD coefficients. The weights, which reflect variation in treatment
changes not explained by covariates or time effects, sum to one but
may be negative in exceptional cases.\footnote{The denominator in (\ref{eq:weight_cov}) equals the sum of $(\Delta_{\ell}\widetilde{D}_{it}-\Delta_{\ell}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})^{2}$,
while the numerator equals the sum of $(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FD},k})(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})$
and is not identical to the sum of $(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\widehat{\delta}^{W}_{\text{FE}})^{2}$
unless $\widehat{\delta}^{W}_{\text{FE}}=\widehat{\delta}^{W}_{\text{FD},k}$
or $\widehat{\delta}^{W}_{\text{FE}}=0$. As such, some of the weights
do not remain positive in highly unusual cases where unexplained treatment
changes given by the TWFE and $k$-period FD regressions are negatively
correlated.} The second is an adjustment term, which arises from discrepancies
between the TWFE and FD regressions in how they account for covariate
effects. While a TWFE regression applies the same coefficients $(\widehat{\beta}^{W}_{\text{FE}}$
and $\widehat{\delta}^{W}_{\text{FE}}$) to all $k$-period covariate
changes, FD regressions allow the coefficients on covariate changes
($\widehat{\beta}^{W}_{\text{FD},k}$ and $\widehat{\delta}^{W}_{\text{FD},k}$)
to vary across $k=1,\ldots,T-1$, leading to a deviation from a clean
weighted-average interpretation. While the adjustment term complicates
this interpretation, it has implications for causal interpretation
of the population TWFE coefficient, as discussed in Section \ref{sec:Causal-Interpretation}.
When there are no time-varying covariates $W_{it}$, the adjustment
term vanishes and the weights are nonnegative.

This decomposition isolates how different time horizons contribute
to the TWFE estimator. Unlike decompositions by between-group comparisons
(e.g., \citealp{goodman2018difference}), which reveal \emph{what units}
are being compared, this decomposition reveals \emph{over what time spans}
comparisons are made. Group-based decompositions presume well-defined
treatment and comparison groups and are most naturally interpreted
when common trends assumptions hold across all time horizons; the
present decomposition applies to arbitrary treatment patterns and
facilitates assessment of whether these assumptions hold uniformly
across different horizons. Systematic variation in $\widehat{\beta}^{D}_{\text{FD},k}$
across different $k$ signals potential violations of key assumptions
underlying TWFE's causal interpretation. Section \ref{sec:Causal-Interpretation}
formalizes this perspective by establishing that each $\widehat{\beta}^{D}_{\text{FD},k}$
admits causal interpretation under different assumption sets, while
Section \ref{sec:Practical-Implications-and} develops diagnostic
tools to assess which assumptions are violated.

\section{Causal Interpretation\label{sec:Causal-Interpretation}}

\subsection{TWFE Interpretation\label{subsec:TWFE-Interpretation}}

This section clarifies the conditions under which the population TWFE
coefficient can be interpreted causally, focusing on the case in which
$X_{it}$ consists of a scalar treatment ($D_{it}$) and a vector
of time-varying covariates ($W_{it}$). I consider a large--$N$
fixed--$T$ setting, with $\left(Y_{it},D_{it},W_{it}\right)^{T}_{t=1}$
and $C_{i}$ being independent and identically distributed across
$i=1,\ldots,N$. The cross-sectional mean is represented by $E\left[Y_{it}\right]$,
while the time-series mean is denoted by $\overline{Y}_{i}$. In addition,
I define
\[
\widetilde{Y}_{it}\equiv Y_{it}-\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)'E\left[\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)'\right]^{-1}E\left[\left(\begin{array}{c}
1\\
C_{i}
\end{array}\right)Y_{it}\right]
\]
to be a residual from the population projection of $Y_{it}$ on $(1,C_{i})$.
This parallels the sample definition in (\ref{eq:Resid}); for simplicity
of notation I reuse the same symbol.

The equivalence result in Section \ref{sec:Numerical-Equivalence}
reduces a TWFE regression to a regression of $\Delta_{k}\widetilde{Y}_{it}$
on $\Delta_{k}\widetilde{D}_{it}$ controlling for $\Delta_{k}\widetilde{W}_{it}$.
The population TWFE coefficient on $D_{it}$ is given by
\begin{equation}
\beta^{D}_{\text{FE}}=\frac{\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[\Delta_{k}Y_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\delta^{W}_{\text{FE}}\right)\right]}{\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[\Delta_{k}D_{it}\left(\Delta_{k}\widetilde{D}_{it}-\Delta_{k}\widetilde{W}_{it}'\delta^{W}_{\text{FE}}\right)\right]},\label{eq:pop_FE}
\end{equation}
where
\begin{equation}
\delta^{W}_{\text{FE}}=\left(\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[\Delta_{k}\widetilde{W}_{it}\Delta_{k}\widetilde{W}_{it}'\right]\right)^{-1}\left(\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[\Delta_{k}\widetilde{W}_{it}\Delta_{k}\widetilde{D}_{it}\right]\right)\label{eq:delta_FE_pop}
\end{equation}
represents the population version of $\widehat{\delta}^{W}_{\text{FE}}$.
With incidental unit-specific effect parameters being eliminated,
this structure becomes analogous to a standard cross-sectional regression,
making the problem tractable with standard causal inference tools.

The following assumptions characterize the conditions under which
$\beta^{D}_{\text{FE}}$ admits causal interpretation. The goal is
not to justify these assumptions but to make explicit what the TWFE
approach relies on, enabling researchers to evaluate when it is appropriate.

\begin{assumptionx}

\label{assu:PO}\textup{(Potential Outcome, Static)} For each $t=1,\ldots,T$,
$\{Y_{it}(d):d\in(\underline{d},\overline{d})\}$ is a stochastic
process that defines a potential outcome associated with each possible
treatment level $d\in(\underline{d},\overline{d})$, where $-\infty\le\underline{d}<\overline{d}\le\infty$.
The observed outcome is given by $Y_{it}=Y_{it}(D_{it})$.

\end{assumptionx}

Assumption \ref{assu:PO} restricts potential outcomes to depend only
on current treatment status, ruling out dynamic treatment effects.
In staggered adoption settings with binary treatments, TWFE regressions
can be analyzed while allowing dynamics \citep[e.g.,][]{callaway2020difference,goodman2018difference}.
In the general setting considered here with continuous treatments
and arbitrary treatment paths, allowing both heterogeneity and dynamics
would make potential outcomes intractably high-dimensional. Appendix
\ref{subsec:Dynamic-Effects} considers dynamics under a homogeneous
treatment effect restriction, demonstrating that the TWFE coefficient
remains difficult to interpret even under this strong assumption.

\begin{assumptionx}

\label{assu:PT}\textup{(Conditional Common Trends)} There exists
$d^{0}\in(\underline{d},\overline{d})$ such that
\[
E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]
\]
 for any $(k,t)$ with $1\le t<t+k\le T$.

\end{assumptionx}

Assumption \ref{assu:PT} requires that trends in potential outcomes
$\Delta_{k}Y_{it}(d^{0})$ at a baseline treatment level $d^{0}$
be mean independent from treatment changes $\Delta_{k}D_{it}$, conditional
on time-invariant covariates $C_{i}$ and the concurrent changes $\Delta_{k}W_{it}$
in time-varying covariates. It requires mean independence from the
concurrent treatment changes $\Delta_{k}D_{it}$, rather than from
the entire treatment path $(D_{i1},\ldots,D_{iT})$. Therefore, when
considered for any individual $k$, the condition is weaker than the
strict exogeneity assumption standard in panel data and commonly invoked
in DID settings.\footnote{In staggered adoption designs, the standard assumption imposes exogeneity
with respect to treatment group (defined by adoption timing), which
is by construction equivalent to exogeneity with respect to the entire
treatment path.} However, requiring this weaker condition to hold at all difference
lengths $k=1,\ldots,T-1$ simultaneously does impose collective restrictions
on treatment dynamics, as detailed in the remarks at the end of this
section.

Assumption \ref{assu:PT} presumes the existence of a natural baseline
treatment level with economic meaning, such as \textquotedbl no treatment\textquotedbl{}
($d^{0}=0$) in a binary treatment setting. The baseline level $d^{0}$
serves as the counterfactual reference point for causal interpretation;
the TWFE coefficient measures treatment effects relative to this baseline.\footnote{One could select any arbitrary $d^{0}$ and assert that the assumption
holds. However, this makes the causal interpretation meaningless,
since effects would be measured relative to an economically irrelevant
reference point. Moreover, the assumption itself becomes difficult
to defend on economic grounds.} The baseline level $d^{0}$ need not be experienced in the data as
long as it represents an economically meaningful reference point and
the parallel trends assumption at $d^{0}$ can be economically justified.\footnote{For instance, a tariff rate of zero in trade policy analysis may provide
a natural baseline even if no country implements complete free trade,
and common trends in the absence of trade barriers may be plausible
on economic grounds.} When no natural baseline exists, such as with minimum wage policy
where all levels represent active intervention, it can be problematic
to assume common trends hold at one arbitrarily chosen level while
potentially failing at all others. Appendix \ref{subsec:Causal-Interpretation-without}
addresses this issue by invoking the common trends assumption for
every treatment level.

\begin{assumptionx}

\label{assu:Variation}\textup{(Variation)} $\sum^{T-1}_{k=1}\sum^{T-k}_{t=1}E\left[Var\left(\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right)\right]>0$.

\end{assumptionx}

Assumption \ref{assu:Variation} requires that $\Delta_{k}D_{it}$
have some variation conditional on $\left(\Delta_{k}W_{it},C_{i}\right)$.

\begin{assumptionx}

\label{assu:Homogeneity}\textup{(Linearity and Time Homogeneity)}
For any $(k,t)$ with $1\le t<t+k\le T$, the conditional expectation
of $\Delta_{k}D_{it}$ given $\left(\Delta_{k}W_{it},C_{i}\right)$
can be expressed as
\[
E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]=\Delta_{k}W_{it}'\delta^{W}+C_{i}'\delta^{C}_{k,t}+\delta^{I}_{k,t},
\]
where the coefficient $\delta^{W}$ does not vary across $(k,t)$.

\end{assumptionx}

Assumption \ref{assu:Homogeneity} imposes two restrictions: linearity
of the conditional mean function, and time homogeneity of the coefficient
$\delta^{W}$ across ($k$,$t$). The linearity component can be made
less restrictive through flexible covariate specifications (\citealp{angrist1999empirical}),
while the time homogeneity component imposes a fundamental restriction
that cannot be relaxed within the TWFE framework.

Assumption \ref{assu:Homogeneity} is closely related to an imperfect
weighted-average formula in Theorem \ref{theorem:FE_not_equal_FD}.
The following lemma clarifies the relationship.
\begin{lem}
\label{lem:asy_equiv}Under Assumption \ref{assu:Homogeneity}, $\delta^{W}_{\text{FE}}=\delta^{W}=\delta^{W}_{\text{FD},k}$
for all $k=1,\ldots,T-1$, where $\delta^{W}_{\text{FE}}$ is as defined
in (\ref{eq:delta_FE_pop}) and $\delta^{W}_{\text{FD},k}$ is a coefficient
from a regression of $\Delta_{k}\widetilde{D}_{it}$ on $\Delta_{k}\widetilde{W}_{it}$.
\end{lem}
Lemma \ref{lem:asy_equiv} suggests that Assumption \ref{assu:Homogeneity}
is a sufficient condition for $\widehat{\delta}^{W}_{\text{FE}}$
and $\widehat{\delta}^{W}_{\text{FD},k}$ to be equivalent in the
population. Therefore, under Assumption \ref{assu:Homogeneity} the
adjustment term in Theorem \ref{theorem:FE_not_equal_FD} vanishes
in the population.

The following theorem characterizes $\beta^{D}_{\text{FE}}$ under
the full set of assumptions. Appendix \ref{subsec:Proof_Causal_TWFE}
shows that Assumption \ref{assu:PT} and both linearity and time homogeneity
components of Assumption \ref{assu:Homogeneity} are necessary: relaxing
any introduces bias terms that preclude causal interpretation.
\begin{thm}
\label{thm:Causal_TWFE}Under Assumptions \ref{assu:PO}, \ref{assu:PT},
\ref{assu:Variation}, and \ref{assu:Homogeneity},
\[
\beta^{D}_{\text{FE}}=\sum^{T}_{t=1}E\left[\tau_{it}\omega_{it}\right],
\]
where
\[
\tau_{it}\equiv\begin{cases}
\frac{Y_{it}(D_{it})-Y_{it}(d^{0})}{D_{it}-d^{0}} & D_{it}\ne d^{0}\\
0 & D_{it}=d^{0}
\end{cases}
\]
is a per-unit effect of a deviation of the treatment $D_{it}$ from
the baseline level $d^{0}$. The weights satisfy $\sum^{T}_{t=1}E\left[\omega_{it}\right]=1$
and
\[
\omega_{it}\propto\left(D_{it}-d^{0}\right)\left\{ \left(\widetilde{D}_{it}-\overline{\widetilde{D}}_{i}\right)-\left(\widetilde{W}_{it}-\overline{\widetilde{W}}_{i}\right)'\delta^{W}_{\text{FE}}\right\} .
\]
\end{thm}
Theorem \ref{thm:Causal_TWFE} interprets $\beta^{D}_{\text{FE}}$
as a weighted average of per-unit treatment effects, with weights
summing to one but possibly negative. The weight function $\omega_{it}$
interacts a deviation of $D_{it}$ from the baseline level $d^{0}$
with a TWFE residual of $D_{it}$.

However, this weighted-average interpretation faces multiple challenges.
Three issues concern the assumptions needed for the weighted-average
interpretation, while a fourth issue arises even when this interpretation
holds.

First, the common trends assumption (Assumption \ref{assu:PT}) must
hold across all difference lengths $k$. This requirement extends
far beyond the canonical two-period DID setting, and becomes increasingly
difficult to justify as the number of time periods grows. Section
\ref{sec:Practical-Implications-and} explores practical implications
from this requirement in detail.

Second, the role of time-varying covariates in a TWFE regression differs
entirely from a two-period DID setting that allows for covariates.
When a TWFE regression includes time-varying covariates, a common
trends assumption must be made conditional on changes in covariates,
not their initial levels, as illustrated by Assumption \ref{assu:PT}.
By contrast, a common trends assumption is typically made conditional
on pretreatment observables in a two-period DID setting \citep{heckman1997matching,heckman1998matching,abadie2005semiparametric}.
As formalized by \citet{caetano2022difference}, a common trends assumption
that uses covariate changes is less theoretically favored than the
one that uses predetermined covariate levels. In particular, if the
treatment change $\Delta_{k}D_{it}$ influences the covariate change
$\Delta_{k}W_{it}$, then variation in $\Delta_{k}D_{it}$ with $\Delta_{k}W_{it}$
being fixed does not correspond to \emph{ceteris paribus} variation
that defines the causal effect of the treatment.\footnote{This type of issue is often called a ``bad control'' problem, a
term popularized by \citet[Section 3.2.3]{angrist2008mostly}. They
describe occupation as a ``bad control'' when examining the effects
of a college degree on earnings. Indeed, occupation is not held constant
when we typically define the \emph{ceteris paribus} effect of a
college degree on earnings.}

Third, the linearity and time homogeneity assumption (Assumption \ref{assu:Homogeneity})
can be restrictive unless all covariates are discrete and time-invariant.
This assumption demands that the relationship between treatment changes
and covariate changes be not only linear but also constant across
all time periods and difference lengths. Despite being essential for
causal interpretation, these requirements are typically assumed implicitly
rather than explicitly defended, unlike common trends assumptions.

Fourth, even when the weighted-average interpretation holds, the weights
$\omega_{it}$ can be negative, as similarly highlighted in a binary
treatment case \citep[e.g.,][]{Borusyak2018,Chaisemartin2020,goodman2018difference}.
Whether this poses a problem depends on the correlation between weights
and treatment effects. If some weights are negative and weights are
strongly correlated with per-unit treatment effects, then $\beta^{D}_{\text{FE}}$
can even lie outside the support of per-unit treatment effects $\tau_{it}$.
Nevertheless, negative weights alone are not necessarily problematic
if treatment effects are uncorrelated with weights, as noted by \citet{Chaisemartin2020}
in their Corollary 2. Appendix \ref{subsec:Causal-Interpretation-without}
shows that the weight function $\omega_{it}$ contains redundant variation
that is orthogonal to treatment effects under a more strict common
trends assumption, which may alleviate the concern about negative
weights in some cases.

Overall, these issues highlight that causal interpretation of TWFE
estimators involves more restrictive assumptions than commonly recognized,
particularly regarding the length over which common trends can be
justified, the role of time-varying covariates, and the need for linearity
and time homogeneity conditions, all of which are rarely scrutinized
in practice.

\subsubsection*{Remarks}

\citet{Chaisemartin2020} provide a weighted-average interpretation
of the TWFE coefficient that parallels Theorem \ref{thm:Causal_TWFE}
in their online appendix. They separately consider cases with nonbinary
treatment without covariates and binary treatment with time-varying
covariates. However, their proof can be extended to allow for nonbinary
treatment with both time-invariant and time-varying covariates. In
the current setup, their counterpart to Assumption \ref{assu:PT}
can be expressed as
\begin{equation}
E\left[\Delta_{1}Y_{it}(d^{0})|D_{i1},\ldots,D_{iT},W_{i1},\ldots,W_{iT},C_{i}\right]=E\left[\Delta_{1}Y_{it}(d^{0})|\Delta_{1}W_{it},C_{i}\right].\label{eq:CD_cond}
\end{equation}
They also assume linearity and time homogeneity of $E\left[\Delta_{1}Y_{it}(d^{0})|\Delta_{1}W_{it},C_{i}\right]$.
This assumption corresponds to Assumption \ref{assu:Homogeneity},
which imposes similar conditions on $E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]$
for $k=1,\ldots,T-1$.

Their assumption (\ref{eq:CD_cond}) conditions on entire treatment
and covariate histories. Thus, this single-period ($k=1$) assumption
implies the corresponding assumption for all difference lengths $k=1,\ldots,T-1$
when combined with their linearity and time homogeneity assumption.
By contrast, Assumption \ref{assu:PT} conditions only on concurrent
changes $(\Delta_{k}D_{it},\Delta_{k}W_{it})$, and the assumption
for $k=1$ does not imply the corresponding ones for longer differences.
For example, $\Delta_{2}Y_{it}(d^{0})$ may correlate with $\Delta_{2}D_{it}$
even when $\Delta_{1}$ terms show no correlation, due to feedback
between outcome changes in one period and treatment changes in another.

The contribution of Theorem \ref{thm:Causal_TWFE} relative to their
result lies in showing that Assumption \ref{assu:PT}, which relies
solely on concurrent changes rather than full histories, is central
to causal interpretation of the TWFE coefficient. The importance of
covariate changes rather than levels for TWFE interpretation is well-recognized
in canonical two-period settings, stemming from the known equivalence
between TWFE and FD. Theorem \ref{thm:Causal_TWFE} shows that concurrent
changes remain the fundamental basis for causal interpretation even
in multiperiod settings. \citet{caetano2024difference} also emphasize
the role of covariate changes for TWFE interpretation but focus on
the two-period case; they condition on full covariate histories in
their multiperiod extension. By contrast, Theorem \ref{thm:Causal_TWFE}
demonstrates the sufficiency of conditioning only on concurrent changes.

While this emphasis on concurrent changes clarifies the role of covariates,
the implications for treatment dynamics warrant closer examination.
Exogeneity of an entire treatment series $\left(D_{i1},\ldots,D_{iT}\right)$
as in (\ref{eq:CD_cond}), often called a strict exogeneity condition,
is a common assumption in panel data models. This assumption rules
out the possibility that a shock to the current outcome influences
future treatment status or is correlated with past treatment status.
Although constraints that Assumption \ref{assu:PT} imposes on treatment
dynamics appear weaker than strict exogeneity when considered for
each $k=1,\ldots,T-1$ individually, similar restrictions arise when
considered collectively. Specifically, the common trends assumption
for periods $t$ to $t+2k$ inherently makes it unlikely for potential
outcome changes $\Delta_{k}Y_{it}(d^{0})$ in the first $k$ periods
to influence treatment changes $\Delta_{k}D_{i,t+k}$ in the subsequent
$k$ periods. It also makes correlations between prior treatment changes
$\Delta_{k}D_{it}$ and subsequent potential outcome changes $\Delta_{k}Y_{i,t+k}(d^{0})$
practically implausible. Without excluding such scenarios, correlations
between $\Delta_{2k}Y_{it}(d^{0})$ and $\Delta_{2k}D_{it}$ would
arise, even in the absence of correlations between $\Delta_{k}Y_{it}(d^{0})$
and $\Delta_{k}D_{it}$ or between $\Delta_{k}Y_{i,t+k}(d^{0})$ and
$\Delta_{k}D_{i,t+k}$.

Yet a critical distinction of Assumption \ref{assu:PT} from strict
exogeneity is that these restrictions manifest as testable departures
from common trends at specific horizons rather than as a failure of
a monolithic condition that must be maintained or abandoned entirely.
This horizon-specific structure enables the diagnostic approach developed
in Section \ref{subsec:Diagnosing-Common-Trends}.

\subsection{FD Interpretation\label{subsec:FD-Interpretation}}

The causal interpretation of the $k$-period FD coefficient follows
directly from the TWFE framework, but invokes assumptions only for
the specific difference length $k$:

\begin{assumptionx}

\label{assu:PT-k}\textup{(Conditional Common Trends, $k$-period FD)}
There exists $d^{0}\in(\underline{d},\overline{d})$ such that the
following holds for any $t\in\{1,\ldots,T-k\}$:
\[
E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}D_{it},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right].
\]


\end{assumptionx}

\begin{assumptionx}

\label{assu:Variation-k}\textup{(Variation, $k$-period FD)} $\sum^{T-k}_{t=1}E\left[Var\left(\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right)\right]>0$.

\end{assumptionx}

\begin{assumptionx}

\label{assu:Homogeneity-k}\textup{(Linearity and Time Homogeneity, $k$-period FD)}
For $t\in\{1,\ldots,T-k\}$, the conditional expectation of $\Delta_{k}D_{it}$
given $\left(\Delta_{k}W_{it},C_{i}\right)$ can be expressed as
\[
E\left[\Delta_{k}D_{it}|\Delta_{k}W_{it},C_{i}\right]=\Delta_{k}W_{it}'\delta^{W}_{k}+C_{i}'\delta^{C}_{k,t}+\delta^{I}_{k,t},
\]
where the coefficient $\delta^{W}_{k}$ does not vary across $t$.

\end{assumptionx}

These assumptions are identical to Assumptions \ref{assu:PT}, \ref{assu:Variation},
and \ref{assu:Homogeneity}, except that they need only hold for the
specific $k$ rather than across all $k=1,\ldots,T-1$. As noted in
the preceding remarks, Assumption \ref{assu:PT-k} for smaller $k$
does not imply it holds for larger $k$.

The population $k$-period FD coefficient $\beta^{D}_{\text{FD},k}$
is obtained from a regression of $\Delta_{k}Y_{it}$ on $\Delta_{k}D_{it}$
controlling for $\Delta_{k}W_{it}$ and $C_{i}$ interacted with period
indicators. Under the stated assumptions, this coefficient has the
following causal interpretation.
\begin{thm}
\label{thm:Causal_FD}Under Assumptions \ref{assu:PO}, \ref{assu:PT-k},
\ref{assu:Variation-k}, and \ref{assu:Homogeneity-k},
\[
\beta^{D}_{\text{FD},k}=\sum^{T}_{t=1}E\left[\tau_{it}\omega^{(k)}_{it}\right],
\]
where $\tau_{it}$ is a per-unit effect defined in Theorem \ref{thm:Causal_TWFE}.
The weights satisfy $\sum^{T}_{t=1}E\left[\omega^{(k)}_{it}\right]=1$
and
\[
\omega^{(k)}_{it}\propto(D_{it}-d^{0})\sum^{T}_{s=1}\mathbbm{1}_{|t-s|=k}\left(\widetilde{D}_{it}-\widetilde{D}_{is}-\left(\widetilde{W}_{it}-\widetilde{W}_{is}\right)'\delta^{W}_{k}\right).
\]
\end{thm}
Theorem \ref{thm:Causal_FD} suggests that the $k$-period FD coefficient
admits a weighted-average interpretation of per-unit treatment effects,
with weights that depend on the unpredictable component of $k$-period
treatment changes.

The FD interpretation inherits several problems from the TWFE framework.
The common trends assumption remains difficult to justify for large
$k$, though researchers can mitigate this concern by choosing smaller
values of $k$. The identification strategy still relies on covariate
changes rather than predetermined covariate levels, creating the same
theoretical concerns about ceteris paribus variation. Covariate impacts
should be homogeneous across time periods $t$, though homogeneity
across difference lengths $k$ is no longer required. Finally, the
possibility of negative weights persists, complicating the interpretation
of the weighted average.

Among these inherited problems, those associated with covariates can
be readily addressed within the FD framework. Rather than controlling
for covariate changes $\Delta_{k}W_{it}$, FD regressions can be modified
to control for the initial covariate levels $W_{it}$ or other predetermined
variables. Moreover, the homogeneity restriction on covariate impacts
across time periods can be relaxed by allowing period-specific coefficients.

The FD framework thus offers greater flexibility than the TWFE approach
in addressing covariate-related concerns while requiring weaker assumptions
overall. Given that a TWFE regression is merely a pooled version of
FD regressions as shown in Section \ref{sec:Numerical-Equivalence},
this raises the question of whether such aggregation is desirable
when an individual FD coefficient may admit causal interpretation
under more plausible conditions.

\section{Practical Implications and Diagnostics\label{sec:Practical-Implications-and}}

\subsection{Discussion: When Should We Use TWFE?\label{subsec:Discussion}}

Section \ref{sec:Numerical-Equivalence} shows that the TWFE estimator
is a weighted average of FD estimators; Section \ref{sec:Causal-Interpretation}
establishes that both of them admit causal interpretations. However,
the FD interpretation invokes assumptions only for the specific difference
length $k$, whereas the TWFE interpretation demands these assumptions
hold simultaneously across all possible difference lengths. This asymmetry
raises a natural question: why aggregate across all available FD estimators
to form a TWFE estimator, when each individual FD estimator may provide
valid causal identification under weaker conditions?

In traditional panel data econometrics, the answer centers on efficiency
gains. Within the textbook fixed effects model, TWFE and all $k$-period
FD estimators identify the same population parameter, and the TWFE
estimator achieves greater efficiency under additional distributional
assumptions about the error term. Indeed, adding a constant effects
assumption to the assumptions in Section \ref{subsec:TWFE-Interpretation}
yields this equivalence:

\begin{assumptionx}

\label{assu:CE}\textup{(Constant Effects)} The per-unit treatment
effect defined in Theorem \ref{thm:Causal_TWFE} satisfies $\tau_{it}=\tau$
for all $(i,t)$.

\end{assumptionx}
\begin{cor}
If Assumptions \ref{assu:PO}, \ref{assu:PT}, \ref{assu:Variation},
\ref{assu:Homogeneity}, and \ref{assu:CE} hold, then $\beta^{D}_{FE}=\tau=\beta^{D}_{\text{FD},k}$
for all $k=1,\ldots,T-1$.
\end{cor}
However, this efficiency justification loses its force when the estimators
identify different parameters. The possibility of heterogeneous treatment
effects alone undermines this foundation, even when all other identifying
assumptions hold. Moreover, potential bias from violated common trends
assumptions at longer horizons can be too substantial to justify the
gains from reduced standard errors. Long differences may be less reliable
for causal identification if unobserved heterogeneity in trends grows
over time.\footnote{\citet{millimet2025mis} formalize a related mechanism within the
linear model: when unit-specific heterogeneity drifts over time, longer
differences introduce greater bias than shorter ones.} In addition, as noted in Section \ref{subsec:TWFE-Interpretation},
the common trends assumption essentially rules out feedback from past
outcome changes to future treatment changes within its time frame.
Over longer horizons, this assumption becomes harder to justify, as
the likelihood of such feedback increases.

One might alternatively justify TWFE aggregation by appealing to interest
in long-run effects, since TWFE regressions exploit long-run differences
as well as short-run differences. Identifying effects over $k$-period
horizons inevitably requires assuming common trends over at least
$k$ periods, making the assumption a necessary compromise on this
ground. However, even $k$-period FD coefficients have little connection
to $k$-period treatment effects when dynamic effects are present,
as formalized in Appendix \ref{subsec:Dynamic-Effects} under an assumption
of homogeneous treatment effects. The TWFE coefficient would therefore
represent a difficult-to-interpret weighted average of these already
complex effect measures, even under the strong homogeneous effect
assumption.

These considerations suggest that TWFE maintains clear advantages
only under three joint conditions: (1) common trends assumptions hold
for all difference lengths $k$, (2) treatment effects are static
rather than dynamic, and (3) treatment effects are linear and homogeneous
across units and time periods. When any of these conditions fails,
the case for TWFE becomes weaker. Researchers may then focus on FD
regressions with specific $k$, which inherit some limitations of
TWFE but allow more targeted identifying assumptions. Alternatively,
they may employ heterogeneity-robust estimators tailored to their
specific setting (e.g., \citealt{Chaisemartin2020,de2020difference};
\citealt{callaway2020difference,wooldridge2021two,de2022difference,Borusyak2018,callaway2021difference}),
though such methods may be more challenging to apply in broader settings
with nonbinary, nonstaggered treatments and time-varying covariates,
and some of them still rely on common trends assumptions across the
entire panel duration.

The relationship between TWFE and FD estimators provides valuable
diagnostic information about assumption violations. If TWFE and $k$-period
FD estimates differ substantially, this signals violation of at least
one of the three justifying conditions. Two of these---violations
of common trends assumptions for specific difference lengths and the
presence of dynamic effects---can be assessed through common trends
diagnostics developed in Section \ref{subsec:Diagnosing-Common-Trends}.
For the third condition concerning treatment effect homogeneity, the
weight diagnostics established in the recent literature (e.g., \citealp{Chaisemartin2020})
are useful for evaluating how severely violations affect the TWFE
coefficient, though with important caveats that arise from assuming
common trends at only one treatment level, as highlighted by \citet{fabre2022robustness}
in binary treatment settings and by Appendix \ref{subsec:Causal-Interpretation-without}
in broader settings.

\subsection{Diagnosing Violations of the Common Trends Assumption\label{subsec:Diagnosing-Common-Trends}}

The common trends assumption cannot be directly tested, since potential
outcome changes $\Delta_{k}Y_{it}(d^{0})$ are unobserved. In staggered
adoption settings, pre-trend tests provide suggestive evidence about
the plausibility of common trends. By contrast, under general treatment
paths there is no obvious way to scrutinize this assumption. Yet the
plausibility of this assumption can be assessed through testable implications
involving the relationship among nonconcurrent treatment and outcome
changes. This section develops diagnostic procedures that exploit
the temporal structure of panel data to flag violations of common
trends across different time horizons.

\subsubsection*{Sufficient Conditions for Common Trends}

The key insight underlying these diagnostics is that the common trends
assumption becomes implausible when treatment changes are systematically
related to nonconcurrent outcome changes. Suppose that for given $k\in\{2,\ldots,T-1\}$,
potential outcome changes in the first $\ell$ periods are correlated
with treatment changes in the subsequent $k-\ell$ periods, or treatment
changes in the first $\ell$ periods are correlated with potential
outcome changes in the subsequent $k-\ell$ periods. This would suggest
intertemporal outcome-to-treatment feedback or unobserved factors
that drive both treatment and outcome dynamics over the horizon of
$k$ periods, indicating a potential violation of the common trends
assumption.

More formally, for each $\ell\in\{1,\ldots,k-1\}$, the following
two conditions jointly comprise a sufficient condition for Assumption
\ref{assu:PT-k}:
\begin{description}
\item [{(Condition\,I)}] $E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]$
for $t\in\{1,T-k\}$.
\item [{(Condition\,II)}] $E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{k}W_{it},C_{i}\right]$
for $t\in\{1,\ldots,T-k\}$.
\end{description}
Condition I requires that potential outcome changes over periods $t$
to $t+\ell$ be mean independent of future treatment changes over
periods $t+\ell$ to $t+k$ as well as concurrent treatment changes,
conditional on covariates. Condition II requires that potential outcome
changes over periods $t+\ell$ to $t+k$ be mean independent of past
treatment changes over periods $t$ to $t+\ell$ as well as concurrent
treatment changes, conditional on covariates.

\subsubsection*{Testable Implications and Diagnostic Regressions}

To overcome the challenge that potential outcome changes are unobserved,
I consider the following decomposition of observed outcome changes,
which follows from Assumption \ref{assu:PO} and the definition of
per-unit treatment effects $\tau_{it}$.
\[
\Delta_{k}Y_{it}=\Delta_{k}Y_{it}(d^{0})+\left(\frac{\tau_{i,t+k}+\tau_{it}}{2}\right)\Delta_{k}D_{it}+\Delta_{k}\tau_{it}\left(\frac{D_{i,t+k}+D_{it}}{2}-d^{0}\right).
\]
While heterogeneity of per-unit treatment effects $\tau_{it}$ complicates
the analysis in general, combining Assumption \ref{assu:CE} and Condition
I yields:
\[
E\left[\Delta_{\ell}Y_{it}|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{\ell}Y_{it}(d^{0})|\Delta_{k}W_{it},C_{i}\right]+\tau\Delta_{\ell}D_{it}.
\]
Similarly, under Assumption \ref{assu:CE} and Condition II:
\[
E\left[\Delta_{k-\ell}Y_{i,t+\ell}|\Delta_{\ell}D_{it},\Delta_{k-\ell}D_{i,t+\ell},\Delta_{k}W_{it},C_{i}\right]=E\left[\Delta_{k-\ell}Y_{i,t+\ell}(d^{0})|\Delta_{k}W_{it},C_{i}\right]+\tau\Delta_{k-\ell}D_{i,t+\ell}.
\]
These expressions provide the foundation for empirical diagnostics:
under the joint hypothesis of Conditions I and II (sufficient for
common trends) and constant effects, regressions of outcome changes
on concurrent treatment changes and lead/lag treatment changes controlling
for covariates ($\Delta_{k}W_{it},C_{i}$) should yield zero coefficients
on the lead/lag terms.

The following regression specifications implement the diagnostic procedures:
\begin{eqnarray}
\Delta_{\ell}Y_{it} & = & \beta_{0}+\beta_{1}\Delta_{k-\ell}D_{i,t+\ell}+\beta_{2}\Delta_{\ell}D_{it}+C_{i}'\gamma_{t}+\Delta_{k}W_{it}'\delta+\epsilon_{it},\label{eq:lead_reg}\\
\Delta_{k-\ell}Y_{i,t+\ell} & = & \beta_{0}+\beta_{1}\Delta_{\ell}D_{it}+\beta_{2}\Delta_{k-\ell}D_{i,t+\ell}+C_{i}'\gamma_{t}+\Delta_{k}W_{it}'\delta+\epsilon_{it}.\label{eq:lag_reg}
\end{eqnarray}

Under the joint hypothesis of Conditions I and II (sufficient for
Assumption \ref{assu:PT-k}) and constant effects (Assumption \ref{assu:CE}),
we should observe $\beta_{1}=0$ for each regression. These diagnostics
can be implemented for increasing values of $k$, starting with $k=2$.
Since Conditions I and II also demand exogeneity of concurrent treatment
changes, violation of common trends for some $k^{*}$ implies violation
of Conditions I and II for larger $k>k^{*}$, which in turn signals
(though does not imply) violation of Assumption \ref{assu:PT-k}.

These diagnostics resemble distributed-lag regressions but differ
in a key respect: these diagnostics exploit treatment changes over
nonoverlapping time periods, whereas distributed-lag regressions,
being a form of multivariate TWFE regression, pool all possible FDs
and therefore incorporate numerous overlapping treatment changes,
obscuring the connection to specific common trends assumptions.

\subsubsection*{Limitations and Interpretation}

These diagnostic procedures should not be interpreted as direct tests
of the common trends assumption---a limitation inherent to all such
diagnostics, including standard pre-trend tests. The link between
diagnostic results and the common trends assumption involves multiple
inferential steps, each with important caveats.

First, Conditions I and II are sufficient but not necessary for the
common trends assumption (Assumption \ref{assu:PT-k}). Their violation
therefore does not in itself imply that the assumption fails. Nevertheless,
for the common trends assumption to hold despite violations of Conditions
I and II, correlations in different subperiods must coincidentally
cancel out when aggregated over the full $k$-period horizon. While
such cancellation is theoretically possible, it would demand a precise
offsetting of intertemporal feedback, which is difficult to justify
on economic grounds.

Second, $\beta_{1}=0$ in the lead and lag diagnostics is not sufficient
for Conditions I and II to hold, since the diagnostics are silent
about the exogeneity of concurrent treatment changes---a key part
of these conditions. This mirrors a well-known limitation of pre-trend
tests in staggered adoption designs, where absence of differential
pre-trends does not guarantee common trends in post-treatment periods.

Third, the purest interpretation of the diagnostics requires the constant
effects assumption (Assumption \ref{assu:CE}). When treatment effects
$\tau_{it}$ vary and depend on treatment paths, the coefficients
$\beta_{1}$ can be contaminated by this heterogeneity rather than
purely reflecting violations of the sufficient conditions. Yet, this
limitation does not undermine diagnostic value: if $\beta_{1}\ne0$
is detected, it indicates either a violation of the sufficient conditions
for common trends or the presence of treatment effect heterogeneity
that complicates causal interpretation. Both scenarios suggest that
standard TWFE estimates may be problematic, making the diagnostic
informative regardless of which interpretation applies.

These caveats notwithstanding, the diagnostics remain informative
for assessing the plausibility of common trends at different horizons,
especially since intertemporal outcome--treatment correlation captures
economically meaningful threats to identification, such as anticipation,
policy responses to past outcomes, or persistent unobserved shocks.
They also have the ancillary benefit of detecting dynamic treatment
effects, since $\beta_{1}\neq0$ can arise when treatment effects
depend on treatment history (Appendix \ref{subsec:Dynamic-Effects}).
Accordingly, the diagnostics help evaluate whether TWFE estimates
are likely to be interpretable as causal effects, and whether estimators
that rely more heavily on short-run (or long-run) variation may be
preferable.

\section{Empirical Illustration\label{sec:Empirical-Illustration}}

This section illustrates insights from the econometric results by
studying the impact of minimum wages on employment using TWFE regressions.
The TWFE regression analysis based on state-level panel data in the
US has attracted considerable attention for decades. The econometric
framework of this paper, applicable to a broad class of multiperiod
panel settings, is particularly well-suited for this empirical context.
The treatment variable, the state minimum wage, is continuous and
changes many times in a given state. The existing insights from the
TWFE regression literature, typically addressing binary treatments
and staggered adoption designs, do not directly apply to this setting.

To present the estimates, I use a state--year panel of employment
outcomes and minimum wage laws in 50 US states and the District of
Columbia from 1979 to 2019. As a dependent variable in a TWFE regression,
I use the log employment rate of teens (ages 16--19) from the Current
Population Survey (CPS).\footnote{These sample and outcome variable choices follow \citet{manning2021elusive},
who uses CPS data from 1979 to 2019, aggregated into quarterly state-level
observations. However, my analysis is based on annual frequency to
incorporate the BDS data.} Another dependent variable is the net job creation rate from the
Business Dynamics Statistics (BDS), using four major sectors with
high shares of workers without college education.\footnote{The four sectors are: Construction, Manufacturing, Retail Trade, and
Accommodation and Food Services. Using the data from all sectors does
not qualitatively change the estimates presented below, but the estimated
effects become smaller in magnitude.} These employment measures, as stock and flow variables, capture different
aspects of labor market adjustment. The focus on age groups or sectors
more likely affected by the minimum wage is practical, since the majority
of workers in the US are not directly affected by the minimum wage.
The minimum wage data are from \citet{Vaghul2019}, and I use the
log of the annual average of the state minimum wage as the treatment
variable in each regression. I use census region indicators as covariates.

I begin by illustrating the weighted-average characterization suggested
in Section \ref{subsec:Weighted-Average-Relationship}. Figure \ref{Fig:TWFE_FD}
presents the FD coefficients $\widehat{\beta}_{\text{FD},k}$ for
$k=1,\ldots,40$ as dots accompanied by standard error bars, the TWFE
coefficient $\widehat{\beta}_{\text{FE}}$ as a horizontal dashed
line, and the weights given by Theorem \ref{theorem:FE_not_equal_FD}
as bars. Since all covariates are time-invariant, the adjustment term
in Theorem \ref{theorem:FE_not_equal_FD} can be ignored. While the
inclusion of time-varying covariates would result in a deviation from
the exact weighted-average interpretation, Appendix \ref{subsec:The-Relationship-between}
finds that the deviation is minimal in this setting even after including
time-varying covariates such as log per-capita income and log teen
population.

\begin{figure}[t]
\caption{The TWFE and FD Coefficients on Log Minimum Wage}

\label{Fig:TWFE_FD}

\vspace{0.5em}
\centering
\begin{threeparttable}
\vspace{0.25em}

\begin{minipage}[t]{0.5\columnwidth}
(a) Log Employment Rate of Teens

\includegraphics[width=1\columnwidth]{Figure1}
\end{minipage}\,\,\,\,\,\,\,\,
\begin{minipage}[t]{0.5\columnwidth}
(b) Net Job Creation Rate (\%)

\includegraphics[width=1\columnwidth]{Figure2}
\end{minipage}

\vspace{0.5em}
\begin{tablenotes}
\small

\item Notes: The FD coefficients are illustrated with standard error
bars. Covariates are census region indicators. The standard errors
are robust to to heteroskedasticity and correlation across observations
on the same state.

\end{tablenotes}
\end{threeparttable}
\end{figure}

In Panel (a) of Figure \ref{Fig:TWFE_FD}, the TWFE coefficient is
$\widehat{\beta}_{\text{FE}}=-0.266$, whereas the FD coefficient
is $\widehat{\beta}_{\text{FD},1}=-0.001$ with a one-year gap and
becomes larger in magnitude as the gap increases. While the difference
among ``short'' FD, ``long'' FD, and TWFE estimates itself has
long been recognized in the minimum wage effects literature \citep[e.g.,][]{neumark1992employment},
an explicit numerical relationship among these estimates has not been
previously established. The theoretical results in Section \ref{sec:Numerical-Equivalence}
demonstrate that a TWFE regression aggregates these heterogeneous
FD coefficients into a single coefficient. The weighted average of
all 40 FD coefficients is indeed $-0.266$. Panel (b) of Figure \ref{Fig:TWFE_FD}
presents analogous plots for a regression of the net job creation
rate. Even though the TWFE coefficient is $\widehat{\beta}_{\text{FE}}=-2.65$,
one-year FD coefficient $\widehat{\beta}_{\text{FD},1}=-5.74$ is
larger in magnitude and some long-run FD coefficients have different
signs. The weighted average of the FD coefficients is again confirmed
to be identical to the TWFE coefficient.

The weights depicted in Figure \ref{Fig:TWFE_FD} are identical across
the two panels. This consistency arises because the weights depend
only on the explanatory variables, not on the dependent variables.
As observed in (\ref{eq:weight_cov}), these weights are associated
with the unexplained variation of treatment changes. Consequently,
the weights exhibit a distinct pattern: they are small for both short-
and long-run FD coefficients, albeit for different reasons. For the
short-run FD coefficients, the weights are small because minimum wage
adjustments tend to be incremental over short periods. By contrast,
for the long-run FD coefficients, the weights are small due to the
limited number of observations of long-run changes.

The substantial differences among FD coefficients documented in Figure
\ref{Fig:TWFE_FD} raise fundamental questions about causal interpretation.
Traditional panel data econometrics justifies TWFE's aggregation of
these estimates on efficiency grounds, but this efficiency justification
is valid only when all FD estimators identify the same population
parameter. The patterns observed here suggest that different FD estimators
likely identify distinct parameters rather than noisy estimates of
a common effect. Indeed, the hypothesis that these coefficients are
all identical can be rejected even at a significance level of 0.01\%.
More concerning, some of these estimates may not identify causal parameters
at all if the common trends assumption is violated. With a 41-year
panel spanning 1979--2019, it seems unlikely that this assumption
holds uniformly across all $k=1,...,40$. The procedures developed
in Section \ref{subsec:Diagnosing-Common-Trends} provide a systematic
way to evaluate these concerns across different time horizons.

Table \ref{Table:Reg1} presents regressions based on equations (\ref{eq:lead_reg})
and (\ref{eq:lag_reg}). These diagnostics test whether past and future
treatment changes have zero coefficients when regressing outcome changes
on them, controlling for concurrent treatment changes and covariates.
For practical presentation, I focus on even values of $k\le10$ with
lead/lag length $\ell=k/2$.

\begin{table}[t]
\caption{Diagnosing Common Trends Violation for Minimum Wage Changes}
\label{Table:Reg1}

\centering
\begin{threeparttable}
\setlength\tabcolsep{0.15em}

\begin{tabular}{llcccccccccccccc}
 &  &  &  &  &  &  &  &  &  &  &  &  &  &  & \tabularnewline
\hline
\multirow{2}{*}{} &  & \multicolumn{2}{c}{$k=2$} &  & \multicolumn{2}{c}{$k=4$} &  & \multicolumn{2}{c}{$k=6$} &  & \multicolumn{2}{c}{$k=8$} &  & \multicolumn{2}{c}{$k=10$}\tabularnewline
\cline{3-4}\cline{6-7}\cline{9-10}\cline{12-13}\cline{15-16}
 &  & Lead & Lag &  & Lead & Lag &  & Lead & Lag &  & Lead & Lag &  & Lead & Lag\tabularnewline
\hline
\multicolumn{16}{l}{\textbf{Panel A. Log Teen Employment Rate}}\tabularnewline
Future change &  & 0.101 &  &  & 0.081 &  &  & --0.013 &  &  & --0.094 &  &  & --0.092 & \tabularnewline
 &  & (0.061) &  &  & (0.056) &  &  & (0.056) &  &  & (0.061) &  &  & (0.050) & \tabularnewline
Past change &  &  & --0.128 &  &  & 0.017 &  &  & 0.024 &  &  & --0.021 &  &  & --0.121\tabularnewline
 &  &  & (0.070) &  &  & (0.040) &  &  & (0.076) &  &  & (0.085) &  &  & (0.116)\tabularnewline
Concurrent &  & --0.034 & 0.027 &  & 0.025 & --0.008 &  & 0.032 & 0.036 &  & --0.002 & 0.052 &  & 0.009 & 0.013\tabularnewline
change &  & (0.061) & (0.058) &  & (0.067) & (0.066) &  & (0.092) & (0.067) &  & (0.107) & (0.062) &  & (0.118) & (0.054)\tabularnewline
P-value &  & \multicolumn{2}{c}{0.019} &  & \multicolumn{2}{c}{0.288} &  & \multicolumn{2}{c}{0.936} &  & \multicolumn{2}{c}{0.048} &  & \multicolumn{2}{c}{0.002}\tabularnewline
\hline
\multicolumn{16}{l}{\textbf{Panel B. Net Job Creation Rate (\%)}}\tabularnewline
Future change &  & 1.88 &  &  & 4.24 &  &  & 2.99 &  &  & 2.79 &  &  & 4.18 & \tabularnewline
 &  & (1.23) &  &  & (1.53) &  &  & (1.16) &  &  & (1.28) &  &  & (1.48) & \tabularnewline
Past change &  &  & 0.19 &  &  & --4.24 &  &  & --3.58 &  &  & --1.45 &  &  & --1.32\tabularnewline
 &  &  & (2.60) &  &  & (1.55) &  &  & (1.53) &  &  & (2.12) &  &  & (2.27)\tabularnewline
Concurrent &  & --6.17 & --5.98 &  & --4.48 & --5.10 &  & --5.42 & --5.55 &  & --4.87 & --5.00 &  & --4.13 & --5.16\tabularnewline
change &  & (2.35) & (2.72) &  & (1.78) & (1.64) &  & (1.87) & (1.46) &  & (1.88) & (1.22) &  & (2.16) & (1.31)\tabularnewline
P--value &  & \multicolumn{2}{c}{0.280} &  & \multicolumn{2}{c}{0.004} &  & \multicolumn{2}{c}{0.017} &  & \multicolumn{2}{c}{0.037} &  & \multicolumn{2}{c}{0.005}\tabularnewline
\hline
\end{tabular}

\vspace{0.2em}
\begin{tablenotes}
\footnotesize

\item Note: Each column reports a regression of outcome changes on
future or past minimum wage changes using equation (\ref{eq:lead_reg})
or (\ref{eq:lag_reg}) with $\ell=k/2$, controlling for concurrent
minimum wage changes and census region indicators. Standard errors
are in parentheses and robust to heteroskedasticity and correlation
across observations on the same state. P--values test the joint null
that future and past change coefficients equal zero.

\end{tablenotes}
\end{threeparttable}
\end{table}

For teen employment, the hypothesis that both past and future treatment
change coefficients equal zero is rejected at the 5\% significance
level for $k=$ 2, 8, and 10. For $k=2$, the future change coefficient
(0.101) and the past treatment change coefficient (--0.128) exceed
the concurrent change coefficient in magnitude, suggesting short-term
feedback mechanisms or dynamic treatment effects. For $k=$ 8 and
10, both past and future minimum wage changes have negative coefficients,
which raises the possibility of unobserved long-term factors that
simultaneously drive minimum wage policies and teen employment trends
over extended periods.

The job creation rate estimates also exhibit patterns that are hard
to reconcile with the common trends assumption. The hypothesis of
zero coefficients is rejected at the 5\% level for $k=$ 4, 6, 8,
and 10. Most notably, the future treatment change coefficients are
consistently positive and significant across these horizons, indicating
that increases in net job creation systematically precede minimum
wage increases. This pattern may reflect intertemporal outcome-to-treatment
feedback, where policymakers respond to improving labor market conditions
by raising minimum wages.

Overall, the diagnostics indicate that the common trends assumption
is most credible only at short horizons. This suggests that estimators
relying exclusively on short-run changes may be more appropriate than
the TWFE estimator for identifying causal effects in this setting.
Yet, because employment is a stock variable that may adjust only gradually,
an exclusive focus on short-run variation may never recover the economically
relevant long-run effect, as emphasized in the literature (e.g., \citealp{meer2016effects}).
This tension underscores the difficulty of designing empirical strategies
that are both credible and substantively informative, and highlights
the value of diagnostics in making such trade-offs explicit.

\section{Conclusion\label{sec:Conclusion}}

Building on a long-overlooked numerical equivalence between TWFE and
pooled FD regressions, this paper provides insights into how TWFE
regressions translate data into the coefficient of interest. A key
observation is that TWFE regressions reflect associations among changes
in the variables of interest, pooling all possible short- and long-run
comparisons. As a natural consequence of this property, causal interpretation
of the TWFE coefficient relies on the common trends assumption holding
across all time horizons. To address this challenge, I develop diagnostic
procedures that exploit testable implications to help researchers
assess whether the common trends assumptions are plausible across
different time horizons.

This paper focuses on diagnosing problems of TWFE regressions rather
than proposing solutions. While recent literature has developed various
alternative estimators in specific contexts such as binary treatments
or staggered adoption designs, extending robust causal inference methods
to broader settings with general treatment paths and time-varying
covariates remains an important challenge for future research.


\bibliographystyle{econ-econometrica}
\bibliography{DID}


\newpage{}