EconBase
← Back to paper

Trading Scope for Credibility in Difference-in-Differences

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

49,750 characters

Trading Scope for Credibility in Difference-in-Differences


\maketitle

\begin{abstract}
\noindent When parallel trends fails for some treated cohorts but not others, the average treatment
effect on the treated (ATT), an average over all of them, is exactly the target that becomes hard to
recover. We propose changing the estimand rather than defending it. The credible-subpopulation local ATT
(LATT) is the effect for the subpopulation of cohorts whose parallel trends is credible, and it is
point-identified under parallel trends for the selected cohorts alone, a weaker requirement that can hold
when the ATT's fails. It is estimated by reweighting standard group-time effects toward those cohorts, and
paired with honest sensitivity bounds on the residual violation that a pre-trend screen cannot rule out. The method's advantage grows with how informative pre-trends are about post-treatment
violations, as simulations confirm. In an application, a significantly positive pooled estimate of the
shale boom's effect on local house prices proves to rest on cohorts already trending before onset, and
the credible subpopulation reveals no effect.
\end{abstract}

\medskip
\noindent\textbf{Keywords:} Difference-in-differences, Staggered adoption, Parallel trends, Robust inference, Sensitivity analysis, Partial identification.

\smallskip
\noindent\textbf{JEL classification:} C1, C23, C51.

\newpage

\section{Introduction}

Difference-in-differences is among the most widely used research designs in empirical economics, and its
credibility rests on the parallel trends assumption, that absent treatment treated and comparison units
would have evolved in parallel. Researchers assess it by testing for pre-treatment differences in trends,
but that practice is fragile. Pre-trends tests are often underpowered against economically meaningful
violations \citep{bilinski2020, freyaldenhoven2019, kahnlang2020, roth2022pretrends}, conditioning on
passing them induces a pre-test selection bias \citep{roth2022pretrends}, and even a clean pre-period
does not guarantee a clean post-period, since nothing prevents a confound from arriving at the moment of
treatment \citep{kahnlang2020}.

A growing literature, surveyed in \citet{roth2023trending},
has responded by relaxing parallel trends itself. Building on \citet{manski2018}, \citet{rambachan2023}
abandon exact parallel trends and instead bound how different the post-treatment violations can be from
the observed pre-trends, reporting a uniformly valid confidence set for the ATT in place of a point
estimate. A Bayesian variant \citep{kwonroth2024} places an empirical-Bayes prior on the violation $\delta$ to
sharpen the set into a posterior mean and credible interval, and \citet{dechaisemartin2026} instead
ranks the post-treatment difference-in-differences against the pre-treatment ones. All keep the ATT as
the target and use the pre-trends to extrapolate the violation across periods.

The ATT is where a structural limitation appears. When parallel trends
fails for some cohorts but holds for others, the ATT, an average over all treated cohorts including the
offending ones, is precisely the object that is hard to recover. A point estimate of it is then biased,
and the sensitivity set, though honest, can be wide and uninformative, because it must accommodate the
worst offenders in the average.

This paper takes a different route. When the population target is not credibly identified, we retreat to
a subpopulation where it is. This is the logic of the local average treatment effect \citep{imbens1994},
of the optimal-subpopulation average treatment effect under limited overlap \citep{crump2009}, and of
overlap weighting \citep{li2018}. In each case one trades the original estimand for a credibly-identified
effect on a selected subpopulation, and defends the trade by characterizing whom the subpopulation
comprises.

We develop the staggered-DiD instance of this principle. The credible-subpopulation LATT is the weighted
average of group-time effects restricted to cohorts whose parallel trends is credible, and it is
point-identified under parallel trends for the selected cohorts alone, a weaker requirement than the ATT's
that can hold when the latter fails, since the excluded cohorts' violations do not enter it.
Estimation is a transparent reweighting of standard group-time effects \citep{callaway2021}, so the
method applies to any heterogeneity-robust estimator that delivers them.\footnote{These estimators
\citep{callaway2021, sun2021, borusyak2024, dechaisemartin2020} repair the aggregation failure of
two-way fixed effects under staggered adoption \citep{goodmanbacon2021}, whose implicit weights can turn
negative under treatment-effect heterogeneity. That repair concerns aggregation, not identification,
since they maintain parallel trends and are biased when it fails, which is the failure we address.} Credibility
is assessed from the pre-trends, making the reweighting a data-dependent selection, whose inference we
treat with the sensitivity machinery of \citet{rambachan2023}.

The method thus occupies a distinct point on a frontier of credibility, precision, and scope. Against a
point estimate of the ATT it buys credibility by narrowing the target, and against a HonestDiD set for
the ATT it buys a sharper, point-identified answer. Its advantage is governed by how informative
pre-trends are about post-treatment violations. When they are uninformative, selection buys nothing and
the assumption must rest on economic context, and there the relative-magnitudes restriction is
self-undermining, reporting false precision when pre-trends are flat where a fixed, researcher-set bound
is not.

The remainder of the paper is organized as follows. Section~\ref{sec:approach} fixes the staggered-DiD
setup, defines the credible-subpopulation LATT, and develops its identification, estimation, and
sensitivity-based inference. Section~\ref{sec:sim} reports the simulation study,
Section~\ref{sec:app} applies the method to the local house-price effect of the shale boom, and
Section~\ref{sec:conc} concludes.

\section{The credible-subpopulation LATT}\label{sec:approach}

\subsection{Setup}\label{sec:setup}

We adopt the staggered adoption framework of \citet{callaway2021}. There are periods
$t=1,\dots,T$ and units $i$ indexed by their first treatment date $G_i \in \{2,\dots,T\}\cup\{\infty\}$,
where $G_i=\infty$ denotes never-treated. Let $Y_{it}(g)$ denote the potential outcome in period $t$
if unit $i$ is first treated in period $g$, and $Y_{it}(\infty)$ the never-treated potential outcome.
The building-block causal parameter is the group-time average treatment effect on the treated,
\begin{equation}
\mathrm{ATT}(g,t) = \mathbb{E}\!\left[\,Y_{it}(g)-Y_{it}(\infty)\mid G_i=g\,\right].
\end{equation}
Summary parameters are weighted aggregations $\theta = \sum_{g}\sum_t w(g,t)\,\mathrm{ATT}(g,t)$ with
researcher-chosen weights \citep{callaway2021}, of which the overall ATT and the event-study path are
special cases.

It is convenient to work with the event-study representation and the decomposition of
\citet{rambachan2023}. Let $\hat\beta=(\hat\beta_{\mathrm{pre}}',\hat\beta_{\mathrm{post}}')'$ be an
asymptotically normal vector of event-study coefficients estimated by any of the heterogeneity-robust
procedures above, with $\sqrt{n}(\hat\beta-\beta)\to\mathcal{N}(0,\Sigma)$. The estimand decomposes as
\begin{equation}
\beta = \tau + \delta, \qquad \tau_{\mathrm{pre}}=0,
\label{eq:decomp}
\end{equation}
where $\tau$ collects the causal effects (zero before treatment, by no anticipation) and $\delta$ is
the differential trend between treated and comparison groups that would have occurred absent
treatment. Parallel trends is the restriction $\delta_{\mathrm{post}}=0$, under which
$\beta_{\mathrm{post}}=\tau_{\mathrm{post}}$, and a pre-trends test is a test of $\delta_{\mathrm{pre}}=0$.

We will require the analogous objects at the cohort level. Write $\delta_g$ for the differential
trend of cohort $g$ and say that parallel trends holds for $g$ if $\delta_{g,\mathrm{post}}=0$.
A key feature of applications is that this may hold for some cohorts and fail for others. Defending the
timing of treatment as unrelated to the outcome path is already delicate in a two-group design, and
staggered adoption raises the burden rather than lowering it, since each cohort's timing must be
defended on its own, one cohort's adoption perhaps as good as randomly timed while another's coincides
with a confounding shock.

Finally, we distinguish two regimes for how cohorts are judged credible, because they carry different
inferential guarantees. Under ex-ante (or independent) selection, credibility is decided from
pre-determined covariates, institutional knowledge, or an independent data split, so that the selected
set is a fixed $S$ statistically independent of the estimation errors in $\hat\beta$. Under data-driven
selection, credibility is assessed from the realized pre-trends $\hat\beta_{\mathrm{pre}}$, so that the
selected set $\hat S$ is random and depends on the estimation sample, with a deterministic population
counterpart $S^\ast$, the set a population pre-trend statistic would select. The identification argument
below holds for any set held fixed, and so identifies the random target $\theta_{\hat S}$ conditional on the
realized $\hat S$. Identifying the fixed population parameter $\theta_{S^\ast}$ instead requires $\hat S$ to
recover $S^\ast$, a selection-consistency condition we return to under inference
(Proposition~\ref{prop:oracle}). The distinction is thus minor for identification but central for inference,
and we treat it explicitly below.

\subsection{Estimand and identification}

Let $S$ denote a set of cohorts judged to have credible parallel trends, produced by a selection rule
$R$ that we specify below. We define the target as the treatment effect for that subpopulation.

\begin{definition}[Credible-subpopulation LATT]
For cohort weights $w_g>0$ (e.g.\ proportional to cohort size) and a within-cohort aggregation $a(\cdot)$, a
linear functional of a cohort's post-treatment vector (a fixed event-time, selecting one coordinate, or an
average over post-treatment periods), the credible-subpopulation local ATT, defined for any nonempty $S$, is
\begin{equation}
\theta_S \;=\; \frac{\sum_{g\in S} w_g\,\mathrm{ATT}(g,\cdot)}{\sum_{g\in S} w_g},
\qquad \mathrm{ATT}(g,\cdot)\;=\;a\!\left(\tau_{g,\mathrm{post}}\right),
\end{equation}
the aggregated causal effect of cohort $g$.
\end{definition}

The value of the target is that it is identified under a weaker condition than the ATT.

\begin{proposition}[Identification]\label{prop:id}
If parallel trends holds for every $g\in S$, i.e.\ $\delta_{g,\mathrm{post}}=0$ for all $g\in S$, then
$\theta_S$ is point-identified by the reweighted aggregated group-time effects,
$\theta_S = \big(\sum_{g\in S} w_g\,a(\beta_{g,\mathrm{post}})\big)/\sum_{g\in S} w_g$, and this holds
irrespective of whether $\delta_{g,\mathrm{post}}\neq 0$ for cohorts $g\notin S$.
\end{proposition}

\begin{proof}
For $g\in S$, $\beta_{g,\mathrm{post}} = \tau_{g,\mathrm{post}}+\delta_{g,\mathrm{post}}
= \tau_{g,\mathrm{post}}$ by the hypothesis and \eqref{eq:decomp}, so $a(\beta_{g,\mathrm{post}})
= a(\tau_{g,\mathrm{post}}) = \mathrm{ATT}(g,\cdot)$ by linearity of $a$. The reweighted sum over $S$
therefore equals the weighted average of the aggregated causal effects, which is $\theta_S$ by definition.
Cohorts outside $S$ receive zero weight, so their differential trends do not enter.
\end{proof}

Proposition~\ref{prop:id} is elementary, but it is the paper's central point. Changing the target to the
credible subpopulation weakens the identifying requirement from parallel trends on all cohorts to parallel
trends on $S$, which can hold when the former fails, at the cost of a narrower question whose subpopulation
$S$ must then be characterized. The weakening is genuine but conditional, and two things must be kept
separate. The identifying content is the substantive restriction $\delta_{g,\mathrm{post}}=0$ for $g\in S$,
an assumption about post-treatment trends. The flatness screen of Section~\ref{sec:select} is a statistical
classification rule that supplies evidence for that restriction from the pre-period, not the restriction
itself, since a cohort flat before treatment may still drift after it. Point identification of the causal
$\theta_S$ thus holds only when the restriction does, and the screen cannot certify it. We therefore treat
the point-identified $\theta_S$ as a benchmark and pair it with the honest sensitivity bounds developed
below, which bound the residual post-treatment violation the screen leaves behind and, in doing so,
return a set rather than a point when that violation is nonzero.

The construction uses staggered timing only to supply the groups. What it requires is a partition of the
treated population into groups whose parallel trends can each be assessed. Adoption cohorts are the
natural groups, and their credibility heterogeneity comes for free from the variation in when and why
units adopt. In a non-staggered design the same construction applies with groups defined by covariates or
geography, provided the groups differ in credibility and there are enough pre-periods to estimate each
group's pre-trend rather than chase noise. The groups should accordingly be subpopulations large enough
to assess, not individual units, whose pre-trends are too noisy to select on without reviving the
pre-test bias in its most severe form.

\subsection{Selection and estimation}\label{sec:select}

The selection rule $R$ maps the data to the set $S$, and its form fixes both which subpopulation
$\theta_S$ describes and the inferential guarantees available. It takes two forms, matching the two
regimes of Section~\ref{sec:setup}. Under ex-ante selection, $S$ is fixed before estimation from
pre-determined covariates, institutional knowledge of which cohorts were plausibly confounded, or an
independent data split. Under data-driven selection, credibility is read from the estimated
pre-trends, and a cohort enters $S$ when its estimated pre-trend is close to flat,
\begin{equation}
S \;=\; \big\{\, g : \max_{e<0}\, \lvert \hat\beta_{g,\mathrm{pre}}(e)\rvert \le c \,\big\},
\label{eq:rule}
\end{equation}
for a threshold $c$.

This is the direct reading of parallel trends. The assumption is that the treated-comparison difference
would have stayed constant absent treatment, so a cohort is credible when its pre-treatment difference
is close to constant, and the screen keeps exactly those. It is posed in the same currency as the honest
inference of the next subsection, which bounds how far the post-treatment difference could have wandered
from flat, so selection and inference test one object, the level of the parallel-trends
violation.\footnote{A researcher willing to trust that a straight pre-trend would have continued linearly
can relax flatness to approximate linearity, screening instead on the pre-trend's curvature (its second
difference) and pairing it with the smoothness sensitivity class of \citet{rambachan2023} and a
linear-extrapolation estimator \citep{armstrong2018}. This buys the steep-but-straight cohorts that
flatness discards, at the cost of the extrapolation assumption. The principle is the same in either
currency. The screen should be posed in whichever functional the sensitivity class bounds, flatness with
the level bound and curvature with the smoothness class, since level and curvature rank cohorts
differently and select different sets.}

The threshold trades scope
against credibility, with a lax $c$ retaining more cohorts but admitting larger residual violations and a
strict $c$ the reverse. Because rule~\eqref{eq:rule} reads $S$ off the noisy estimates $\hat\beta_{\mathrm{pre}}$, the selected set
is random and dependent on the estimation sample, which is the source of the pre-test bias below and of the
inference subtleties that follow it.

We do not recommend committing to a single $c$. The sensitivity interval of
Section~\ref{sec:approach} is applied to whichever set $c$ selects, and its bound $M$ must dominate that
set's residual \emph{post}-treatment violation, so the effect of $c$ on the required $M$ runs through how
far screening on flat pre-trends also flattens post-trends. When pre-trends are informative about
post-treatment violations, a stricter $c$ yields a cleaner set that a smaller $M$ suffices to cover, while
a laxer $c$ buys scope at the cost of a larger $M$, so the threshold trades scope for precision rather than
credibility for nothing. When they are not, tightening $c$ need not shrink the post-treatment violation and
the required $M$ can be nonmonotone in $c$. Either way the honest object to report is the estimand together
with its interval traced as a function of $c$, the selection analogue of the breakdown curve we report for
$M$. A reader then sees the scope-credibility-precision frontier directly and need not trust a single
cutoff or a presumed monotone mapping. When a
single value is nonetheless wanted, we recommend calibrating it to negligibility rather than to
significance, setting $c$ to the largest pre-trend that would bias $\theta_S$ by less than a stated
tolerance, in the spirit of the non-inferiority tests of \citet{bilinski2020}, rather than to a critical
value of a pre-trends test, which would reinherit the low power that motivates the paper. The
application of Section~\ref{sec:app} reports such a path.

The estimator $\hat\theta_S$ simply reweights, over the selected cohorts, the sample group-time effects
of whichever heterogeneity-robust estimator one uses, and nothing below depends on that choice.

\begin{proposition}[Consistency and efficiency]\label{prop:cons}
Under ex-ante selection and the regularity conditions of the underlying group-time estimator,
$\hat\theta_S$ is consistent for $\theta_S$. If in addition parallel trends holds for all cohorts,
$\hat\theta_S$ is consistent for the ATT when either treatment effects are homogeneous or $S$ comprises all
cohorts, and otherwise for the selected-cohort LATT, which differs from the ATT by the gap between the
selected- and full-population weighted averages of the cohort effects. In either case there is an
efficiency loss relative to estimators that use all cohorts.
\end{proposition}

Under data-driven selection the rule~\eqref{eq:rule} couples $S$ to the estimation errors: a confounded
cohort whose pre-trend noise happens to fall below $c$ is admitted, and its differential trend enters
$\hat\theta_S$. This is a pre-test selection bias in the sense of \citet{roth2022pretrends}, with a
definite source and a definite fate.

\begin{proposition}[Pre-test bias]\label{prop:pretest}
Under data-driven selection, $\hat\theta_S$ carries a finite-sample bias equal to the weighted average
of the admitted confounded cohorts' post-treatment differential trends, times their probability of
admission. Whether it vanishes turns on detectability. Call a confounded cohort detectable if its
population pre-trend statistic exceeds the threshold, $\max_{e<0}\lvert\beta_{g,\mathrm{pre}}(e)\rvert>c$,
and undetectable if it is flat in the pre-period, $\max_{e<0}\lvert\beta_{g,\mathrm{pre}}(e)\rvert\le c$,
yet violates parallel trends in the post-period. As per-cohort information grows, a detectable cohort's
admission probability tends to zero and its contribution to the bias vanishes, whereas an undetectable
cohort's admission probability tends to one and its contribution persists at every sample size. Hence
$\hat\theta_S$ is consistent for $\theta_S$ if and only if every confounded cohort is detectable, and
otherwise retains an asymptotic bias equal to the residual violation of the undetectable cohorts, which is
the identification gap that the honest inference of the next subsection bounds.
\end{proposition}

\subsection{Honest inference on the selected set}

Selecting credible cohorts does not by itself deliver parallel trends, since a cohort with a flat
pre-trend may still have a nonzero post-treatment differential trend, and selection removes gross
offenders but not subtle residual violations. We therefore accompany the point estimate with the
sensitivity analysis of \citet{rambachan2023}, applied to the selected-cohort aggregate, imposing a
restriction $\delta_{S,\mathrm{post}}\in\Delta$ on the residual differential trend and reporting the
implied confidence set for $\theta_S$. Matching the flatness criterion of the screen, we adopt the level
bound $\Delta^{\mathrm{Level}}(M)=\{\delta:\lvert\delta_{\mathrm{post}}(e)\rvert\le M\ \forall e\}$, which
asks how far the treated-comparison difference could have wandered from its flat pre-treatment level,
absent treatment. The estimator is then the raw reweighted post-treatment average, with no extrapolation,
and the associated fixed-length confidence interval (FLCI) is near-optimal \citep{armstrong2018}. Because
$M$ is measured in the outcome's own units, the breakdown value $M^\ast$ reads directly as the size
of post-treatment parallel-trends violation that would just overturn the stated conclusion.

Validity depends on the selection regime. Under ex-ante selection $S$ is independent of $\hat\beta$, so
the uniform validity of \citet{rambachan2023} applies verbatim. Under data-driven selection $S$ is
random, and two issues arise. The first is post-selection inference. When cohorts are well separated
from the threshold, selection is asymptotically harmless.

\begin{proposition}[Selection consistency and oracle coverage]\label{prop:oracle}
Suppose the statistic on which selection is based is consistent for its population counterpart, that
the following separation condition holds, that there exists $\epsilon>0$ such that, for every cohort, the
population value of that statistic lies at distance at least $\epsilon$ from the threshold $c$, and that the
population credible set's residual violation satisfies the maintained restriction
$\delta_{S^\ast,\mathrm{post}}\in\Delta^{\mathrm{Level}}(M)$. Then $P(\hat S = S^\ast)\to 1$, where
$S^\ast$ denotes the population credible set, and the FLCI computed on the estimated set $\hat S$ has the
same asymptotic coverage as the oracle interval computed on $S^\ast$. Selection consistency delivers this
equivalence to the oracle interval, and the maintained restriction delivers the oracle interval's nominal
coverage of the causal target, so the two together give nominal coverage.
\end{proposition}

\begin{proof}
By consistency of the selection statistic together with the separation condition, the probability that
any given cohort is misclassified tends to zero. Since the number of cohorts is finite, a union bound
gives $P(\hat S\neq S^\ast)\to 0$. On the complementary event the estimator, its covariance estimate and
the resulting interval coincide exactly with their oracle counterparts computed on the non-random set
$S^\ast$. The two sequences therefore share a limiting distribution, and the coverage of the interval
computed on $\hat S$ converges to that of the oracle interval, which is nominal by the validity of
\citet{rambachan2023} applied to a fixed set whose residual violation lies in $\Delta^{\mathrm{Level}}(M)$.
\end{proof}

Two caveats attach to Proposition~\ref{prop:oracle}. First, the separation condition is untestable, and
is nearly the parallel trends question in disguise, since to assert that every cohort is plainly credible
or plainly not is close to asserting one already knows which cohorts are credible, so the result
relocates the identifying content rather than supplying it. Second, the guarantee is pointwise, not
uniform. \citet{leeb2005} show that a post-selection estimator's distribution cannot be estimated uniformly
consistently, so naive post-selection intervals lack uniform validity \citep{leeb2008}. The
obstruction is confined to local-to-threshold configurations, where the selection statistic drifts
toward the boundary at the sampling rate. Away from the boundary, selection is asymptotically
deterministic and inference behaves conventionally. A researcher unwilling to assume separation can
sidestep both caveats by selecting on information independent of the estimation errors, such as
pre-determined covariates, a sample split, or randomized (data-carving) selection \citep{fithian2014}.
These buy validity through independence between selection and estimation rather than correct selection,
at some cost in efficiency.

The second issue is that the bound $M$ must accommodate a random selected set. The FLCI at $M$ is valid
only if the selected set's residual violation lies in $\Delta^{\mathrm{Level}}(M)$. Since that violation
is itself random, unconditional coverage requires $M$ to dominate it with high probability.

\begin{lemma}[Calibration to the random selected set]\label{lem:calib}
Under data-driven selection the selected aggregate's residual violation is random, and the FLCI's
unconditional coverage is governed by the joint distribution of that violation, the selection event, and
the sampling error. When the violation's dispersion is non-negligible relative to the FLCI's sampling
slack, setting $M$ to the mean residual violation undercovers, because the realized violation exceeds $M$
in a nontrivial fraction of samples, and restoring nominal unconditional coverage requires $M$ to exceed
the mean and approach an upper quantile of the residual-violation distribution, with the sampling slack
providing a partial buffer. When the slack instead dominates the dispersion the buffer alone can suffice.
\end{lemma}

The practical implication, in the regime where the violation's spread is the binding consideration, is that
$M$ should be calibrated to the residual violation of a typical selected set, not to its mean.

Finally, we caution against one otherwise-natural choice of $\Delta$. The relative-magnitudes
restriction $\Delta^{RM}(\bar M)$ of \citet{rambachan2023}, following \citet{manski2018}, bounds
post-treatment violations by $\bar M$ times the observed maximal pre-treatment violation, so its
identified-set half-width is proportional to the observed pre-trend magnitude. When pre-trends are
uninformative about post-treatment violations, the observed pre-trends are small while the true violation
need not be, so the restriction produces intervals whose width tends to zero even as the bias does not,
yielding undercoverage precisely in the regime where our method's value is in question. A fixed
restriction such as $\Delta^{\mathrm{Level}}(M)$, whose bound is set by the researcher rather than read
off the observed pre-trends, does not share this defect.

\section{Simulation study}\label{sec:sim}

\subsection{Design}\label{sec:design}

We study the method in two environments. In the reduced-form model we draw event-study
coefficients directly from their asymptotic distribution, $\hat\beta\sim\mathcal{N}(\tau+\delta,\Sigma)$
with $\Sigma$ known, which isolates the identification and inference mechanics in the setting where the
underlying theory is exact. In the panel model we generate individual outcomes for a staggered
design and estimate both the coefficients and their covariance from the data, which confirms that the
conclusions survive realistic estimation. Unless stated otherwise, results are averages over eight to
twenty thousand replications.

The cohort structure follows a common template. Some cohorts are clean, with $\delta_g=0$, and the
remainder confounded, with a nonzero differential trend, and Table~\ref{tab:dgp} gives the count and
clean-confounded split used in each exercise. Every cohort is given the same true treatment effect,
normalized to one. This normalization is deliberate, making
the true ATT and the true LATT both equal to one, so that any departure of an estimator from one is
attributable to a violation of parallel trends rather than to treatment-effect heterogeneity. It
therefore isolates the identification problem that the paper is about.

Throughout, each cohort carries a four-period pre-treatment and four-period post-treatment event study
with the reference period normalized to zero. A confounded cohort carries a level violation, its
treated-comparison difference shifting by $V>0$ in the post-period and by $\phi V$ in the pre-period, for
a scalar $\phi\ge 0$ we call the informativeness parameter and that is the master axis of the study. When
$\phi=0$ the confound is flat before treatment and no pre-trend rule can detect it, and as $\phi$ grows
the pre-period shift announces it and it becomes detectable. Because $V$ is the post-treatment violation it
is directly comparable to the level bound $M$ of $\Delta^{\mathrm{Level}}(M)$. In the controlled coverage
exercise below the aggregate being tested carries this violation exactly, which makes the boundary
statement ``covers if and only if $M\ge V$'' checkable. More generally, coverage requires only that $M$
dominate the selected aggregate's violation, which for a set mixing clean and confounded cohorts is weakly
below the largest single-cohort $V$, so $M\ge V$ is sufficient and is necessary only in the worst case of an
all-confounded set. This also makes $\phi$ the operational version of the informativeness of pre-trends
discussed in Section~\ref{sec:approach}.

The selection rule is the flatness screen, posed in the same currency as the honest inference, so cohort
$g$ is retained if its estimated pre-trend is close to flat, $\max_e\lvert\hat\beta_{g,\mathrm{pre}}(e)\rvert\le c$.
The estimator of the ATT averages the post-treatment effect over all cohorts, as a heterogeneity-robust
estimator would, and the LATT averages it over the selected cohorts only. Because the screen reads the selected
set off the same noisy pre-trends, the selection is genuinely data-dependent
in the sense of Section~\ref{sec:approach}, and the pre-test bias of Proposition~\ref{prop:pretest} is
operative.

\begin{table}[t]
\centering
\caption{Simulation parameters. Every exercise uses a four-period pre and four-period post event study with
reference period zero, equal cohort weights, and true effect one for every cohort. The reduced-form model
draws coefficients directly, $\hat\beta\sim\mathcal{N}(\tau+\delta,\Sigma)$ with within-cohort AR(1)
covariance $\Sigma_{ij}=s^2\rho^{|i-j|}$. The panel model draws individual untreated outcomes
$Y_{it}(0)=\delta_g(e)+\epsilon_{it}$, $\epsilon\sim\mathcal{N}(0,\sigma_\epsilon^2)$, for $500$ treated
units per cohort and $1500$ shared controls over $T=14$ periods, and estimates the coefficients and their
covariance from the data.}
\label{tab:dgp}
\begin{tabular}{lcccc}
\toprule
& Estimand & Scope & Honest cov.\ & Panel \\
& (Table~\ref{tab:tier1}) & (Fig.~\ref{fig:scope}) & (Table~\ref{tab:flci}, Fig.~\ref{fig:layer2}) & (Sec.~\ref{sec:calib}) \\
\midrule
Cohorts (clean/confounded) & $6$ ($4/2$) & $12$ ($6/6$) & --- & $6$ ($3/3$) \\
Post-treatment violation $V$ & $0.8$ & $0.6$ & $0.3,\ 0.6$ & $0.6$ \\
Informativeness $\phi$ & $1.0$ & $[0,1.5]$ & --- & $[0,1.5]$ \\
Screen threshold $c$ & $0.40$ & $0.40$ & --- & $0.40$ \\
Noise & $s{=}0.03\text{--}0.40$ & $\sigma{=}0.20$ & $\sigma_{\mathrm{agg}}{=}0.10$ & $\sigma_\epsilon{=}0.5$ \\
Pre/post correlation $\rho$ & $0.5$ & $0.5$ & --- & estimated \\
Replications & $20{,}000$ & $8{,}000$ & $4{,}000$ & $2000\text{--}3000$ \\
\bottomrule
\end{tabular}
\end{table}

\subsection{Recovering a clean effect where the ATT is biased}

We first illustrate Proposition~\ref{prop:id} in the simplest configuration, with a subset of cohorts
violating parallel trends in the post-period. Table~\ref{tab:tier1} reports the result. The estimator
of the ATT is biased by $+0.267$, since it averages over all cohorts, so the offending cohorts'
differential trends enter it directly. The LATT lands at $1.000$, on the truth, because the selection
removes those cohorts from the average and, by Proposition~\ref{prop:id}, their violations then do not
enter the estimand at all.

The lower panel of the table sweeps the sampling noise $s$, and is the more informative comparison.
As precision improves the LATT's bias falls monotonically from $+0.013$ to zero, and this residual is the
pre-test bias of Proposition~\ref{prop:pretest}, which vanishes as the selection rule becomes reliable.
Two sources could in principle generate it, a noisy screen that admits genuinely confounded cohorts, which
would persist even under independent selection, and the correlation between a cohort's pre- and
post-treatment sampling errors, which distorts the retained cohorts' estimates only under same-sample
selection. The final column isolates the two by selecting on an independent pre-trend draw. It tracks the
full-sample column almost exactly, so the residual is the first source, admission of confounded cohorts,
and the second is negligible here because the two-sided flatness screen conditions the retained estimates
symmetrically. The ATT estimator's bias, by contrast, is frozen at $+0.267$ regardless of precision. The
contrast is the paper's basic point in miniature, the LATT facing a noise problem, which more data cures,
whereas the ATT estimator faces a wrong-target problem, which no amount of data cures.

\begin{table}[t]
\centering
\caption{The estimand demonstration (reduced-form model). True effect $=1$ for all cohorts. In the lower
panel LATT (full) selects on the same pre-trends used for estimation, while LATT (split) selects on an
independent draw. Their near-equality shows the residual bias is admission of confounded cohorts, not
same-sample conditioning.}
\label{tab:tier1}
\begin{tabular}{lccc}
\toprule
& ATT & LATT & LATT \\
& & (full) & (split) \\
\midrule
Point estimate & $1.267$ & $1.000$ & \\
Bias vs.\ truth & $+0.267$ & $+0.000$ & \\
\midrule
\multicolumn{4}{l}{Residual LATT bias as sampling noise $s$ falls}\\
$s=0.40$ & $+0.265$ & $+0.013$ & $+0.014$ \\
$s=0.20$ & $+0.267$ & $-0.001$ & $-0.000$ \\
$s=0.12$ & $+0.267$ & $-0.000$ & $-0.000$ \\
$s=0.03$ & $+0.267$ & $+0.000$ & $+0.000$ \\
\bottomrule
\end{tabular}
\end{table}

\subsection{The scope condition}

Figure~\ref{fig:scope} traces the method across the informativeness axis $\phi$ described in
Section~\ref{sec:design}, and its four panels should be read together.

The upper-left panel plots the bias of the two point estimators against the truth. The bias of the ATT
estimator is a flat line at $+0.30$, since it never consults pre-trends, so its performance cannot depend
on how informative they are. The LATT's bias begins on top of that line at $\phi=0$ and falls to
zero as pre-trends become informative. The vertical gap between the two curves is therefore exactly the
method's advantage, and reading it off from left to right gives the scope condition, with the advantage
nil when pre-trends carry no information about post-treatment violations and maximal when they carry a
great deal.

That the two curves coincide at the left edge is not a defect but an honest reflection of
the fact, argued in Section~\ref{sec:approach}, that selecting on uninformative pre-trends buys nothing.

The upper-right panel explains why the LATT's bias falls, by plotting the identification gap of
the selected set, the average post-treatment violation among the cohorts that survive selection. It
tracks the LATT's bias almost exactly, falling from $0.30$ to essentially zero. This confirms that the
LATT's remaining bias is not an artifact of the estimator but simply the residual violation that
selection has failed to remove. The lower-right panel gives the mechanism behind both, the average
number of confounded cohorts that slip through the screen falling from $5.09$ of six when the confound is
invisible to essentially zero when it is plainly visible.

The lower-left panel turns to inference and reports the coverage of a point-based confidence interval
for the causal target, computed both in the naive way and by sample-splitting. Both collapse at low
informativeness and recover only at high informativeness. The comparison is diagnostic, since
sample-splitting removes any contamination from data-dependent selection by construction, so the fact
that it collapses
just as badly establishes that the failure is not selection distortion but the identification gap
of the upper-right panel, since the surviving cohorts genuinely still violate parallel trends. This is what
motivates the sensitivity bounds of the next subsection, since no purely inferential fix can repair a
violation of the identifying assumption.

(At the right edge the naive interval slightly outperforms the split one. This is simply because
splitting discards half the data, and by that point selection is nearly deterministic, so there is no
contamination left to correct.)

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{master_axis.png}
\caption{Scope condition. Everything is governed by pre-trend informativeness. The ATT estimator's
bias is constant, and the LATT's advantage is the gap between the lines. Causal-coverage collapse at low
informativeness is an identification gap (surviving cohorts still violate parallel trends), not
selection distortion, since sample-splitting collapses just as badly.}
\label{fig:scope}
\end{figure}

\subsection{Honest inference restores coverage}\label{sec:flci}

We now validate the level-bound FLCI of Section~\ref{sec:approach} applied to the selected-cohort
aggregate, using the level violation of Section~\ref{sec:design} whose size $V$ is directly comparable
to the assumed bound $M$.

Table~\ref{tab:flci} is the validity check. Read down the rows. When the assumed bound equals the true
violation ($V=M$) the FLCI covers at $95.0\%$, that is, at its nominal level. When the bound is generous
($M>V$) it covers conservatively at $100\%$. When the bound is violated ($M<V$) coverage collapses. This
is precisely the behavior an honest sensitivity interval should display. It is valid exactly on the range
of violations it claims to accommodate, and it fails visibly outside it, rather than failing silently.

The final column records that the naive point interval collapses to near-zero coverage the instant any
violation is present, because it is centered on a biased estimate with a width reflecting sampling noise
alone.

Figure~\ref{fig:layer2} repeats the exercise across a continuum of violation sizes and adds the practical
output. In the left panel the point interval falls away rapidly as the violation grows, while each FLCI
holds its nominal level until the true violation reaches its own assumed bound, marked by the vertical
dotted lines, and declines thereafter. The middle panel records the price. Interval half-width is
governed by the assumed $M$ rather than by the realized violation, so honesty about a wider class of
violations is paid for in width, immediately and proportionately.

The right panel is what a practitioner would actually report. Holding the data fixed, it traces the
confidence band as the assumed bound is relaxed and identifies the breakdown value $M^\ast$ at which the
band first admits the truth. Because $M$ is in the outcome's own units, this quantity answers the
question a reader of a DiD paper wants answered directly, how large a post-treatment parallel-trends
violation one must be willing to countenance before the stated conclusion no longer follows. And because
the bound is fixed by the researcher rather than read off the pre-trends, the exercise remains
informative even when pre-trends are uninformative, in direct contrast to the relative-magnitudes
restriction, whose bounds would shrink toward zero in exactly that case and deliver false precision.

\begin{table}[t]
\centering
\caption{FLCI coverage of the causal target (reduced-form model), for an aggregate whose violation equals
$V$, valid iff $M\ge V$.}
\label{tab:flci}
\begin{tabular}{cccc}
\toprule
True violation $V$ & Assumed $M$ & FLCI coverage & Point-estimate coverage \\
\midrule
$0.30$ & $0.30\ (=V)$ & $95.0\%$ & $15.9\%$ \\
$0.30$ & $0.60\ (>V)$ & $100\%$ & $14.9\%$ \\
$0.60$ & $0.30\ (<V)$ & $9.3\%$ & $0.0\%$ \\
$0.60$ & $0.60\ (=V)$ & $95.0\%$ & $0.0\%$ \\
\bottomrule
\end{tabular}
\end{table}

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{layer2_full.png}
\caption{Sensitivity bounds restore honest coverage. The point estimate collapses as the violation grows.
The FLCI is honest exactly while $M\ge V$. Interval width is the price of honesty, set by $M$. The right
panel is the sensitivity/breakdown curve, with $M^\ast$ in outcome units.}
\label{fig:layer2}
\end{figure}

\subsection{Selection, calibration, and panel confirmation}\label{sec:calib}

Two questions remain for the data-driven regime, whether choosing the selected set from the same data
distorts the sensitivity interval, and how to choose $M$ when the selected set, and hence its residual
violation, is random. On the first, coverage against $M$ for the full-data and sample-splitting
procedures crosses the nominal level at nearly the same bound ($M\approx0.048$ and $0.056$), so
data-dependent selection introduces no gross distortion and splitting buys little, consistent with the
pointwise reading of \citet{leeb2005}. On the second, the residual violation of the selected aggregate is
not a constant but a spread, multi-modal random variable across replications, running from zero to about
$0.06$ with a thin tail to $0.11$, each mode a different number of borderline cohorts surviving the
screen, which is the content of Lemma~\ref{lem:calib}. A fixed $M$ must dominate it, and setting $M$ at
the mean residual violation ($0.043$) undercovers, while nominal coverage is reached near the
ninety-fifth percentile ($0.057$) and well below the maximum ($0.112$). The recommendation follows, to
calibrate $M$ to the residual violation of a typical selected set, not its average, and treat it as a
floor.

Finally, we repeat the estimand demonstration on generated panel data, where individual outcomes are
simulated for a staggered design with a level confound and idiosyncratic noise, cohort event-study
coefficients are estimated from genuine difference-in-differences contrasts against a never-treated
group, and their covariance is estimated rather than supplied. The conclusions are unchanged, since
across the informativeness axis the ATT estimator's bias stays flat near $+0.30$ while the LATT's falls
toward zero, reproducing Figure~\ref{fig:scope} with estimated quantities, and the estimated standard
error matches its Monte Carlo counterpart where selection is effectively deterministic. One discrepancy
is informative. At intermediate informativeness, where selection is stochastic, the estimated standard
error understates its Monte Carlo counterpart ($0.016$ against $0.024$ at $\phi=0.75$), the missing
component being the variance induced by the randomness of selection, which a standard error conditional
on the realized set does not capture. This confirms, from the panel side, that conditioning on a
data-dependent selection understates uncertainty, and why the honest object to report is the sensitivity
interval of Section~\ref{sec:flci} rather than a conventional confidence interval.

\section{Empirical application}\label{sec:app}

We illustrate the method with the local effect of the shale boom on house prices, a staggered design
with many counties per cohort. The treatment date is the year a county's oil and gas production first
rises sharply, from county production records \citep{ers2015}, which sorts the boom counties into
onset cohorts, and the outcome is the log county house-price index \citep{bogin2019}, over 1998--2019. We
estimate \citet{callaway2021} group-time effects against never-treated counties, those with negligible
production, screen cohorts on the flatness of their estimated pre-trend, and report the level bound of
Section~\ref{sec:approach} for the selected aggregate.

Onset is gradual and partly endogenous, which is what makes the setting instructive. A county's
production ramps up over several years, and where it does so during a regional housing upswing its
prices are already climbing before the boom is dated. Figure~\ref{fig:frack}(a) shows the result. The
early cohorts have flat pre-trends and are retained, while the later cohorts, whose onset coincides
with the mid-2000s house-price run-up, carry pre-treatment trends several times larger and are dropped.

The two aggregates part ways (Figure~\ref{fig:frack}b). Pooled across all cohorts, the event study
climbs through the onset date to a house-price effect of $+6.7$ log points ($t=8$), the kind of estimate
a practitioner would report as a clear positive effect of the boom. But the pooled path is already
trending before treatment, and the effect is inherited from the dropped cohorts. Restricted to the
credible subpopulation, the three cohorts whose pre-trends are essentially flat (the screen at
$c\approx0.03$), the event study is flat both before and after onset, and the LATT is $-0.2$ log points,
indistinguishable from zero. The gap between the two, $+6.9$ log points, is itself sharply
estimated ($t=5$). Honest inference confirms the reading rather than resting on the point estimate.
Applying the level bound of Section~\ref{sec:approach} to the credible aggregate at $M$ equal to the screen
tolerance $c\approx0.03$, the largest residual post-treatment violation the screen leaves room for, the
LATT's fixed-length interval runs from $-5.9$ to $+5.5$ log points, which contains zero and excludes the
pooled $+6.7$. The breakdown value is $M^\ast\approx4.2$ log points, so reconciling the credible near-zero
reading with the pooled effect would require countenancing a post-treatment violation about $1.4$ times the
screen tolerance and larger than any retained cohort's own pre-trend. The near-zero reading is not an artifact of a
hand-picked threshold. Panel (c) traces
the estimate along the credibility path of Section~\ref{sec:select}, and it stays flat and close to zero
across the whole range of genuinely flat pre-trends, cohorts with $\max_e|\hat\beta_{g,\mathrm{pre}}(e)|$
below about $0.03$, and climbs toward the pooled $+6.7$ only once the threshold is loosened enough to
readmit cohorts whose own pre-trends, around $0.07$--$0.08$, are as large as the effect in question.
Reading the LATT at that lax end simply rebuilds the pooled estimate out of its least credible cohorts,
and the positive figure survives only by abandoning the credibility standard, not by tolerating a
residual violation of the credible set.

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{fracking_diagnostic.png}
\caption{The shale boom and county house prices. (a) Cohort credibility map, the maximum pre-treatment
violation by onset cohort with the flatness screen, where the later, trending cohorts are dropped. (b)
Event study, all cohorts (pooled estimate) versus the credible subpopulation (LATT) at the illustrative
threshold $c\approx0.03$, with $95\%$ bands. (c) The credibility path, the LATT as the screen threshold
$c$ is relaxed, near zero across the flat-pre-trend cohorts and rising toward the pooled estimate only as
trending cohorts are readmitted.}
\label{fig:frack}
\end{figure}

The reading is that the pooled effect reflects a pre-existing trend in the counties that adopt late rather
than a credibly causal effect, and a near-zero net effect for the credible subpopulation is what the
offsetting amenity and disamenity channels of local energy development would predict
\citep{muehlenbachs2015}. The method identifies the effect for the credible cohorts alone, and the excess
in the dropped cohorts blends their steep pre-trends with any genuine difference in their treatment
effects, a mixture the design cannot separate and need not, since the credible reading stands on the
retained cohorts. The application thus shows the method working in the direction opposite to
point-identification, not recovering an effect the pooled estimate misses but withdrawing one it spuriously
reports, and locating that withdrawal exactly in the cohorts the screen flags.

\section{Conclusion}\label{sec:conc}

We have argued that when parallel trends fails for some treated cohorts, a productive response is to
change the estimand rather than to defend or bound the ATT. The credible-subpopulation LATT is
point-identified under parallel trends for the selected cohorts alone, a condition that can hold when the
ATT's fails, is estimated by a transparent reweighting of standard group-time effects, and can be
accompanied by honest sensitivity bounds on the residual violation that the pre-trend screen cannot rule
out. This is a specific instance of the general principle, familiar from instrumental
variables and limited overlap, of retreating to a credibly-identified subpopulation.

The method delivers a sharp answer where the ATT estimator is biased and its sensitivity set is wide,
at the price of a narrower, subpopulation target whose composition must be described rather than
assumed. Its advantage is contingent on pre-trends being informative about post-treatment violations.
Where they are not, only economic context can carry the assumption. Under data-driven selection the
estimator carries a pre-test bias that vanishes as pre-trends grow informative and a pointwise, rather than
uniform, guarantee, which a researcher can trade for uniform validity by selecting on independent
information. The principal open
question is a decision-theoretic characterization of the region in which the point-identified LATT
dominates the set-valued ATT, which would turn the scope condition documented here into a formal
recommendation.

\clearpage
\bibliographystyle{aer}
\bibliography{refs}