EconBase
← Back to paper

Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,132 characters · 31 sections · 40 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 422 chars of source]

\footnotetext[1]{I am extremely grateful to Jonathan Roth for detailed comments that improved the paper. I thank Scott Cunningham for preparing and making publicly available the cleaned Medicaid expansion panel used in the empirical application; the underlying microdata are from the American Community Survey. All errors are my own.}

abstractIn difference-in-differences designs, the parallel trends assumption requires that the outcome gap between treated and control units would have remained flat absent treatment. Pre-treatment event studies frequently reject this flat-gap requirement. Existing responses include parametric trend controls and bounds on the treatment effect under assumptions about the magnitude of the violation. This paper shows that point identification of cohort-specific and aggregate treatment effects in staggered designs remains achievable under strictly weaker assumptions. I replace the flat-gap requirement with a hierarchy of higher-order conditions, Parallel[$p$], embed this framework in the group-time average treatment effect structure of callaway2021difference, and prove an aggregation theorem for the case where different cohorts are identified under different feasible polynomial orders, a challenge unique to staggered designs that has not been previously addressed (Theorem (ref)). A sequential order-selection procedure guides applied practice. Monte Carlo evidence confirms that post-selection bootstrap coverage remains near-nominal and that inference is robust to realistic serial correlation. Applied to Medicaid expansion data, the method yields point estimates resting on an assumption the pre-treatment data do not reject, in contrast to the flat-gap requirement which those same data decisively reject. \noindentKeywords: difference-in-differences, staggered adoption, parallel trends, higher-order parallelism, group-time ATT, Medicaid expansion. \noindentJEL codes: C21, C23.

Introduction

Difference-in-differences (DiD) is among the most widely used designs for program and policy evaluation. In staggered adoption settings, recent work has shown that the conventional two-way fixed effects regression can deliver misleading estimates when treatment effects vary across cohorts, and has proposed alternative estimators that recover interpretable average treatment effects callaway2021difference,sun2021estimating,goodman2021difference,de2020two,borusyak2024revisiting. All of these frameworks, however, rest on the same core identifying assumption: parallel trends.

Parallel trends requires that, absent treatment, the average outcome gap between treated and control units would have remained constant over time. In practice, pre-treatment event studies frequently show systematic trends that are difficult to reconcile with this requirement, especially when pre-periods are long. When this happens, treatment effect estimates are biased, and the magnitude of that bias is generally unknown. The common response is to test for pre-trends and proceed only if the test passes, but this practice is itself problematic: pre-trends tests are often underpowered, and conditioning the analysis on having passed one distorts subsequent inference roth2022pretest. A researcher who finds a trending pre-period is left without a clear way forward.

This paper asks: when parallel trends fails in a staggered adoption design, can we still point-identify average treatment effects without discarding some cohorts or bounding the effect under assumptions about the magnitude of the violation? I propose a middle path. Standard parallel trends implicitly assumes that whatever shape the pre-treatment gap takes, it would have become flat after treatment. I instead use the information in the pre-treatment trajectory: when the gap evolves smoothly, that evolution is informative about what it would have done absent treatment. I replace Parallel[1] with a hierarchy of higher-order assumptions, Parallel[$p$], and project the pre-treatment trajectory forward as the counterfactual. Where rambachan2023more trace out an identified set indexed by the strength of a smoothness restriction, I instead select an order from the pre-treatment data and report a point estimate under that order.

Formally, Parallel[$p$] requires only that the $p$-th order time difference of the untreated outcome gap is common across groups. Parallel[1] is the standard assumption: the gap must be flat. Parallel[2] allows the gap to trend linearly; Parallel[3] allows a quadratic trajectory; higher orders allow richer curvature. Each step up in $p$ is strictly weaker than the previous one and remains sufficient for point identification of group-time average treatment effects. When a cohort has more than $p$ pre-treatment periods, Parallel[$p$] generates testable implications about the pre-treatment gap that can be assessed with observable data.

These implications do not, however, eliminate extrapolation. Every DiD estimator relies on an extrapolation assumption: Parallel[1] extrapolates a flat gap into the post-treatment period, while Parallel[$p$] extrapolates the pre-existing polynomial trajectory. Neither is testable in the post-treatment period. The advantage of Parallel[$p$] is that its key implication, that the pre-treatment gap lies on a low-degree polynomial, is testable with observed data, whereas the flatness required by Parallel[1] is often visibly violated. Parallel[$p$] thus replaces an assumption the data may reject with one they do not, recovering point identification where standard methods cannot proceed.

The paper makes three contributions. First, I formally identify $\mathrm{ATT}(g,t)$ under Parallel[$p$] for arbitrary $p \geq 1$ in the staggered adoption framework of callaway2021difference, using the polynomial structure of the pre-treatment gap to construct a cohort-specific counterfactual in each post-treatment period.

Second, I address a problem that arises only in staggered settings. Because treatment timing varies, cohorts differ in how many pre-treatment periods they have, and therefore in the maximum order each can support: a cohort treated two periods into the panel can support only Parallel[1], while one treated ten periods in can support up to Parallel[9]. The natural question is how to combine cohorts that are identified under different orders into a single summary effect. Theorem (ref) resolves this, showing that coherent weighted averages of group-time ATTs remain identified even when each cohort $g$ is identified under its own order $p_g$. To my knowledge, this is the first aggregation result for staggered DiD under cohort-heterogeneous orders.

Third, I develop a sequential order-selection procedure tailored to staggered settings, provide Monte Carlo evidence on finite-sample performance and post-selection inference, and implement the method in a companion Stata command anddp, which produces the aggregate DD[$p$] ATT with cluster bootstrap inference. The selection procedure responds directly to the pre-testing critique: rather than testing the flat-gap null and conditioning on passing it, it selects the lowest order the pre-treatment data do not reject and estimates under that order. The Monte Carlo evidence shows that bootstrap confidence intervals retain near-nominal coverage after this selection step.

I apply the method to the staggered adoption of Medicaid expansion under the Affordable Care Act, using state-level insurance coverage data. Expanding and non-expanding states were on visibly diverging insurance trajectories before any expansion took effect, and a joint pre-trends test rejects flatness decisively. Standard DiD, which assumes the gap would have stayed flat, is therefore hard to defend here. The proposed estimator, DD[$p$], nonetheless recovers point estimates under Parallel[$p$], an assumption the same pre-treatment data do not reject, and yields: expansion increased insurance coverage by roughly six percentage points.

The higher-order parallelism idea builds on mora2019alternative, who develop the Parallel[$p$] hierarchy and prove identification in a two-group, multiple-period setting. This paper extends their framework to staggered adoption. egami2023using who adapt the Mora--Reggio idea to staggered designs using a generalised method of moments (GMM) estimator combining Parallel[1] and Parallel[2]. I differ in three respects. First, I work with the full hierarchy for arbitrary $p \geq 1$, not only $p \in \{1,2\}$. Second, I select a single order per cohort by a sequential test and report which was used, rather than combining fixed orders. Third, because cohorts may be identified under different selected orders, aggregating them raises a problem that does not arise under their fixed-order combination, the problem Theorem (ref) resolves.

Prior responses to parallel trends violations include the following. Sensitivity bounds rambachan2023more restrict how far the post-treatment violation of parallel trends can depart from the pre-treatment trend, and report how the identified set for the treatment effect widens as that restriction is relaxed; the approach is rigorous and transparent, and rather than committing to a single counterfactual it traces out a range of estimates indexed by the strength of the assumption. The closest connection is to their smoothness restriction: assuming the post-treatment violation is exactly linear coincides with Parallel[2] in the present hierarchy. The approaches diverge in what they do with that restriction. They treat linearity as one end of a continuum and report how the identified set grows as it is relaxed, whereas I select an order from the data and report a point estimate under it. Pre-trend tests roth2022pretest diagnose violations but do not provide a corrected estimator. A complementary approach is the non-inferiority framework of bilinski2026nothing, which recasts pre-trend assessment as an equivalence test: rather than asking whether pre-trends are exactly zero, it asks whether deviations are small enough not to matter for the estimated effect. Their procedure operates under Parallel[1] and provides a decision rule for tolerating small violations. roth2023s provide a comprehensive review of these and other responses to parallel trends violations.

A related practice in applied work is to include unit-specific linear time trends in TWFE regressions to absorb differential pre-trends angrist2009mostly. dobkin2018economic adopt a conceptually related strategy in a single-timing event study: they include a linear trend in event time alongside a saturated set of post-treatment dummies. Because the post-treatment dummies absorb all post-treatment variation, the slope coefficient is identified from pre-treatment data only. This is formally equivalent to the $p=2$ case of the present framework applied to a single cohort, where no aggregation is required. Dobkin's linear extrapolation is therefore a single one-step relaxation of flat parallel trends. DD[$p$], the estimator I develop under Parallel[$p$], generalises this in three ways: it extends to arbitrary order $p \geq 1$ with a sequential procedure to select the order the data support; it handles staggered adoption with multiple cohorts; and it provides formal aggregation across cohorts identified under different feasible orders, together with asymptotic inference (see Appendix (ref)).

Other approaches to parallel-trends failures include synthetic control methods abadie2010synthetic,xu2017generalized,arkhangelsky2021synthetic,ben2022synthetic, which construct a weighted comparison group that matches treated units' pre-treatment trajectories and work well with long pre-treatment series and relatively few treated units. Triple-differences designs strezhnev2023decomposing,ortiz2025better require a placebo stratum known to be unaffected by treatment. Change-in-changes athey2006identification relies on rank preservation. Partial identification approaches manski2018right deliver bounds on treatment effects. Each of these methods is suited to a different empirical context; the present approach targets the common setting where the pre-treatment gap evolves smoothly and a polynomial counterfactual is credible.

The remainder of the paper is organised as follows. Section (ref) establishes the framework. Section (ref) introduces the hierarchy of higher-order parallelism assumptions. Section (ref) derives the identification results, including the aggregation theorem. Section (ref) presents estimation and inference. Section (ref) develops the order-selection procedure. Section (ref) reports Monte Carlo evidence. Section (ref) presents the Medicaid expansion application. Section (ref) concludes.

Setup

Panel, Treatment, and Notation

The notation and staggered-adoption setup follow callaway2021difference. Consider a balanced panel of $N$ units over $T$ periods $t = 1,\ldots,T$. Unit $i$ has adoption date $G_i \in \{g_1,\ldots,g_K,\infty\}$. $G_i = g$ means unit $i$ is first treated in period $g$; $G_i = \infty$ means never treated. Treatment is absorbing: once a unit is treated it remains treated in all subsequent periods, so a unit's adoption date $G_i$ fully summarises its treatment path. Units with $G_i = g$ form cohort $g$. Let $N_g$ be the size of cohort $g$ and $N_\infty$ the number of never-treated units, with $N = \sum_g N_g + N_\infty$.

Asymptotics are large-$N$ with $T$ fixed. The cohort fractions $\pi_g = N_g/N$ are held fixed as $N$ grows, which is achieved by requiring each $N_g$ to grow proportionally with $N$, a standard large-sample device that prevents any cohort from vanishing or dominating in the limit.

The number of pre-treatment periods for cohort $g$ is:

equation[equation omitted — 51 chars of source]

where $t_{\min}$ is the first observed period.

Potential Outcomes

For each unit $i$ and period $t$, let $Y_{it}(g)$ denote the potential outcome if first treated in period $g$, and $Y_{it}(\infty)$ the never-treated potential outcome neyman1923application,rubin1974estimating. The observed outcome is:

equation[equation omitted — 118 chars of source]

Before treatment, observed outcomes equal untreated potential outcomes, making pre-treatment data informative about the untreated trajectory.

assumption[No Anticipation] For all $i$, $g$, and $t < g$: $Y_{it}(g) = Y_{it}(\infty)$.

Units do not alter behaviour in anticipation of future treatment, so their pre-treatment observed outcomes equal their untreated potential outcomes.

assumption[Overlap] For each cohort $g$, the probability of belonging to that cohort is strictly positive and strictly less than one: $0 < \Pr(G_i = g) < 1$.

Target Parameters

The group-time average treatment effect is:

equation[equation omitted — 123 chars of source]

The expectation averages over all units $i$ belonging to cohort $g$, comparing their actual post-treatment outcome to what they would have experienced had treatment never occurred.

A scalar summary is the weighted aggregate:

equation[equation omitted — 95 chars of source]

for non-negative weights $w_{g,t}$ summing to one.

Higher-Order Parallel Trends

Standard Parallel Trends

assumption[Parallel{[1]}] For all cohorts $g$ and periods $t$: $\mathbb{E}[\Delta Y_{it}(\infty)\mid G_i=g] = \mathbb{E}[\Delta Y_{it}(\infty)\mid G_i=\infty]$, where $\Delta Y_{it} = Y_{it} - Y_{i,t-1}$.

Standard parallel trends says that, absent treatment, the average untreated outcome would have changed by the same amount each period for every group, treated cohorts and never-treated alike. Equivalently, the gap between any cohort and the never-treated group would have stayed constant over time. This has a testable pre-treatment implication: the pre-treatment gap should be flat. When the event-study plot shows a visible pre-treatment trend, that implication fails, and Parallel[$1$] is implausible.

Higher-Order Differences

The $p$-th order difference operator is defined recursively: $\Delta^{1}Y_{it} = Y_{it} - Y_{i,t-1}$ and $\Delta^{p}Y_{it} = \Delta^{1}(\Delta^{p-1}Y_{it})$ for $p \geq 2$. Written out:

equation[equation omitted — 98 chars of source]

$\Delta^{1}$ measures the year-on-year change; $\Delta^{2}$ measures how that change is itself changing; $\Delta^{3}$ measures the rate of change of $\Delta^{2}$, and so on.

To build intuition, suppose the gap between treated and control groups rises by 0.01 per year before treatment. Parallel[1] requires the gap to be constant, so a steadily rising gap is inconsistent with it. Parallel[2] requires only that this 0.01 annual change is common to the treated and never-treated groups, a weaker restriction that the pre-treatment data may well support.

assumption[Parallel{[$p$]}] For a given integer $p \geq 1$, all cohorts $g$, and all periods $t$: \[ \mathbb{E}[\Delta^{p}Y_{it}(\infty)\mid G_i=g] = \mathbb{E}[\Delta^{p}Y_{it}(\infty)\mid G_i=\infty]. \]

Parallel[1] is the special case $p=1$. Parallel[2] allows different levels and slopes across groups, requiring only that the acceleration is common. Each step up in $p$ is a strictly weaker assumption and requires one additional pre-treatment period to test.

remark[Parallel{[$p$]} does not imply Parallel{[$p-1$]}] The hierarchy is nested in the sense that data satisfying Parallel[$p-1$] also satisfy Parallel[$p$] (a flat function is a special case of a linear function). But the converse does not hold: data with a stable linear pre-trend satisfy Parallel[2] but not Parallel[1]. The appropriate order is determined by the data-generating process and estimated via the sequential test in Section (ref).

Identification in Staggered Settings

Pre-Treatment Gap and Polynomial Structure

Define the observed gap between cohort $g$ and never-treated units in period $t$:

equation[equation omitted — 117 chars of source]

Under Assumption (ref), this equals the untreated potential outcome gap in the pre-treatment period.

lemma[Polynomial Gap Structure] Under Assumptions (ref) and (ref), for all pre-treatment periods $t < g$ with $t \geq t_{\min}+p$: $\Delta^{p}\gamma_{g,t} = 0$. Equivalently, $\gamma_{g,t}$ is a polynomial of degree $p-1$ in $t$ for all $t < g$: \begin{equation} \gamma_{g,t} = c_{g,0} + c_{g,1}t + \cdots + c_{g,p-1}t^{p-1}, \quad t < g. \end{equation} The cohort-specific constants $c_{g,0},\ldots,c_{g,p-1}$ are identified from $p$ pre-treatment gap observations.
proofUnder Assumption (ref), for $t < g$: $\mathbb{E}[Y_{it}\mid G_i=g] = \mathbb{E}[Y_{it}(\infty)\mid G_i=g]$. Therefore: \[ \gamma_{g,t} = \mathbb{E}[Y_{it}(\infty)\mid G_i=g] - \mathbb{E}[Y_{it}(\infty)\mid G_i=\infty]. \] Since $\Delta^{p}$ is a linear operator: \begin{align*} \Delta^{p}\gamma_{g,t} &= \mathbb{E}[\Delta^{p}Y_{it}(\infty)\mid G_i=g] - \mathbb{E}[\Delta^{p}Y_{it}(\infty)\mid G_i=\infty] = 0, \end{align*} where the equality uses Assumption (ref). A sequence on $\mathbb{Z}$ with identically zero $p$-th difference is a polynomial of degree at most $p-1$, the discrete analogue of the fact that a function whose $p$-th derivative is identically zero is a polynomial of degree $p-1$.\footnote{This is a standard result in finite-difference calculus; see, e.g., Jordan (1965, Calculus of Finite Differences) for a classical treatment.} Formally, let $\mathcal{P}_{p-1}$ denote the space of polynomials of degree $\leq p-1$ on $\mathbb{Z}$. The kernel of $\Delta^{p}$ on $\mathbb{Z}$-sequences is exactly $\mathcal{P}_{p-1}$ (this follows by induction: $\Delta^{1}\gamma=0$ iff $\gamma$ is constant; if $\Delta^{p}\gamma=0$ then $\Delta^{p-1}(\Delta\gamma)=0$, so $\Delta\gamma\in\mathcal{P}_{p-2}$, which implies $\gamma\in\mathcal{P}_{p-1}$). The $p$ coefficients $c_{g,0},\ldots,c_{g,p-1}$ are uniquely determined by any $p$ distinct values of $\gamma_{g,t}$ from the pre-treatment period, obtained by solving the resulting linear system or, when $m_g > p$, by OLS.

Lemma (ref) makes the identifying content of Parallel[$p$] concrete and testable: the pre-treatment gap must lie on a polynomial of degree $p-1$. For $p=1$: flat (constant). For $p=2$: straight line. For $p=3$: quadratic curve. The residuals from fitting this polynomial to the pre-treatment data are the testable implications used in Section (ref).

Counterfactual Identification

The counterfactual gap, what the gap would have been in the post-treatment period absent treatment, is:

equation[equation omitted — 156 chars of source]

This is unobservable for $t \geq g$ because treated units observed outcomes include the treatment effect.

proposition[Counterfactual Identification] Under Assumptions (ref) and (ref), with $m_g \geq p$: \begin{enumerate}[label=(\roman*)] • $\gamma_{g,t}(0)$ is a polynomial of degree $p-1$ in $t$ for all periods $t$, including post-treatment. • This polynomial is uniquely identified by the $p$ most recent pre-treatment observations $\{\gamma_{g,g-k}\}_{k=1}^{p}$, which are observable under Assumption (ref) (see Remark (ref) for the finite-sample estimator, which uses all $m_g$ pre-treatment observations for efficiency). • Evaluating the polynomial at any $t \geq g$ gives $\gamma_{g,t}(0)$. \end{enumerate} In particular: for $p=1$, $\gamma_{g,t}(0) = \gamma_{g,g-1}$ (flat at last pre-period value, the standard DiD counterfactual); for $p=2$, $\gamma_{g,t}(0) = (t-g+2)\gamma_{g,g-1}-(t-g+1)\gamma_{g,g-2}$ (linear extrapolation of the pre-existing trend).
proofPart (i): Assumption (ref) states $\Delta^{p}\gamma_{g,t}(0) = 0$ for all $t$, not just pre-treatment periods. By the same argument as Lemma (ref), $\gamma_{g,t}(0) \in \mathcal{P}_{p-1}$ for all $t$. Part (ii): A polynomial of degree $p-1$ is uniquely determined by $p$ values. Under Assumption (ref), $\gamma_{g,g-k} = \mathbb{E}[Y_{i,g-k}(\infty)\mid G_i=g] - \mathbb{E}[Y_{i,g-k}(\infty)\mid G_i=\infty]$ for $k=1,\ldots,p$, which equals $\gamma_{g,g-k}(0)$ (since no treatment has occurred). These $p$ values uniquely pin down the polynomial.\footnote{The coefficient vector $\mathbf{c}_g$ is uniquely determined because the $p \times p$ Vandermonde matrix formed from $p$ distinct integer time indices has full column rank; distinct values of $t$ guarantee non-vanishing Vandermonde determinant.} Part (iii): Evaluate the identified polynomial at $t \geq g$. Since $\gamma_{g,t}(0)$ is a polynomial everywhere and we have identified it from pre-treatment data, its post-treatment values are also identified. For $p=1$: the unique degree-0 polynomial through one point is the constant $\gamma_{g,g-1}(0) = \gamma_{g,g-1}$. For $p=2$: the unique degree-1 polynomial through two points $\{(g-2, \gamma_{g,g-2}), (g-1, \gamma_{g,g-1})\}$ has slope $\gamma_{g,g-1}-\gamma_{g,g-2}$ and evaluates at $t$ to $\gamma_{g,g-2} + (t-g+2)(\gamma_{g,g-1}-\gamma_{g,g-2}) = (t-g+2)\gamma_{g,g-1}-(t-g+1)\gamma_{g,g-2}$.

Group-Time ATT Identification

theorem[Identification under Parallel{[$p$]}] Under Assumptions (ref), (ref), and (ref) (for the relevant order $p$), with $m_g \geq p$, for any post-treatment period $t \geq g$: \begin{equation} \mathrm{ATT}(g,t) = \gamma_{g,t} - \gamma_{g,t}(0), \end{equation} where $\gamma_{g,t}$ is the observed post-treatment gap and $\gamma_{g,t}(0)$ is the polynomial counterfactual from Proposition (ref). Both are identified from observable data under Assumption (ref).
proofWrite: \begin{align*} \mathrm{ATT}(g,t) &= \mathbb{E}[Y_{it}\mid G_i=g] - \mathbb{E}[Y_{it}(\infty)\mid G_i=g]. \end{align*} By definition of $\gamma_{g,t}(0)$: $\mathbb{E}[Y_{it}(\infty)\mid G_i=g] = \mathbb{E}[Y_{it}(\infty)\mid G_i=\infty] + \gamma_{g,t}(0)$. Never-treated units are never treated, so $\mathbb{E}[Y_{it}(\infty)\mid G_i=\infty] = \mathbb{E}[Y_{it}\mid G_i=\infty]$, which is directly observable. Therefore: \begin{align*} \mathrm{ATT}(g,t) &= \mathbb{E}[Y_{it}\mid G_i=g] - \mathbb{E}[Y_{it}\mid G_i=\infty] - \gamma_{g,t}(0) \\ &= \gamma_{g,t} - \gamma_{g,t}(0). \qedhere \end{align*}

The treatment effect is the vertical distance between the observed post-treatment gap and the polynomial counterfactual: the hollow dots above the dashed line in Figure (ref).

The Short Pre-Period Problem and Cohort-Heterogeneous Orders

In a two-group setting, the number of pre-treatment periods is fixed for the single treated cohort, and the researcher simply chooses the largest $p$ the panel supports. In staggered settings, $m_g = g - t_{\min}$ varies across cohorts. Early-treated cohorts have small $m_g$; late-treated cohorts have large $m_g$.

definition[Feasible Set and Cohort-Specific Orders] For target order $p$, the feasible cohort set is $\mathcal{F}(p) = \{g : m_g \geq p\}$. The cohort-specific applied order is $p_g \in \{1,\ldots,m_g\}$, the order assumed for cohort $g$.

When $p_g < p$ for some cohorts, applying a uniform $p$ either requires excluding those cohorts (changing the estimand) or over-relaxing to a lower order than the data of later-treated cohorts can support. The following theorem resolves this by allowing heterogeneous orders across cohorts.

theorem[Aggregation under Cohort-Heterogeneous Orders] Suppose for each cohort $g$, Assumptions (ref)--(ref) hold, and Assumption (ref) holds for cohort-specific order $p_g$, where $1 \leq p_g \leq m_g$. Let $\mathbf{p} = (p_g)_g$. Then: \begin{enumerate}[label=(\roman*)] • (Cohort identification) $\mathrm{ATT}(g,t)$ is identified for each $g$ and $t \geq g$ by Theorem (ref) applied with order $p_g$. • (Aggregate identification) $\theta(\mathbf{p}) = \sum_{g}\sum_{t\geq g} w_{g,t}\,\mathrm{ATT}^{(p_g)}(g,t)$ is identified for any non-negative weights summing to one. • (Callaway--Sant'Anna as special case) When $p_g = 1$ for all $g$ and the weights $w_{g,t}$ are chosen to match the callaway2021difference aggregation scheme, $\theta(\mathbf{p})$ reduces to the Callaway--Sant'Anna aggregate ATT under standard parallel trends. \end{enumerate}
proofPart (i) follows from Theorem (ref) applied independently to each cohort under its own order $p_g$. Each cohort's counterfactual is constructed from that cohort's own pre-treatment gaps; although all cohorts share the same never-treated group as the comparison group, the Parallel[$p_g$] restriction for cohort $g$ is a population condition on cohort $g$'s potential outcomes and can be imposed independently of restrictions placed on other cohorts. Hence identification for one cohort does not require or restrict the identifying assumption of another. Part (ii): since each $\mathrm{ATT}^{(p_g)}(g,t)$ is identified (part i), and a finite weighted average of identified quantities is identified, $\theta(\mathbf{p})$ is identified. Formally, write $\theta(\mathbf{p}) = \boldsymbol{w}^\top \boldsymbol{a}$, where $\boldsymbol{a}$ collects all identified $\mathrm{ATT}^{(p_g)}(g,t)$ and $\boldsymbol{w}$ are the corresponding non-negative weights summing to one. Since each component of $\boldsymbol{a}$ is identified, so is $\theta$. Part (iii): when $p_g=1$ for all $g$, Proposition (ref) gives $\gamma_{g,t}(0) = \gamma_{g,g-1}$ for all $t \geq g$. This is the flat counterfactual of standard parallel trends. When the weights $w_{g,t}$ are furthermore chosen to match the callaway2021difference aggregation scheme, $\theta(\mathbf{p})$ identifies the same population parameter as the Callaway--Sant'Anna aggregate ATT under standard parallel trends.
remark[Economic interpretation of the aggregate] The structural parameter is $\mathrm{ATT}(g,t)$, defined in Section (ref) without reference to any identifying order. The superscript $(p_g)$ on $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ denotes the estimator constructed under order $p_g$, not a different population quantity: the target is always $\mathrm{ATT}(g,t)$. When different cohorts are identified under different orders $p_g$, the aggregate $\theta(\mathbf{p})$ remains interpretable because each cohort's $\mathrm{ATT}(g,t)$ is a well-defined structural object, the average difference between actual and counterfactual outcomes for that cohort, regardless of how it is identified. The identification strategy does not change the parameter; it only changes the assumption under which it is recovered. Transparency requires reporting which order was applied to each cohort; the three reporting strategies in Section (ref) formalise this.
remark[Relation to prior work] mora2019alternative prove the identification result in Theorem (ref) for the two-group case under Parallel[$p$]. callaway2021difference provide aggregation of group-time ATTs under standard parallel trends ($p_g = 1$ for all cohorts). egami2023using adapt higher-order parallelism to staggered designs. Theorem (ref) extends these results by allowing cohort-specific orders $p_g \in \{1,\ldots,m_g\}$; to my knowledge it is the first aggregation result for staggered DiD under cohort-heterogeneous feasible orders.

Estimation and Inference

The DD[\texorpdfstring{$p$}{p}] Estimator

Let $\bar Y_{g,t} = N_g^{-1}\sum_{i:G_i=g}Y_{it}$ and $\bar Y_{\infty,t}$ be the cohort and never-treated sample means.

Step 1. Sample gaps. $\hat\gamma_{g,t} = \bar Y_{g,t} - \bar Y_{\infty,t}$.

Step 2. Polynomial fit. Fit a polynomial of degree $p_g - 1$ by OLS to the $m_g$ pre-treatment gap observations:

equation[equation omitted — 161 chars of source]

When $m_g = p_g$, this is exact interpolation; when $m_g > p_g$, the OLS residuals provide the over-identifying restrictions used in Section (ref).\footnote{To improve numerical stability, particularly at orders $p \geq 3$, time should be centred at the mean pre-treatment year before constructing the polynomial basis. Results are invariant to this reparametrisation but condition numbers are substantially reduced.}

remark[OLS estimator versus minimum-point counterfactual] Proposition (ref) identifies the counterfactual from any $p_g$ pre-treatment observations, most naturally the $p_g$ most recent. Step 2 uses all $m_g$ pre-treatment observations in an OLS fit. Under exact Parallel[$p_g$], both approaches recover the same population polynomial, but they are different finite-sample estimators when $m_g > p_g$: the OLS fit exploits all available pre-treatment data and is generally more efficient. In particular, at $p_g = 1$ the OLS estimator averages all pre-treatment gaps rather than using only the last pre-period observation, so it does not reduce to a base-period comparison in finite samples. Applied researchers wishing strict numerical equivalence to Callaway--Sant'Anna at $p_g=1$ should use csdid directly; the present estimator delivers the same identification but uses a different finite-sample implementation.

Step 3 --- Counterfactual and ATT. $\hat\gamma_{g,t}(0) = \sum_{k=0}^{p_g-1}\hat c_{g,k}\,t^k$ and:

equation[equation omitted — 117 chars of source]

Step 4 --- Aggregate. $\hat\theta(\mathbf{p}) = \sum_{g,t\geq g} \hat w_{g,t}\,\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$, using cohort-share weights $\hat w_g = N_g / N_{\text{treated}}$, uniform across post-treatment event times within each cohort (see Table (ref) for sensitivity to alternative schemes).

Three reporting strategies. Three canonical order assignments provide a sensitivity check. Strategy I (Uniform Conservative): $p_g = 1$ for all cohorts, applying the strictest common assumption uniformly. Strategy II (Uniform Feasible): choose target $p \leq \min_g m_g$, restrict to $\mathcal{F}(p)$, same assumption for all included cohorts. Strategy III (Cohort Maximum): $p_g = m_g - 1$, the highest feasible order for each cohort given its pre-treatment series, ensuring at least one residual degree of freedom for the polynomial fit. This strategy fully exploits the heterogeneous-order aggregation of Theorem (ref) and is the recommended default when pre-treatment series are long and R$^2$ diagnostics support higher-order fits across cohorts. When pre-treatment series are short, the cohort-maximum order may be poorly identified; in such cases the sensitivity table and R$^2$ diagnostics should guide the choice between Strategy II and Strategy III.\footnote{Reporting all three strategies provides a sensitivity check on the choice of cohort weighting scheme.}

remark[Uniform versus cohort-heterogeneous orders in practice] Theorem (ref) establishes identification when cohorts are identified under different feasible orders $p_g$. In practice, the sequential algorithm in Section (ref) selects a single uniform order $p^\star$ applied to all cohorts, particularly when pre-treatment series are of similar length. This simplifies interpretation and comparison across cohorts without sacrificing the generality of the heterogeneous-order aggregation in Theorem (ref). Cohort-heterogeneous orders are appropriate when there is strong prior reason to believe different cohorts face different pre-trend structures, or when later-treated cohorts have substantially more pre-treatment periods and the researcher wishes to exploit additional data available for those cohorts. In either case, transparency requires reporting which order was applied to each cohort.

Asymptotic Properties

proposition[Asymptotic Normality] Under Assumptions (ref), (ref), and (ref) (for the relevant order per cohort), bounded second moments, independent sampling across units, and $\pi_g > 0$ for all $g$: \begin{enumerate}[label=(\roman*)] • $\sqrt{N}(\widehat{\mathrm{ATT}}^{(p_g)}(g,t) - \mathrm{ATT}(g,t)) \xrightarrow{d} \mathcal{N}(0, V_{g,t})$. • $\sqrt{N}(\hat\theta - \theta) \xrightarrow{d} \mathcal{N}(0, V)$. • $V$ is consistently estimated by the cluster bootstrap described below. \end{enumerate}
proofThe sample gap $\hat\gamma_{g,t}$ is a difference of sample means. By independence across units and bounded second moments: $\sqrt{N}(\hat\gamma_{g,t} - \gamma_{g,t}) \xrightarrow{d} \mathcal{N}(0, \sigma^2_{g,t})$ by the central limit theorem, where $\sigma^2_{g,t} = \operatorname{Var}(Y_{it}\mid G_i=g)/\pi_g + \operatorname{Var}(Y_{it}\mid G_i=\infty)/\pi_\infty$. The OLS polynomial coefficients $\hat{\mathbf{c}}_g$ are a linear function of the pre-treatment gaps $\{\hat\gamma_{g,t}\}_{t<g}$: $\hat{\mathbf{c}}_g = (\mathbf{V}^\top\mathbf{V})^{-1} \mathbf{V}^\top\hat{\boldsymbol{\gamma}}_g^{\mathrm{pre}}$, where $\mathbf{V}$ is the $(m_g\times p_g)$ Vandermonde matrix of pre-treatment time polynomials. By the delta method applied to the jointly normal vector of pre-treatment sample gaps, $\sqrt{N}(\hat{\mathbf{c}}_g - \mathbf{c}_g) \xrightarrow{d} \mathcal{N}(\mathbf{0}, \boldsymbol{\Sigma}_c)$ where $\boldsymbol{\Sigma}_c = (\mathbf{V}^\top\mathbf{V})^{-1}\mathbf{V}^\top \boldsymbol{\Sigma}_\gamma^{\mathrm{pre}} \mathbf{V}(\mathbf{V}^\top\mathbf{V})^{-1}$ and $\boldsymbol{\Sigma}_\gamma^{\mathrm{pre}}$ is the covariance matrix of the pre-treatment gaps. The counterfactual $\hat\gamma_{g,t}(0) = \mathbf{v}_t^\top\hat{\mathbf{c}}_g$, where $\mathbf{v}_t = (1,t,\ldots,t^{p_g-1})^\top$, is linear in $\hat{\mathbf{c}}_g$, hence also asymptotically normal by the delta method. The ATT estimate $\widehat{\mathrm{ATT}}^{(p_g)}(g,t) = \hat\gamma_{g,t} - \hat\gamma_{g,t}(0)$ is the difference of two jointly asymptotically normal quantities, hence asymptotically normal. Part (ii): the aggregate $\hat\theta$ is a finite weighted sum of cohort-level ATT estimates. All cohorts share the same never-treated control group, so their ATT estimators are not independent: sampling variation in the common control group enters every $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$. The correct aggregate variance, as the limit of $N$ times the finite-sample covariance matrix, is: $V = \lim_{N\to\infty} N \cdot \sum_{g,t,g',t'} w_{g,t}w_{g',t'} \operatorname{Cov}(\widehat{\mathrm{ATT}}^{(p_g)}(g,t), \widehat{\mathrm{ATT}}^{(p_{g'})}(g',t'))$, which includes cross-cohort covariance terms arising through the shared control. Asymptotic normality of $\hat\theta$ nonetheless follows because $\hat\theta$ is a linear function of the jointly asymptotically normal vector of all gap estimates, and the delta method applies. The full variance $V$, including cross-cohort terms, is consistently estimated by the cluster bootstrap (Part iii). Part (iii): the cluster bootstrap (resampling entire units) consistently estimates the asymptotic variance under the stated regularity conditions, because it correctly replicates the joint sampling distribution of $(\hat\gamma_{g,t})_{g,t}$ across cohorts; see callaway2021difference for details of the argument in the analogous setting.\footnote{Note that serial correlation within units is accommodated by the cluster bootstrap, which resamples entire units and thereby preserves within-unit dependence across time. The analytical variance formula in Part (i) assumes independence across time periods within units and should not be used directly under serial correlation; inference should rely on the bootstrap throughout.}
remark[Small cohorts and finite-sample fragility] Proposition (ref) assumes each cohort fraction $\pi_g$ remains positive as $N \to \infty$. In practice, the delayed-expansion cohorts in the Medicaid application contain 3, 2, 1, and 2 states respectively. A percentile cluster bootstrap cannot establish a central limit theorem for a singleton treated cohort. Inferential results for cells involving the 2017 cohort (one state) and other small cohorts should be treated with caution; the reported estimates for these cohorts are best understood as descriptive comparisons rather than asymptotically valid inference.

Influence function. The influence function for $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ has two components. Let $\psi_{it}^{\mathrm{post}}$ be the unit's contribution to the post-treatment gap $\hat\gamma_{g,t}$, and let $\psi_{it}^{\mathrm{pre}}$ capture its contribution to the polynomial coefficients through the pre-treatment gaps. The counterfactual evaluated at $t$ is linear in the pre-treatment gaps via the Vandermonde projection: $\hat\gamma_{g,t}(0) = \mathbf{v}_t^\top(\mathbf{V}^\top\mathbf{V})^{-1} \mathbf{V}^\top\hat{\boldsymbol{\gamma}}_g^{\mathrm{pre}}$. The influence function for $\hat\gamma_{g,t}(0)$ is therefore a linear combination of influence functions for the pre-treatment gaps, with weights $\mathbf{v}_t^\top(\mathbf{V}^\top\mathbf{V})^{-1} \mathbf{V}^\top$. The full influence function for $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ is $\psi_{it} = \psi_{it}^{\mathrm{post}} - \psi_{it}^{\mathrm{cf}}$, where $\psi_{it}^{\mathrm{cf}}$ is the component arising from polynomial estimation. When $p_g=1$, $\psi_{it}^{\mathrm{cf}}$ is analogous to the pre-treatment contribution in the Callaway--Sant'Anna influence function, though the two differ when more than one pre-treatment period is used: the present estimator averages all pre-treatment gaps, while the Callaway--Sant'Anna construction uses the relevant base period.

Cluster bootstrap. For state-level panel data, all inference uses the cluster bootstrap: resample entire states with replacement ($B$ draws), compute $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ on each draw, and construct 95\,% confidence intervals using the normal approximation (estimate $\pm z_{0.025} \times$ bootstrap standard error). Section (ref) shows that $B = 999$ draws yields stable standard errors for datasets of the size studied here.

\paragraph{Dependence across cohorts.} When states belong to the same geographical or institutional cluster (e.g., Census division), their outcomes may be correlated across cohorts. The cluster bootstrap at the state level accounts for within-state dependence across time but assumes independence across states. In settings with strong cross-state dependence (e.g., common macro shocks affecting all states simultaneously), the standard errors may be understated. One recommendation is to supplement with randomisation inference, assigning treatment dates randomly to never-treated states and checking whether the estimated ATT exceeds the permutation distribution rambachan2023more.

Order Selection and Diagnostic Tests

Testable Restrictions

Lemma (ref) implies that when $m_g > p$, Parallel[$p$] generates $m_g - p$ testable restrictions on pre-treatment data. For cohort $g$, define:

equation[equation omitted — 199 chars of source]

Under the null Parallel[$p$] and the additional assumption that the higher-order differences $\Delta^{p}\hat\gamma_{g,t}$ are asymptotically uncorrelated across $t$, $T_g(p) \xrightarrow{d} \chi^2(m_g-p)$. This uncorrelatedness condition is not automatically satisfied even under serial independence of raw outcomes: applying the $p$-th difference operator to an i.i.d.\ sequence mechanically induces a moving-average dependence of order $p-1$ among the differenced terms. In practice, $T_g(p)$ is therefore best treated as a descriptive diagnostic rather than a formal test; researchers should complement it with the pre-period $R^2$ and visual residual inspection, and rely on the cluster bootstrap for inference. Pooling across cohorts:

equation[equation omitted — 150 chars of source]

In settings with serial correlation or strong cross-cohort dependence, the chi-square approximation may be unreliable and practitioners should rely on permutation-based critical values or treat the statistic as a descriptive diagnostic rather than a formal test. The cluster bootstrap used for inference does not depend on these distributional assumptions.

Sequential Algorithm

Let $\alpha$ be the significance level and $p_{\max} = \min_g m_g$.

enumerate[label=Step \arabic*.] • For each cohort $g$, test $T_g(1)$ at level $\alpha$. If not rejected, adopt $p_g=1$ for that cohort. • If $T_g(1)$ rejected, test $T_g(2)$. If not rejected, adopt $p_g=2$. • Continue until either the current test is not rejected or $p_g = p_{\max}$.

The procedure is applied independently to each cohort, allowing different cohorts to be identified under different selected orders. The pre-period $R^2$ from the polynomial fit provides a complementary diagnostic: low $R^2$ at a given order signals poor polynomial fit regardless of the test outcome, which can occur when pre-treatment series are short and the F-test has limited power. In practice, $R^2$ and the sensitivity of the aggregate ATT across orders should guide order selection alongside the formal test.

remark[Polynomial misspecification diagnostic] Before applying the sequential test, I recommend inspecting the polynomial fit via the in-sample $R^2$ from the pre-period OLS regression. Low $R^2$ signals that the pre-treatment gap does not follow a polynomial of the chosen degree potentially indicating structural breaks, seasonal patterns, or logistic growth. Section (ref) shows that under such misspecification, $R^2$ deteriorates noticeably, providing an early warning. In such cases, the polynomial extrapolation may not be reliable and researchers may find it useful to complement the point estimates with sensitivity bounds rambachan2023more to assess robustness to departures from the polynomial structure.
remark[Extrapolation horizon and polynomial behaviour] A potential limitation of polynomial extrapolation is that fitted polynomials can behave erratically when projected beyond the support of the estimation data. An analogous concern motivates the critique of high-degree polynomial regression in the regression discontinuity literature gelman2019high, though the setting here differs: the polynomial is fitted to pre-treatment gaps in time rather than to outcomes near a discontinuity threshold. Appendix (ref) quantifies how performance changes with the extrapolation horizon. Variance rather than bias is the primary cost of longer horizons. Applied researchers working with long post-treatment windows may find it useful to complement DD[$p$] estimates with the derivative-bounded sensitivity bounds of rambachan2023more, which do not rely on polynomial extrapolation.
remark[Post-selection inference] I recommend reporting estimates for $p=1,2,3$ regardless of the sequential test outcome. The selected order is the primary specification; others are robustness checks. Section (ref) demonstrates that when the correct order is selected, bootstrap confidence intervals achieve approximately nominal coverage. When over-selection occurs (using $p=2$ when $p=1$ is true), coverage deteriorates, this is the cost of unnecessary flexibility. The sequential test minimises over-selection by adopting the lowest order the data do not reject.

Monte Carlo Evidence

Main Simulation Design

I evaluate DD[$p$] for $p \in \{1,2,3\}$ under three data-generating processes. The panel has $N=300$ units, $T=12$ periods, three cohorts ($g=5,7,9$) with 4, 6, and 8 pre-treatment periods, and 40\,% never-treated. True ATT $= 0.5$ throughout. Results are based on 500 replications.

DGP-1 (True Parallel[1]). $Y_{it}(\infty) = \alpha_i + \lambda_t + \varepsilon_{it}$, $\alpha_i \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,1)$, $\varepsilon_{it} \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,1)$, $\lambda_t$ a common trend. DD[1] is the efficient estimator.

DGP-2 (True Parallel[2], not Parallel[1]). $Y_{it}(\infty) = \alpha_i + \beta_g t + \lambda_t + \varepsilon_{it}$, $\beta_g \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,0.5)$. Cohort-specific linear trends violate Parallel[1]. DD[2] is the correctly specified estimator.

DGP-3 (True Parallel[3]). $Y_{it}(\infty) = \alpha_i + \beta_g t + \gamma_g t^2 + \lambda_t + \varepsilon_{it}$, $\gamma_g \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,0.15)$. Cohort-specific quadratic trends.

Table (ref) reports bias and RMSE; Figure (ref) displays them.

table[table omitted — 1,216 chars of source]
figure[figure omitted — 481 chars of source]

Post-Selection Inference

Table (ref) reports order-selection frequencies and empirical coverage of 95\,% confidence intervals after the sequential order-selection procedure, based on 500 simulation draws with $B=99$ cluster bootstrap replications per draw. Each row corresponds to a different data-generating process; the selection frequencies show which order the sequential F-test chooses under each DGP, and the coverage column reports how often the resulting confidence interval contains the true ATT.

table[table omitted — 1,532 chars of source]

Selection frequencies are consistent with the test's design: it selects the lowest order whose pre-treatment implications the data do not reject, defaulting toward parsimony. Coverage after selection is near-nominal for DGP-2 and DGP-3 (93.6% and 93.8%) and slightly below nominal for the flat and borderline DGPs (89.8% and 89.2%), where finite-sample bootstrap imprecision at $B=99$ accounts for the shortfall.

Serial Correlation and Placebo

Table (ref) reports two additional robustness checks. Panel A shows that DD[2] maintains near-nominal coverage under AR(1) serial correlation with persistence $\rho$ up to 0.7, ranging from 91.4% at $\rho=0$ to 94.4% at $\rho=0.7$. This reflects that the estimator operates on cohort-level averages, which attenuate within-unit serial correlation. Panel B shows a placebo false-positive rate of 7.2%, modestly above the nominal 5%, consistent with the finite-sample bootstrap behaviour at $B=99$ documented in Table (ref); the mean estimated ATT under the null is essentially zero, confirming the absence of systematic bias.

table[table omitted — 1,364 chars of source]

Failure Modes and the Polynomial Diagnostic

The simulations above calibrate DGPs to exact polynomial trends. When the pre-treatment gap follows a non-polynomial trajectory, such as a structural break or logistic growth, Parallel[$p$] is misspecified regardless of the chosen order.

The practical diagnostic is the pre-period $R^2$ from the OLS polynomial fit: under a genuine polynomial trend $R^2$ is high; under structural breaks or other non-polynomial dynamics $R^2$ deteriorates noticeably, providing an early warning before the sequential test is applied. See Remark (ref). When the $R^2$ diagnostic signals misspecification, sensitivity bounds rambachan2023more may be more appropriate. Table (ref) shows how the cluster bootstrap standard error for the aggregate DD[2] ATT stabilises as $B$ increases, using the Medicaid expansion application data.

table[table omitted — 523 chars of source]

Empirical Application: Medicaid Expansion

Setting and Data

I use a state-year panel covering 46 states observed over 2008--2019, yielding 552 observations.\footnote{Data are drawn from the Mixtape Sessions Advanced DiD repository, made publicly available by Scott Cunningham and accessible at \url{https://raw.githubusercontent.com/Mixtape-Sessions/Advanced-DID/main/Exercises/Data/ehec_data.dta}. The underlying microdata source is the American Community Survey (ACS). } The outcome (dins) is the share of low-income childless adults with health insurance in each state and year, derived from the American Community Survey (ACS). Treatment is Medicaid expansion under the Affordable Care Act of 2010, which gave states the option to expand Medicaid eligibility beginning in 2014. Five expansion cohorts are present: 2014 (22 states, $m_g=6$), 2015 (3, $m_g=7$), 2016 (2, $m_g=8$), 2017 (1, $m_g=9$), and 2019 (2, $m_g=11$). Sixteen states never expanded and serve as the never-treated comparison group throughout. Population weights are normalised to mean one to ensure interpretable test statistics.\footnote{Raw weights average 647,395 per state-year, inflating chi-squared statistics proportionally to the sum of weights rather than the number of observations when used un-normalised. Normalising preserves relative weighting across states.}

Pre-Trend Diagnostics

The Callaway--Sant'Anna aggregate ATT under Parallel[1] is $\hat\theta_{\mathrm{CS}} = 0.075$ (95\,% CI: [0.051, 0.099], $p < 0.001$). The joint pre-trend test yields $\chi^2(36) = 65{,}921$, $p < 0.001$, decisively rejecting standard parallel trends.\footnote{The large chi-squared reflects importance-weight scaling in csdid; see footnote 1. The qualitative rejection is unaffected.} The pre-treatment gap for the 2014 cohort rises steadily from 0.032 in 2008 to 0.042 in 2013, consistent with a linear trend. I select $p^\star = 2$ as the primary specification based on pre-period $R^2$ diagnostics and sensitivity stability across orders (Table (ref)).

figure[figure omitted — 879 chars of source]
table[table omitted — 920 chars of source]
table[table omitted — 918 chars of source]

Main Results

Table (ref) reports DD[1], DD[2], and DD[3] estimates at event times $\tau = 0,\ldots,4$ with 95\,% confidence intervals from the state-level cluster bootstrap ($B=999$). Figure (ref) displays the event-study comparison between DD[1] and the primary DD[2] specification. Figure (ref) extends this to include DD[3] as a robustness check.\footnote{The DD[1] estimates in Table (ref) use the last pre-treatment gap as the flat counterfactual, following the callaway2021difference convention. The unweighted CS simple aggregate (0.068) is close to the DD[1] cohort-weighted average (0.065), since both use the same identifying assumption and the same baseline convention. The weighted CS estimate (0.075) differs because it uses state population weights; the DD[$p$] estimates in this table use equal cohort-share weights.}

The findings show that first, all three estimators find a positive, significant, and growing effect. Insurance coverage rose by approximately 4--5 percentage points in the year of expansion and 7--8 points four years later. This qualitative conclusion is robust across identification strategies.

Second, the DD[1] cohort-weighted average (0.065) is nearly identical to the unweighted Callaway--Sant'Anna simple aggregate (0.068), confirming that DD[1] as implemented here --- using the last pre-treatment gap as the flat counterfactual --- reproduces the CS estimator closely. The gap between the weighted CS estimate (0.075) and these figures reflects population weighting, not the identifying assumption.

Third, DD[2] consistently exceeds DD[1] at every event time, with differences of 0.4--1.0 percentage points (10--22 percent of the DD[1] estimate). This reflects that later-treated cohorts (2015--2019) were on a modest downward trajectory relative to never-expanding states before their expansion. The standard estimator, which assumes a flat counterfactual, underestimates the treatment effect for these cohorts. DD[2] corrects for this by projecting the downward pre-trend forward.

Fourth, the confidence intervals of DD[1] and DD[2] overlap substantially at every event time: I cannot reject the null that the two estimators produce the same treatment effect. The contribution of DD[2] in this application is therefore not a reversal of the standard finding but a more credible quantification of it. The same qualitative conclusion, Medicaid expansion increased insurance coverage, is supported by an assumption the pre-treatment data do not reject, rather than one they reject.

table[table omitted — 1,674 chars of source]
figure[figure omitted — 392 chars of source]
figure[figure omitted — 405 chars of source]

Table (ref) summarises estimates across all three specifications and records which pass the pre-trend test.

table[table omitted — 946 chars of source]

Weighting Sensitivity

Table (ref) shows the aggregate ATT under three weighting strategies. Results are insensitive to weighting choice: DD[2] ranges from 0.064 to 0.067 and DD[1] from 0.060 to 0.065. The cohort-level results in the penultimate row reveal that the direction of the DD[2] correction varies by cohort: for the large 2014 cohort, DD[2] is slightly smaller than DD[1] (its pre-trend was upward), while for later cohorts DD[2] exceeds DD[1] (their pre-trends were downward). This heterogeneity in direction is precisely what motivates cohort-specific counterfactuals.

table[table omitted — 683 chars of source]

Relation to Sensitivity Bounds

A natural benchmark is the sensitivity bounds approach of rambachan2023more, which restricts the magnitude of parallel trends violations and reports how estimated effects vary within that restriction. Their approach delivers valid intervals whose width reflects the assumed maximum violation magnitude. Notably, rambachan2023more use the same dataset as an illustrative example in their software package HonestDiD. Their smoothness restriction at $M=0$ imposes that the post-treatment counterfactual gap follows a linear extrapolation of the pre-existing trend which is conceptually equivalent to DD[2] at Parallel[2]. For $M > 0$, their approach allows deviations from linearity up to a bound of size $M$, providing a natural complement to the point estimates reported here: DD[2] delivers a point estimate under the polynomial structure, while rambachan2023more deliver robust intervals that remain valid even if the linear structure is imperfect. DD[2] complements this analysis by showing that if the specific polynomial structure is credible, as the pre-treatment $R^2$ values and visual evidence suggest for most cohorts, one can achieve point identification without widening the bounds. The two approaches are not competitors: if the $R^2$ diagnostic signals misspecification, the rambachan2023more bounds remain the appropriate tool.

Conclusion

When pre-treatment event studies reject standard parallel trends, the applied researcher faces a difficult choice: impose a maintained assumption the data contradict, or report sensitivity bounds that are valid but do not deliver a point estimate. This paper offers a third option: replacing the flatness requirement with a structured polynomial extrapolation whose pre-treatment implications are directly checkable.

The central theoretical contribution is Theorem (ref) an aggregation result allowing cohorts to be identified under different feasible orders, arising naturally from variation in pre-treatment period counts across staggered adoption cohorts. This problem has no analogue in two-group settings and is not addressed by any existing paper, including egami2023using, who extend higher-order parallelism to staggered designs at $p=2$ under a uniform order.

Monte Carlo evidence shows near-nominal coverage when the correct order is selected, robustness to AR(1) serial correlation, and a placebo false positive rate modestly above the nominal 5\,%, consistent with finite-sample bootstrap behaviour at $B=99$. Applied to Medicaid expansion, the approach yields estimates resting on an assumption the pre-treatment data do not reject.

Two directions merit future work. First, a formal uniform validity result for the complete selection-estimation pipeline would strengthen the inferential foundations. The extended version of roth2022pretest on post-selection inference provides a starting point, though adapting those results to the sequential polynomial order selection here requires additional work. Second, replacing the polynomial order restriction with derivative-bounded smoothness restrictions rambachan2023more would extend the framework to settings where the pre-treatment gap does not follow a polynomial, at the cost of point identification.