Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,132 characters · 31 sections · 40 citation commands
\footnotetext[1]{I am extremely grateful to Jonathan Roth for detailed comments that improved the paper. I thank Scott Cunningham for preparing and making publicly available the cleaned Medicaid expansion panel used in the empirical application; the underlying microdata are from the American Community Survey. All errors are my own.}
Difference-in-differences (DiD) is among the most widely used designs for program and policy evaluation. In staggered adoption settings, recent work has shown that the conventional two-way fixed effects regression can deliver misleading estimates when treatment effects vary across cohorts, and has proposed alternative estimators that recover interpretable average treatment effects callaway2021difference,sun2021estimating,goodman2021difference,de2020two,borusyak2024revisiting. All of these frameworks, however, rest on the same core identifying assumption: parallel trends.
Parallel trends requires that, absent treatment, the average outcome gap between treated and control units would have remained constant over time. In practice, pre-treatment event studies frequently show systematic trends that are difficult to reconcile with this requirement, especially when pre-periods are long. When this happens, treatment effect estimates are biased, and the magnitude of that bias is generally unknown. The common response is to test for pre-trends and proceed only if the test passes, but this practice is itself problematic: pre-trends tests are often underpowered, and conditioning the analysis on having passed one distorts subsequent inference roth2022pretest. A researcher who finds a trending pre-period is left without a clear way forward.
This paper asks: when parallel trends fails in a staggered adoption design, can we still point-identify average treatment effects without discarding some cohorts or bounding the effect under assumptions about the magnitude of the violation? I propose a middle path. Standard parallel trends implicitly assumes that whatever shape the pre-treatment gap takes, it would have become flat after treatment. I instead use the information in the pre-treatment trajectory: when the gap evolves smoothly, that evolution is informative about what it would have done absent treatment. I replace Parallel[1] with a hierarchy of higher-order assumptions, Parallel[$p$], and project the pre-treatment trajectory forward as the counterfactual. Where rambachan2023more trace out an identified set indexed by the strength of a smoothness restriction, I instead select an order from the pre-treatment data and report a point estimate under that order.
Formally, Parallel[$p$] requires only that the $p$-th order time difference of the untreated outcome gap is common across groups. Parallel[1] is the standard assumption: the gap must be flat. Parallel[2] allows the gap to trend linearly; Parallel[3] allows a quadratic trajectory; higher orders allow richer curvature. Each step up in $p$ is strictly weaker than the previous one and remains sufficient for point identification of group-time average treatment effects. When a cohort has more than $p$ pre-treatment periods, Parallel[$p$] generates testable implications about the pre-treatment gap that can be assessed with observable data.
These implications do not, however, eliminate extrapolation. Every DiD estimator relies on an extrapolation assumption: Parallel[1] extrapolates a flat gap into the post-treatment period, while Parallel[$p$] extrapolates the pre-existing polynomial trajectory. Neither is testable in the post-treatment period. The advantage of Parallel[$p$] is that its key implication, that the pre-treatment gap lies on a low-degree polynomial, is testable with observed data, whereas the flatness required by Parallel[1] is often visibly violated. Parallel[$p$] thus replaces an assumption the data may reject with one they do not, recovering point identification where standard methods cannot proceed.
The paper makes three contributions. First, I formally identify $\mathrm{ATT}(g,t)$ under Parallel[$p$] for arbitrary $p \geq 1$ in the staggered adoption framework of callaway2021difference, using the polynomial structure of the pre-treatment gap to construct a cohort-specific counterfactual in each post-treatment period.
Second, I address a problem that arises only in staggered settings. Because treatment timing varies, cohorts differ in how many pre-treatment periods they have, and therefore in the maximum order each can support: a cohort treated two periods into the panel can support only Parallel[1], while one treated ten periods in can support up to Parallel[9]. The natural question is how to combine cohorts that are identified under different orders into a single summary effect. Theorem (ref) resolves this, showing that coherent weighted averages of group-time ATTs remain identified even when each cohort $g$ is identified under its own order $p_g$. To my knowledge, this is the first aggregation result for staggered DiD under cohort-heterogeneous orders.
Third, I develop a sequential order-selection procedure tailored to staggered settings, provide Monte Carlo evidence on finite-sample performance and post-selection inference, and implement the method in a companion Stata command anddp, which produces the aggregate DD[$p$] ATT with cluster bootstrap inference. The selection procedure responds directly to the pre-testing critique: rather than testing the flat-gap null and conditioning on passing it, it selects the lowest order the pre-treatment data do not reject and estimates under that order. The Monte Carlo evidence shows that bootstrap confidence intervals retain near-nominal coverage after this selection step.
I apply the method to the staggered adoption of Medicaid expansion under the Affordable Care Act, using state-level insurance coverage data. Expanding and non-expanding states were on visibly diverging insurance trajectories before any expansion took effect, and a joint pre-trends test rejects flatness decisively. Standard DiD, which assumes the gap would have stayed flat, is therefore hard to defend here. The proposed estimator, DD[$p$], nonetheless recovers point estimates under Parallel[$p$], an assumption the same pre-treatment data do not reject, and yields: expansion increased insurance coverage by roughly six percentage points.
The higher-order parallelism idea builds on mora2019alternative, who develop the Parallel[$p$] hierarchy and prove identification in a two-group, multiple-period setting. This paper extends their framework to staggered adoption. egami2023using who adapt the Mora--Reggio idea to staggered designs using a generalised method of moments (GMM) estimator combining Parallel[1] and Parallel[2]. I differ in three respects. First, I work with the full hierarchy for arbitrary $p \geq 1$, not only $p \in \{1,2\}$. Second, I select a single order per cohort by a sequential test and report which was used, rather than combining fixed orders. Third, because cohorts may be identified under different selected orders, aggregating them raises a problem that does not arise under their fixed-order combination, the problem Theorem (ref) resolves.
Prior responses to parallel trends violations include the following. Sensitivity bounds rambachan2023more restrict how far the post-treatment violation of parallel trends can depart from the pre-treatment trend, and report how the identified set for the treatment effect widens as that restriction is relaxed; the approach is rigorous and transparent, and rather than committing to a single counterfactual it traces out a range of estimates indexed by the strength of the assumption. The closest connection is to their smoothness restriction: assuming the post-treatment violation is exactly linear coincides with Parallel[2] in the present hierarchy. The approaches diverge in what they do with that restriction. They treat linearity as one end of a continuum and report how the identified set grows as it is relaxed, whereas I select an order from the data and report a point estimate under it. Pre-trend tests roth2022pretest diagnose violations but do not provide a corrected estimator. A complementary approach is the non-inferiority framework of bilinski2026nothing, which recasts pre-trend assessment as an equivalence test: rather than asking whether pre-trends are exactly zero, it asks whether deviations are small enough not to matter for the estimated effect. Their procedure operates under Parallel[1] and provides a decision rule for tolerating small violations. roth2023s provide a comprehensive review of these and other responses to parallel trends violations.
A related practice in applied work is to include unit-specific linear time trends in TWFE regressions to absorb differential pre-trends angrist2009mostly. dobkin2018economic adopt a conceptually related strategy in a single-timing event study: they include a linear trend in event time alongside a saturated set of post-treatment dummies. Because the post-treatment dummies absorb all post-treatment variation, the slope coefficient is identified from pre-treatment data only. This is formally equivalent to the $p=2$ case of the present framework applied to a single cohort, where no aggregation is required. Dobkin's linear extrapolation is therefore a single one-step relaxation of flat parallel trends. DD[$p$], the estimator I develop under Parallel[$p$], generalises this in three ways: it extends to arbitrary order $p \geq 1$ with a sequential procedure to select the order the data support; it handles staggered adoption with multiple cohorts; and it provides formal aggregation across cohorts identified under different feasible orders, together with asymptotic inference (see Appendix (ref)).
Other approaches to parallel-trends failures include synthetic control methods abadie2010synthetic,xu2017generalized,arkhangelsky2021synthetic,ben2022synthetic, which construct a weighted comparison group that matches treated units' pre-treatment trajectories and work well with long pre-treatment series and relatively few treated units. Triple-differences designs strezhnev2023decomposing,ortiz2025better require a placebo stratum known to be unaffected by treatment. Change-in-changes athey2006identification relies on rank preservation. Partial identification approaches manski2018right deliver bounds on treatment effects. Each of these methods is suited to a different empirical context; the present approach targets the common setting where the pre-treatment gap evolves smoothly and a polynomial counterfactual is credible.
The remainder of the paper is organised as follows. Section (ref) establishes the framework. Section (ref) introduces the hierarchy of higher-order parallelism assumptions. Section (ref) derives the identification results, including the aggregation theorem. Section (ref) presents estimation and inference. Section (ref) develops the order-selection procedure. Section (ref) reports Monte Carlo evidence. Section (ref) presents the Medicaid expansion application. Section (ref) concludes.
The notation and staggered-adoption setup follow callaway2021difference. Consider a balanced panel of $N$ units over $T$ periods $t = 1,\ldots,T$. Unit $i$ has adoption date $G_i \in \{g_1,\ldots,g_K,\infty\}$. $G_i = g$ means unit $i$ is first treated in period $g$; $G_i = \infty$ means never treated. Treatment is absorbing: once a unit is treated it remains treated in all subsequent periods, so a unit's adoption date $G_i$ fully summarises its treatment path. Units with $G_i = g$ form cohort $g$. Let $N_g$ be the size of cohort $g$ and $N_\infty$ the number of never-treated units, with $N = \sum_g N_g + N_\infty$.
Asymptotics are large-$N$ with $T$ fixed. The cohort fractions $\pi_g = N_g/N$ are held fixed as $N$ grows, which is achieved by requiring each $N_g$ to grow proportionally with $N$, a standard large-sample device that prevents any cohort from vanishing or dominating in the limit.
The number of pre-treatment periods for cohort $g$ is:
where $t_{\min}$ is the first observed period.
For each unit $i$ and period $t$, let $Y_{it}(g)$ denote the potential outcome if first treated in period $g$, and $Y_{it}(\infty)$ the never-treated potential outcome neyman1923application,rubin1974estimating. The observed outcome is:
Before treatment, observed outcomes equal untreated potential outcomes, making pre-treatment data informative about the untreated trajectory.
Units do not alter behaviour in anticipation of future treatment, so their pre-treatment observed outcomes equal their untreated potential outcomes.
The group-time average treatment effect is:
The expectation averages over all units $i$ belonging to cohort $g$, comparing their actual post-treatment outcome to what they would have experienced had treatment never occurred.
A scalar summary is the weighted aggregate:
for non-negative weights $w_{g,t}$ summing to one.
Standard parallel trends says that, absent treatment, the average untreated outcome would have changed by the same amount each period for every group, treated cohorts and never-treated alike. Equivalently, the gap between any cohort and the never-treated group would have stayed constant over time. This has a testable pre-treatment implication: the pre-treatment gap should be flat. When the event-study plot shows a visible pre-treatment trend, that implication fails, and Parallel[$1$] is implausible.
The $p$-th order difference operator is defined recursively: $\Delta^{1}Y_{it} = Y_{it} - Y_{i,t-1}$ and $\Delta^{p}Y_{it} = \Delta^{1}(\Delta^{p-1}Y_{it})$ for $p \geq 2$. Written out:
$\Delta^{1}$ measures the year-on-year change; $\Delta^{2}$ measures how that change is itself changing; $\Delta^{3}$ measures the rate of change of $\Delta^{2}$, and so on.
To build intuition, suppose the gap between treated and control groups rises by 0.01 per year before treatment. Parallel[1] requires the gap to be constant, so a steadily rising gap is inconsistent with it. Parallel[2] requires only that this 0.01 annual change is common to the treated and never-treated groups, a weaker restriction that the pre-treatment data may well support.
Parallel[1] is the special case $p=1$. Parallel[2] allows different levels and slopes across groups, requiring only that the acceleration is common. Each step up in $p$ is a strictly weaker assumption and requires one additional pre-treatment period to test.
Define the observed gap between cohort $g$ and never-treated units in period $t$:
Under Assumption (ref), this equals the untreated potential outcome gap in the pre-treatment period.
Lemma (ref) makes the identifying content of Parallel[$p$] concrete and testable: the pre-treatment gap must lie on a polynomial of degree $p-1$. For $p=1$: flat (constant). For $p=2$: straight line. For $p=3$: quadratic curve. The residuals from fitting this polynomial to the pre-treatment data are the testable implications used in Section (ref).
The counterfactual gap, what the gap would have been in the post-treatment period absent treatment, is:
This is unobservable for $t \geq g$ because treated units observed outcomes include the treatment effect.
The treatment effect is the vertical distance between the observed post-treatment gap and the polynomial counterfactual: the hollow dots above the dashed line in Figure (ref).
In a two-group setting, the number of pre-treatment periods is fixed for the single treated cohort, and the researcher simply chooses the largest $p$ the panel supports. In staggered settings, $m_g = g - t_{\min}$ varies across cohorts. Early-treated cohorts have small $m_g$; late-treated cohorts have large $m_g$.
When $p_g < p$ for some cohorts, applying a uniform $p$ either requires excluding those cohorts (changing the estimand) or over-relaxing to a lower order than the data of later-treated cohorts can support. The following theorem resolves this by allowing heterogeneous orders across cohorts.
Let $\bar Y_{g,t} = N_g^{-1}\sum_{i:G_i=g}Y_{it}$ and $\bar Y_{\infty,t}$ be the cohort and never-treated sample means.
Step 1. Sample gaps. $\hat\gamma_{g,t} = \bar Y_{g,t} - \bar Y_{\infty,t}$.
Step 2. Polynomial fit. Fit a polynomial of degree $p_g - 1$ by OLS to the $m_g$ pre-treatment gap observations:
When $m_g = p_g$, this is exact interpolation; when $m_g > p_g$, the OLS residuals provide the over-identifying restrictions used in Section (ref).\footnote{To improve numerical stability, particularly at orders $p \geq 3$, time should be centred at the mean pre-treatment year before constructing the polynomial basis. Results are invariant to this reparametrisation but condition numbers are substantially reduced.}
Step 3 --- Counterfactual and ATT. $\hat\gamma_{g,t}(0) = \sum_{k=0}^{p_g-1}\hat c_{g,k}\,t^k$ and:
Step 4 --- Aggregate. $\hat\theta(\mathbf{p}) = \sum_{g,t\geq g} \hat w_{g,t}\,\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$, using cohort-share weights $\hat w_g = N_g / N_{\text{treated}}$, uniform across post-treatment event times within each cohort (see Table (ref) for sensitivity to alternative schemes).
Three reporting strategies. Three canonical order assignments provide a sensitivity check. Strategy I (Uniform Conservative): $p_g = 1$ for all cohorts, applying the strictest common assumption uniformly. Strategy II (Uniform Feasible): choose target $p \leq \min_g m_g$, restrict to $\mathcal{F}(p)$, same assumption for all included cohorts. Strategy III (Cohort Maximum): $p_g = m_g - 1$, the highest feasible order for each cohort given its pre-treatment series, ensuring at least one residual degree of freedom for the polynomial fit. This strategy fully exploits the heterogeneous-order aggregation of Theorem (ref) and is the recommended default when pre-treatment series are long and R$^2$ diagnostics support higher-order fits across cohorts. When pre-treatment series are short, the cohort-maximum order may be poorly identified; in such cases the sensitivity table and R$^2$ diagnostics should guide the choice between Strategy II and Strategy III.\footnote{Reporting all three strategies provides a sensitivity check on the choice of cohort weighting scheme.}
Influence function. The influence function for $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ has two components. Let $\psi_{it}^{\mathrm{post}}$ be the unit's contribution to the post-treatment gap $\hat\gamma_{g,t}$, and let $\psi_{it}^{\mathrm{pre}}$ capture its contribution to the polynomial coefficients through the pre-treatment gaps. The counterfactual evaluated at $t$ is linear in the pre-treatment gaps via the Vandermonde projection: $\hat\gamma_{g,t}(0) = \mathbf{v}_t^\top(\mathbf{V}^\top\mathbf{V})^{-1} \mathbf{V}^\top\hat{\boldsymbol{\gamma}}_g^{\mathrm{pre}}$. The influence function for $\hat\gamma_{g,t}(0)$ is therefore a linear combination of influence functions for the pre-treatment gaps, with weights $\mathbf{v}_t^\top(\mathbf{V}^\top\mathbf{V})^{-1} \mathbf{V}^\top$. The full influence function for $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ is $\psi_{it} = \psi_{it}^{\mathrm{post}} - \psi_{it}^{\mathrm{cf}}$, where $\psi_{it}^{\mathrm{cf}}$ is the component arising from polynomial estimation. When $p_g=1$, $\psi_{it}^{\mathrm{cf}}$ is analogous to the pre-treatment contribution in the Callaway--Sant'Anna influence function, though the two differ when more than one pre-treatment period is used: the present estimator averages all pre-treatment gaps, while the Callaway--Sant'Anna construction uses the relevant base period.
Cluster bootstrap. For state-level panel data, all inference uses the cluster bootstrap: resample entire states with replacement ($B$ draws), compute $\widehat{\mathrm{ATT}}^{(p_g)}(g,t)$ on each draw, and construct 95\,% confidence intervals using the normal approximation (estimate $\pm z_{0.025} \times$ bootstrap standard error). Section (ref) shows that $B = 999$ draws yields stable standard errors for datasets of the size studied here.
\paragraph{Dependence across cohorts.} When states belong to the same geographical or institutional cluster (e.g., Census division), their outcomes may be correlated across cohorts. The cluster bootstrap at the state level accounts for within-state dependence across time but assumes independence across states. In settings with strong cross-state dependence (e.g., common macro shocks affecting all states simultaneously), the standard errors may be understated. One recommendation is to supplement with randomisation inference, assigning treatment dates randomly to never-treated states and checking whether the estimated ATT exceeds the permutation distribution rambachan2023more.
Lemma (ref) implies that when $m_g > p$, Parallel[$p$] generates $m_g - p$ testable restrictions on pre-treatment data. For cohort $g$, define:
Under the null Parallel[$p$] and the additional assumption that the higher-order differences $\Delta^{p}\hat\gamma_{g,t}$ are asymptotically uncorrelated across $t$, $T_g(p) \xrightarrow{d} \chi^2(m_g-p)$. This uncorrelatedness condition is not automatically satisfied even under serial independence of raw outcomes: applying the $p$-th difference operator to an i.i.d.\ sequence mechanically induces a moving-average dependence of order $p-1$ among the differenced terms. In practice, $T_g(p)$ is therefore best treated as a descriptive diagnostic rather than a formal test; researchers should complement it with the pre-period $R^2$ and visual residual inspection, and rely on the cluster bootstrap for inference. Pooling across cohorts:
In settings with serial correlation or strong cross-cohort dependence, the chi-square approximation may be unreliable and practitioners should rely on permutation-based critical values or treat the statistic as a descriptive diagnostic rather than a formal test. The cluster bootstrap used for inference does not depend on these distributional assumptions.
Let $\alpha$ be the significance level and $p_{\max} = \min_g m_g$.
The procedure is applied independently to each cohort, allowing different cohorts to be identified under different selected orders. The pre-period $R^2$ from the polynomial fit provides a complementary diagnostic: low $R^2$ at a given order signals poor polynomial fit regardless of the test outcome, which can occur when pre-treatment series are short and the F-test has limited power. In practice, $R^2$ and the sensitivity of the aggregate ATT across orders should guide order selection alongside the formal test.
I evaluate DD[$p$] for $p \in \{1,2,3\}$ under three data-generating processes. The panel has $N=300$ units, $T=12$ periods, three cohorts ($g=5,7,9$) with 4, 6, and 8 pre-treatment periods, and 40\,% never-treated. True ATT $= 0.5$ throughout. Results are based on 500 replications.
DGP-1 (True Parallel[1]). $Y_{it}(\infty) = \alpha_i + \lambda_t + \varepsilon_{it}$, $\alpha_i \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,1)$, $\varepsilon_{it} \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,1)$, $\lambda_t$ a common trend. DD[1] is the efficient estimator.
DGP-2 (True Parallel[2], not Parallel[1]). $Y_{it}(\infty) = \alpha_i + \beta_g t + \lambda_t + \varepsilon_{it}$, $\beta_g \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,0.5)$. Cohort-specific linear trends violate Parallel[1]. DD[2] is the correctly specified estimator.
DGP-3 (True Parallel[3]). $Y_{it}(\infty) = \alpha_i + \beta_g t + \gamma_g t^2 + \lambda_t + \varepsilon_{it}$, $\gamma_g \overset{\mathrm{iid}}{\sim} \mathcal{N}(0,0.15)$. Cohort-specific quadratic trends.
Table (ref) reports bias and RMSE; Figure (ref) displays them.
Table (ref) reports order-selection frequencies and empirical coverage of 95\,% confidence intervals after the sequential order-selection procedure, based on 500 simulation draws with $B=99$ cluster bootstrap replications per draw. Each row corresponds to a different data-generating process; the selection frequencies show which order the sequential F-test chooses under each DGP, and the coverage column reports how often the resulting confidence interval contains the true ATT.
Selection frequencies are consistent with the test's design: it selects the lowest order whose pre-treatment implications the data do not reject, defaulting toward parsimony. Coverage after selection is near-nominal for DGP-2 and DGP-3 (93.6% and 93.8%) and slightly below nominal for the flat and borderline DGPs (89.8% and 89.2%), where finite-sample bootstrap imprecision at $B=99$ accounts for the shortfall.
Table (ref) reports two additional robustness checks. Panel A shows that DD[2] maintains near-nominal coverage under AR(1) serial correlation with persistence $\rho$ up to 0.7, ranging from 91.4% at $\rho=0$ to 94.4% at $\rho=0.7$. This reflects that the estimator operates on cohort-level averages, which attenuate within-unit serial correlation. Panel B shows a placebo false-positive rate of 7.2%, modestly above the nominal 5%, consistent with the finite-sample bootstrap behaviour at $B=99$ documented in Table (ref); the mean estimated ATT under the null is essentially zero, confirming the absence of systematic bias.
The simulations above calibrate DGPs to exact polynomial trends. When the pre-treatment gap follows a non-polynomial trajectory, such as a structural break or logistic growth, Parallel[$p$] is misspecified regardless of the chosen order.
The practical diagnostic is the pre-period $R^2$ from the OLS polynomial fit: under a genuine polynomial trend $R^2$ is high; under structural breaks or other non-polynomial dynamics $R^2$ deteriorates noticeably, providing an early warning before the sequential test is applied. See Remark (ref). When the $R^2$ diagnostic signals misspecification, sensitivity bounds rambachan2023more may be more appropriate. Table (ref) shows how the cluster bootstrap standard error for the aggregate DD[2] ATT stabilises as $B$ increases, using the Medicaid expansion application data.
I use a state-year panel covering 46 states observed over 2008--2019, yielding 552 observations.\footnote{Data are drawn from the Mixtape Sessions Advanced DiD repository, made publicly available by Scott Cunningham and accessible at \url{https://raw.githubusercontent.com/Mixtape-Sessions/Advanced-DID/main/Exercises/Data/ehec_data.dta}. The underlying microdata source is the American Community Survey (ACS). } The outcome (dins) is the share of low-income childless adults with health insurance in each state and year, derived from the American Community Survey (ACS). Treatment is Medicaid expansion under the Affordable Care Act of 2010, which gave states the option to expand Medicaid eligibility beginning in 2014. Five expansion cohorts are present: 2014 (22 states, $m_g=6$), 2015 (3, $m_g=7$), 2016 (2, $m_g=8$), 2017 (1, $m_g=9$), and 2019 (2, $m_g=11$). Sixteen states never expanded and serve as the never-treated comparison group throughout. Population weights are normalised to mean one to ensure interpretable test statistics.\footnote{Raw weights average 647,395 per state-year, inflating chi-squared statistics proportionally to the sum of weights rather than the number of observations when used un-normalised. Normalising preserves relative weighting across states.}
The Callaway--Sant'Anna aggregate ATT under Parallel[1] is $\hat\theta_{\mathrm{CS}} = 0.075$ (95\,% CI: [0.051, 0.099], $p < 0.001$). The joint pre-trend test yields $\chi^2(36) = 65{,}921$, $p < 0.001$, decisively rejecting standard parallel trends.\footnote{The large chi-squared reflects importance-weight scaling in csdid; see footnote 1. The qualitative rejection is unaffected.} The pre-treatment gap for the 2014 cohort rises steadily from 0.032 in 2008 to 0.042 in 2013, consistent with a linear trend. I select $p^\star = 2$ as the primary specification based on pre-period $R^2$ diagnostics and sensitivity stability across orders (Table (ref)).
Table (ref) reports DD[1], DD[2], and DD[3] estimates at event times $\tau = 0,\ldots,4$ with 95\,% confidence intervals from the state-level cluster bootstrap ($B=999$). Figure (ref) displays the event-study comparison between DD[1] and the primary DD[2] specification. Figure (ref) extends this to include DD[3] as a robustness check.\footnote{The DD[1] estimates in Table (ref) use the last pre-treatment gap as the flat counterfactual, following the callaway2021difference convention. The unweighted CS simple aggregate (0.068) is close to the DD[1] cohort-weighted average (0.065), since both use the same identifying assumption and the same baseline convention. The weighted CS estimate (0.075) differs because it uses state population weights; the DD[$p$] estimates in this table use equal cohort-share weights.}
The findings show that first, all three estimators find a positive, significant, and growing effect. Insurance coverage rose by approximately 4--5 percentage points in the year of expansion and 7--8 points four years later. This qualitative conclusion is robust across identification strategies.
Second, the DD[1] cohort-weighted average (0.065) is nearly identical to the unweighted Callaway--Sant'Anna simple aggregate (0.068), confirming that DD[1] as implemented here --- using the last pre-treatment gap as the flat counterfactual --- reproduces the CS estimator closely. The gap between the weighted CS estimate (0.075) and these figures reflects population weighting, not the identifying assumption.
Third, DD[2] consistently exceeds DD[1] at every event time, with differences of 0.4--1.0 percentage points (10--22 percent of the DD[1] estimate). This reflects that later-treated cohorts (2015--2019) were on a modest downward trajectory relative to never-expanding states before their expansion. The standard estimator, which assumes a flat counterfactual, underestimates the treatment effect for these cohorts. DD[2] corrects for this by projecting the downward pre-trend forward.
Fourth, the confidence intervals of DD[1] and DD[2] overlap substantially at every event time: I cannot reject the null that the two estimators produce the same treatment effect. The contribution of DD[2] in this application is therefore not a reversal of the standard finding but a more credible quantification of it. The same qualitative conclusion, Medicaid expansion increased insurance coverage, is supported by an assumption the pre-treatment data do not reject, rather than one they reject.
Table (ref) summarises estimates across all three specifications and records which pass the pre-trend test.
Table (ref) shows the aggregate ATT under three weighting strategies. Results are insensitive to weighting choice: DD[2] ranges from 0.064 to 0.067 and DD[1] from 0.060 to 0.065. The cohort-level results in the penultimate row reveal that the direction of the DD[2] correction varies by cohort: for the large 2014 cohort, DD[2] is slightly smaller than DD[1] (its pre-trend was upward), while for later cohorts DD[2] exceeds DD[1] (their pre-trends were downward). This heterogeneity in direction is precisely what motivates cohort-specific counterfactuals.
A natural benchmark is the sensitivity bounds approach of rambachan2023more, which restricts the magnitude of parallel trends violations and reports how estimated effects vary within that restriction. Their approach delivers valid intervals whose width reflects the assumed maximum violation magnitude. Notably, rambachan2023more use the same dataset as an illustrative example in their software package HonestDiD. Their smoothness restriction at $M=0$ imposes that the post-treatment counterfactual gap follows a linear extrapolation of the pre-existing trend which is conceptually equivalent to DD[2] at Parallel[2]. For $M > 0$, their approach allows deviations from linearity up to a bound of size $M$, providing a natural complement to the point estimates reported here: DD[2] delivers a point estimate under the polynomial structure, while rambachan2023more deliver robust intervals that remain valid even if the linear structure is imperfect. DD[2] complements this analysis by showing that if the specific polynomial structure is credible, as the pre-treatment $R^2$ values and visual evidence suggest for most cohorts, one can achieve point identification without widening the bounds. The two approaches are not competitors: if the $R^2$ diagnostic signals misspecification, the rambachan2023more bounds remain the appropriate tool.
When pre-treatment event studies reject standard parallel trends, the applied researcher faces a difficult choice: impose a maintained assumption the data contradict, or report sensitivity bounds that are valid but do not deliver a point estimate. This paper offers a third option: replacing the flatness requirement with a structured polynomial extrapolation whose pre-treatment implications are directly checkable.
The central theoretical contribution is Theorem (ref) an aggregation result allowing cohorts to be identified under different feasible orders, arising naturally from variation in pre-treatment period counts across staggered adoption cohorts. This problem has no analogue in two-group settings and is not addressed by any existing paper, including egami2023using, who extend higher-order parallelism to staggered designs at $p=2$ under a uniform order.
Monte Carlo evidence shows near-nominal coverage when the correct order is selected, robustness to AR(1) serial correlation, and a placebo false positive rate modestly above the nominal 5\,%, consistent with finite-sample bootstrap behaviour at $B=99$. Applied to Medicaid expansion, the approach yields estimates resting on an assumption the pre-treatment data do not reject.
Two directions merit future work. First, a formal uniform validity result for the complete selection-estimation pipeline would strengthen the inferential foundations. The extended version of roth2022pretest on post-selection inference provides a starting point, though adapting those results to the sequential polynomial order selection here requires additional work. Second, replacing the polynomial order restriction with derivative-bounded smoothness restrictions rambachan2023more would extend the framework to settings where the pre-treatment gap does not follow a polynomial, at the cost of point identification.