Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
21,839 characters · 10 sections · 11 citation commands
Doubly Robust Local Projections Difference-in-Differences
\noindentJEL codes: C14, C21, C23, C12
Event-study designs are standard tools for tracing policy dynamics. In staggered-adoption settings with heterogeneous treatment effects, however, the usual two-way fixed-effects event-study regression need not recover a causal path because it can mix already-treated comparisons and non-transparent weights deChaisemartinDHaultfoeuille2020,GoodmanBacon2021,CallawaySantAnna2021,SunAbraham2021,BorusyakJaravelSpiess2024,Gardner2022,deChaisemartinDHaultfoeuille2023,Rothetal2023.
LP-DiD estimates horizon-specific effects Jorda2005LocalProjections on clean-control stacks that exclude unclean comparisons, thereby avoiding contamination from already-treated controls DubeGirardiJordaTaylor2025. But each horizon-specific stack remains an observational comparison. When untreated outcome dynamics or treatment timing depend on covariates, the horizon-specific LP-DiD problem can be cast as a semiparametric Average Treatment Effect on the Treated (ATT) problem under conditional parallel trends DubeGirardiJordaTaylor2025. Although doubly robust estimators for staggered adoption already exist CallawaySantAnna2021,SantAnnaZhao2020, they operate on the full panel rather than on horizon-specific clean-control stacks; a researcher committed to the LP-DiD framework, therefore, needs nuisance adjustment derived within that framework, with local propensity scores and outcome regressions that are stack- and horizon-specific by construction.
This article extends the LP-DiD framework by incorporating robustness to misspecification through a doubly robust approach in staggered adoption settings. It shows that the existing LP-DiD horizon problem can be enriched with Regression-Adjusted (RA), Inverse-Probability-Tilting (IPT), and Doubly Robust (DR) structures that admit influence-function-based inference and multiplier-bootstrap bands Hajek1971,GrahamPintoEgel2012,SantAnnaZhao2020,CallawaySantAnna2021. The Monte Carlo compares DRLPDID with LP-DiD benchmarks and its RA and IPT counterparts across nuisance-specification scenarios designed to diagnose the value of double robustness. The application then shows how the resulting dynamic path compares with leading staggered-adoption estimators in practice. The appendix provides the general base-period formulation, proofs, and supplementary results.
LP-DiD compares post-entry outcomes with a pre-treatment benchmark constructed using a specific base-period rule, indexed by $b$. For a given entry date $t$, let $\mathcal{B}(t;b) \subset \{1, \dots, t-1\}$ denote the set of active pre-treatment periods and let $\omega_{tl}^{(b)}$ be weights summing to one. The generic benchmark operator is defined as:
This paper focuses on the standard last-pre-treatment benchmark (denoted by $b = -1$), which assigns a weight of one to period $t-1$ and zero otherwise, yielding $B_{it}^{(-1)} = Y_{i,t-1}$. The horizon-$h$ long difference is therefore:
Write the local treatment-entry indicator as
Here $D_{it}$ denotes local treatment entry, not calendar-time treatment status. For $h\geq 0$, the clean-control stack compares treated entries at $t$ with observations untreated at $t+h$. For $h<0$, it reuses the clean-control stack at $h=0$ and changes only the outcome transformation.
The horizon-specific local target is
Let $X_{it}$ be predetermined covariates. For each $h$, assume no anticipation or overlap:
and local conditional parallel trends,
Then each horizon is an ATT problem on its own LP-DiD stack.
Define the local nuisance functions
In the implemented estimators, both nuisance steps use the same horizon-specific linear basis $Q_{it,h}$, built from the same predetermined covariates and calendar-time controls; the difference is the fitting criterion, with inverse-probability tilting for the treatment-probability working model and least squares or weighted least squares for the untreated-outcome regression.
The local regression-adjusted and inverse-probability-weighting representations are
where $\omega_h(X_{it},t)=\pi_h(X_{it},t)/(1-\pi_h(X_{it},t))$. Combining them yields
Under (ref), (ref) identifies $\theta_h$ if either $m_{0h}$ or $\pi_h$ is correctly specified. Here, IPW refers to the population weighting identity in (ref), whereas the implemented weighted estimator obtains $\pi_h$ by inverse-probability tilting (IPT). DRLPDID estimates (ref) horizon by horizon using inverse-probability tilting for the local propensity score GrahamPintoEgel2012 and weighted least squares for the untreated-outcome regression, following the improved doubly robust logic of SantAnnaZhao2020 in an LP-DiD environment.
Let $\hat\Theta=(\hat\theta_h)_{h\in\mathcal H}$ collect the horizon-specific estimates. For each $h$, the improved estimator solves stacked local moments for the outcome-regression coefficients, the tilting parameters, and the treated and weighted-control means. Standard M-estimation then gives
where $N_C$ is the number of clusters. Equation (ref) yields cluster-robust standard errors for each horizon and for any linear contrast built from the horizon-specific effects.
For dynamic paths, the paper uses a multiplier bootstrap based on the stacked influence array. If $\hat\Phi_c=(\hat\mathrm{IF}_{c,h})_{h\in\mathcal H}$ and $\xi_c$ are i.i.d. multipliers with mean zero and variance one, bootstrap draws are
The main text reports pointwise cluster-robust intervals in the Monte Carlo and multiplier-bootstrap bands in the empirical event-study figures. The appendix provides the full moment system and proofs.
The Monte Carlo uses a staggered-adoption panel with covariate-driven treatment timing, heterogeneous dynamic treatment effects, and a never-treated group. The complete data-generating process is reported in the appendix. We compare five ATT-oriented estimators that share the same treated-entry-weighted target: LPDID-RW, LPDID-RW + X, LPDID-RA, DRLPDID-IPT, and DRLPDID. The designs vary regarding which nuisance component is aligned with the DGP: both nuisance components aligned, outcome-regression misspecified, propensity-score misspecified, and both nuisance components misspecified. Table (ref) reports bias, RMSE, and pointwise 95% coverage for the average post-treatment effect over 500 replications at $N=500$; the corresponding $N=250$ results are reported in Appendix Table (ref).
Three patterns stand out. First, LPDID-RW is consistently unstable across scenarios once local selection and untreated trends depend on covariates. Second, when the outcome-regression component is aligned with the DGP, LPDID-RA remains highly competitive, and DRLPDID stays very close to it, rather than uniformly dominating it. Third, propensity-score misspecification is the setting in which the doubly robust construction is most informative: DRLPDID stays close to LPDID-RA, whereas DRLPDID-IPT deteriorates markedly. The low coverage of DRLPDID-IPT even in the aligned design appears to reflect bias rather than underestimated uncertainty: supplementary diagnostics show that its cluster-robust standard errors track the empirical dispersion closely, but the pure IPT step remains centered below the treated-entry-weighted target. This is consistent with the fact that covariate-balancing odds weights do not by themselves recover the cohort-size treated-entry weighting embedded in $\theta_h$, whereas the regression adjustment in DRLPDID corrects that discrepancy. Under joint misspecification, DRLPDID remains among the best scalar estimators, especially at $N=500$, but the Monte Carlo does not support a claim of uniform dominance in every design.
The application uses the no-fault-divorce panel from StevensonWolfers2006, as revisited by GoodmanBacon2021. The outcome, asmrs, is the state-level female suicide rate; treatment is the enactment of a no-fault divorce law. The sample covers the 48 continental states plus Washington, D.C., from 1964 to 1996. The baseline adjusted specification includes asmrh, the homicide mortality rate.\footnote{The original application also considers pcinc and cases. We omit them from the main specification because they may respond to the legal environment. The appendix reports a richer specification.}
Table (ref) reports average post-treatment effects under the parsimonious specification. DRLPDID yields $-5.69$, very close to Borusyak--Jaravel--Spiess and Gardner (both $-5.83$), less negative than Callaway--Sant'Anna ($-6.81$), and materially less negative than the two local LP-DiD comparators LPDID-RA ($-6.31$) and especially LPDID-RW + X ($-9.30$). Sun--Abraham remains clearly more negative at $-8.03$. The application therefore mirrors the Monte Carlo: covariate adjustment moves local LP-DiD toward the robust staggered-adoption comparators, and the doubly robust version aligns most closely with the imputation-based estimators Borusyak--Jaravel--Spiess and Gardner.
Figure (ref) shows the dynamic paths. DRLPDID remains close to LPDID-RA and also tracks Borusyak--Jaravel--Spiess and Gardner more closely than the more negative LPDID-RW + X and Sun--Abraham paths over most post-treatment horizons. Callaway--Sant'Anna is also broadly in the same range, though somewhat more negative in the scalar summary.
DRLPDID extends LP-DiD with horizon-specific regression adjustment, inverse-probability tilting, and doubly robust moments without changing the local-stack target. Under the last pre-treatment benchmark used here, that target has the usual horizon-specific ATT interpretation. In the Monte Carlo, DRLPDID is consistently competitive and is especially well behaved under propensity-score misspecification, while LPDID-RA remains highly competitive when the outcome-regression component is well aligned with the DGP. In the no-fault-divorce application, DRLPDID delivers dynamic effects close to LPDID-RA and to the main robust staggered-adoption estimators, while remaining materially less negative than the IPW-only local path and the unadjusted LPDID-RW benchmark.
The practical implication is simple. When covariates plausibly shape local selection and untreated dynamics, unadjusted LP-DiD should not be the default specification. Extensions to continuous treatments, alternative base-period constructions, and stabilized treated controls are left for future work.
This work was supported by the Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), grants 301475/2025-3 and 310327/2022-9.
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
The data and simulation materials underlying this article are available from the corresponding author upon reasonable request by email at [email removed].
During the preparation of this work, the authors used ChatGPT and Claude Code for English-language editing and structural consistency checks. After using these tools, the authors reviewed and edited the manuscript as needed and take full responsibility for the content of the published article.