Lorenzo Testa, Edward H. Kennedy, Matthew Reimherr
arXiv 29 Sep 2025 · Statistics — Methodology
arXiv:2509.25009 · PDF · Extracted main text
The Difference-in-Differences (DiD) method is a fundamental tool for causal inference, yet its application is often complicated by missing data. Although recent work has developed robust DiD estimators for complex settings like staggered treatment adoption, these methods typically assume complete data and fail to address the critical challenge of outcomes that are missing at random (MAR) -- a common problem that invalidates standard estimators. We develop a rigorous framework, rooted in semiparametric theory, for identifying and efficiently estimating the Average Treatment Effect on the Treated (ATT) when either pre- or post-treatment (or both) outcomes are missing at random. We first establish nonparametric identification of the ATT under two minimal sets of sufficient conditions. For each, we derive the semiparametric efficiency bound, which provides a formal benchmark for asymptotic optimality. We then propose novel estimators that are asymptotically efficient, achieving this theoretical bound. A key feature of our estimators is their multiple robustness, which ensures consistency even if some nuisance function models are misspecified. We validate the properties of our estimators and showcase their broad applicability through an extensive simulation study.
appendix boundary found by appendix_command · 46% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sant’Anna, Pedro HC and Zhao, Jun (2020) Doubly robust difference-in-differences estimators | 0.843 | 4 | 4 | 75% |
| 2 | Kennedy, Edward H (2024) Semiparametric doubly robust targeted double machine learning: a review self | 0.737 | 3 | 3 | 67% |
| 3 | Tsiatis, Anastasios A (2006) Semiparametric theory and missing data | 0.737 | 3 | 2 | 100% |
| 4 | Kennedy, Edward H (2023) Towards optimal doubly robust estimation of heterogeneous causal effects self | 0.644 | 3 | 2 | 67% |
| 5 | Bickel, Peter J and Klaassen, Chris AJ and Ritov, Ya’acov and Wellne… (1993) Efficient and adaptive estimation for semiparametric models | 0.511 | 2 | 2 | 50% |
| 6 | Callaway, Brantly and Sant’Anna, Pedro HC (2021) Difference-in-differences with multiple time periods | 0.511 | 2 | 1 | 100% |
| 7 | Bellégo, Christophe and Benatia, David and Dortet-Bernadet, Vincent (2025) The chained difference-in-differences | 0.405 | 1 | 1 | 100% |
| 8 | Bickel, Peter J and Ritov, Yaacov (1988) Estimating integrated squared density derivatives: sharp best order of convergence estimates | 0.405 | 1 | 1 | 100% |
| 9 | Callaway, Brantly (2023) Difference-in-differences for policy evaluation | 0.405 | 1 | 1 | 100% |
| 10 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 38 scored citations.