arXiv 27 Nov 2024 · Statistics — Methodology
arXiv:2411.18772 · PDF · DOI · OpenAlex · Extracted main text
This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case analysis, is to drop cases with any missing values over time. A more principled approach involves using nonparametric bounds on causal effects or applying inverse probability weighting based on baseline covariates. Yet, these methods are general remedies that often under-utilize the assumptions already imposed on panel structure for causal identification. In this paper, I outline the pitfalls of complete case analysis and propose an alternative identification strategy based on principal strata. To be specific, I impose parallel trends assumption within each latent group that shares the same missingness pattern (e.g., always-respondents, if-treated-respondents) and leverage missingness rates over time to estimate the proportions of these groups. Building on this, I tailor Lee bounds, a well-known nonparametric bounds under selection bias, to partially identify the causal effect within the DID design. Unlike complete case analysis, the proposed method does not require independence between treatment selection and missingness patterns, nor does it assume homogeneous effects across these patterns.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Dukes, Richardson \ Tchetgen Tchetgen (2022) Alternative approaches for analysing repeated measures data that are missing not at random | 1.000 | 6 | 3 | 100% |
| 2 | Lee (2009) Training, wages, and sample selection: Estimating sharp bounds on treatment effects | 1.000 | 6 | 3 | 100% |
| 3 | Sexton \ Zürcher (2024) Aid, Attitudes, and Insurgency: Evidence from Development Projects in Northern Afghanistan | 0.737 | 3 | 2 | 100% |
| 4 | Zhang \ Rubin (2003) Estimation of causal effects via principal stratification when some outcomes are truncated by “death” | 0.737 | 3 | 2 | 100% |
| 5 | Frangakis \ Rubin (2002) Principal stratification in causal inference | 0.644 | 2 | 2 | 100% |
| 6 | Wang, Zhou \ Richardson (2017) Identification and estimation of causal effects with outcomes truncated by death | 0.644 | 2 | 2 | 100% |
| 7 | Tchetgen Tchetgen \ Wirth (2017) A general instrumental variable framework for regression analysis with outcome missing not at random | 0.585 | 3 | 1 | 100% |
| 8 | Bisgaard \ Slothuus (2018) Partisan Elites as Culprits? How Party Cues Shape Partisan Perceptual Gaps | 0.511 | 2 | 1 | 100% |
| 9 | Ghanem, Hirshleifer, Kedagni \ Ortiz-Becerra (2022) Correcting Attrition Bias using Changes-in-Changes | 0.511 | 2 | 1 | 100% |
| 10 | Chiu, Lan, Liu \ Xu (2023) What to do (and not to do) with causal panel analysis under parallel trends: Lessons from a large reanalysis study | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 15 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.