EconBase
← All papers

Difference-in-differences Design with Outcomes Missing Not at Random

Sooahn Shin

arXiv 27 Nov 2024 · Statistics — Methodology

arXiv:2411.18772 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case analysis, is to drop cases with any missing values over time. A more principled approach involves using nonparametric bounds on causal effects or applying inverse probability weighting based on baseline covariates. Yet, these methods are general remedies that often under-utilize the assumptions already imposed on panel structure for causal identification. In this paper, I outline the pitfalls of complete case analysis and propose an alternative identification strategy based on principal strata. To be specific, I impose parallel trends assumption within each latent group that shares the same missingness pattern (e.g., always-respondents, if-treated-respondents) and leverage missingness rates over time to estimate the proportions of these groups. Building on this, I tailor Lee bounds, a well-known nonparametric bounds under selection bias, to partially identify the causal effect within the DID design. Unlike complete case analysis, the proposed method does not require independence between treatment selection and missingness patterns, nor does it assume homogeneous effects across these patterns.

Citation extraction

15
references
35
in-text mentions
15
distinct cited
0
self-citations
9,922
main-text words

appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Dukes, Richardson \ Tchetgen Tchetgen (2022) Alternative approaches for analysing repeated measures data that are missing not at random1.00063100%
2Lee (2009) Training, wages, and sample selection: Estimating sharp bounds on treatment effects1.00063100%
3Sexton \ Zürcher (2024) Aid, Attitudes, and Insurgency: Evidence from Development Projects in Northern Afghanistan0.73732100%
4Zhang \ Rubin (2003) Estimation of causal effects via principal stratification when some outcomes are truncated by “death”0.73732100%
5Frangakis \ Rubin (2002) Principal stratification in causal inference0.64422100%
6Wang, Zhou \ Richardson (2017) Identification and estimation of causal effects with outcomes truncated by death0.64422100%
7Tchetgen Tchetgen \ Wirth (2017) A general instrumental variable framework for regression analysis with outcome missing not at random0.58531100%
8Bisgaard \ Slothuus (2018) Partisan Elites as Culprits? How Party Cues Shape Partisan Perceptual Gaps0.51121100%
9Ghanem, Hirshleifer, Kedagni \ Ortiz-Becerra (2022) Correcting Attrition Bias using Changes-in-Changes0.51121100%
10Chiu, Lan, Liu \ Xu (2023) What to do (and not to do) with causal panel analysis under parallel trends: Lessons from a large reanalysis study0.40511100%

Showing the top 10 of 15 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Difference-in-Differences with Sample Selection0.81142
2Semiparametric Difference-in-Differences Estimation With Missing Not at Random Data: A Shadow Variable Approach0.73743
3Estimating the Intensive Margin Effect in Panel Data Settings0.73732
4Conformal Inference for Experimental Attrition in Social Science Research0.51121
5Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random0.40511