arXiv 12 Feb 2025 · Econometrics
arXiv:2502.08614 · PDF · DOI · OpenAlex · Extracted main text
Sample selection arises endogenously in causal research when the treatment affects whether certain units are observed. It is a common pitfall in longitudinal studies, particularly in settings where treatment assignment is confounded. In this paper, I highlight the drawbacks of one of the most popular identification strategies in such settings: Difference-in-Differences (DiD). Specifically, I employ principal stratification analysis to show that the conventional ATT estimand may not be well defined, and the DiD estimand cannot be interpreted causally without additional assumptions. To address these issues, I develop an identification strategy to partially identify causal effects on the subset of units with well-defined and observed outcomes under both treatment regimes. I adapt Lee bounds to the Changes-in-Changes (CiC) setting (Athey & Imbens, 2006), leveraging the time dimension of the data to relax the unconfoundedness assumption in the original trimming strategy of Lee (2009). This setting has the DiD identification strategy as a particular case, which I also implement in the paper. Additionally, I explore how to leverage multiple sources of sample selection to relax the monotonicity assumption in Lee (2009), which may be of independent interest. Alongside the identification strategy, I present estimators and inference results. I illustrate the relevance of the proposed methodology by analyzing a job training program in Colombia.
appendix boundary found by appendix_command · 42% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Attanasio, Orazio and Kugler, Adriana and Meghir, Costas Subsidizing Vocational Training for Disadvantaged Youth in Colombia: Evidence from a Randomized Trial | 0.874 | 5 | 2 | 100% |
| 2 | Lee, David S Training, Wages, and Sample Selection: Estimating Sharp Bounds on Treatment Effects | 0.822 | 9 | 5 | 56% |
| 3 | Rathnayake, Gayani and Negi, Akanksha and Bartalotti, Otavio and Zha… Difference-in-Differences with Sample Selection | 0.811 | 4 | 2 | 100% |
| 4 | Athey, Susan and Imbens, Guido W Identification and Inference in Nonlinear Difference-in-Differences Models | 0.751 | 26 | 5 | 42% |
| 5 | Grilli, Leonardo and Mealli, Fabrizia Nonparametric Bounds on the Causal Effect of University Studies on Job Opportunities Using Principal Stratification | 0.737 | 4 | 3 | 50% |
| 6 | Long, Dustin M. and Hudgens, Michael G Sharpening Bounds on Principal Effects with Covariates | 0.737 | 4 | 3 | 50% |
| 7 | Shin, Sooahn Difference-in-differences Design with Outcomes Missing Not at Random | 0.737 | 3 | 2 | 100% |
| 8 | Horowitz, Joel L. and Manski, Charles F Identification and Robustness with Contaminated and Corrupted Data | 0.721 | 8 | 3 | 38% |
| 9 | Angrist, Joshua D. and Pischke, Jörn-Steffen Mostly Harmless Econometrics: An Empiricist's Companion | 0.644 | 2 | 2 | 100% |
| 10 | Arkhangelsky, Dmitry and Imbens, Guido Causal models for longitudinal and panel data: a survey | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 65 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Difference-in-Differences with Sample Selection | 0.644 | 4 | 1 |
| 2 | Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random | 0.405 | 1 | 1 |