Edward McFowland III, Sriram Somanchi, Daniel B. Neill
arXiv 24 Mar 2018 · Statistics — Methodology · 11 citations (OpenAlex)
arXiv:1803.09159 · PDF · DOI · OpenAlex · Extracted main text
In the recent literature on estimating heterogeneous treatment effects, each proposed method makes its own set of restrictive assumptions about the intervention's effects and which subpopulations to explicitly estimate. Moreover, the majority of the literature provides no mechanism to identify which subpopulations are the most affected--beyond manual inspection--and provides little guarantee on the correctness of the identified subpopulations. Therefore, we propose Treatment Effect Subset Scan (TESS), a new method for discovering which subpopulation in a randomized experiment is most significantly affected by a treatment. We frame this challenge as a pattern detection problem where we efficiently maximize a nonparametric scan statistic (a measure of the conditional quantile treatment effect) over subpopulations. Furthermore, we identify the subpopulation which experiences the largest distributional change as a result of the intervention, while making minimal assumptions about the intervention's effects or the underlying data generating process. In addition to the algorithm, we demonstrate that under the sharp null hypothesis of no treatment effect, the asymptotic Type I and II error can be controlled, and provide sufficient conditions for detection consistency--i.e., exact identification of the affected subpopulation. Finally, we validate the efficacy of the method by discovering heterogeneous treatment effects in simulations and in real-world data from a well-known program evaluation study.
appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | E. McFowland III, S. D. Speakman, and D. B. Neill Fast generalized subset scan for anomalous pattern detection self | 1.000 | 8 | 3 | 100% |
| 2 | S. Athey and G. Imbens (2016) Recursive partitioning for heterogeneous causal effects | 0.874 | 5 | 2 | 100% |
| 3 | S. Wager and S. Athey (2018) Estimation and inference of heterogeneous treatment effects using random forests | 0.811 | 4 | 2 | 100% |
| 4 | X. Su, C.-L. Tsai, H. Wang, D. M. Nickerson, and B. Li (2009) Subgroup analysis via recursive partitioning | 0.737 | 3 | 2 | 100% |
| 5 | A. B. Krueger (1999) Experimental estimates of education production functions | 0.693 | 8 | 1 | 100% |
| 6 | E. R. Word, J. Johnston, H. P. Bain, and Others (1990) The state of Tennessee's Student/Teacher Achievement Ratio (STAR) project: Technical report 1985–1990 | 0.693 | 6 | 1 | 100% |
| 7 | F. Chen and D. B. Neill (2014) Non-parametric scan statistics for event detection and forecasting in heterogeneous social media graphs | 0.644 | 2 | 2 | 100% |
| 8 | H. I. Weisberg and V. P. Pontes (2015) Post hoc subgroups in clinical trials: Anathema or analytics? | 0.644 | 2 | 2 | 100% |
| 9 | V. Chernozhukov, I. Fernández-Val, and B. Melly (2013) Inference on counterfactual distributions | 0.585 | 3 | 1 | 100% |
| 10 | J. Folger and C. Breda (1989) Evidence from project STAR about class size and student achievement | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 54 scored citations.