EconBase
← All papers

Efficient Discovery of Heterogeneous Quantile Treatment Effects in Randomized Experiments via Anomalous Pattern Detection

Edward McFowland III, Sriram Somanchi, Daniel B. Neill

arXiv 24 Mar 2018 · Statistics — Methodology · 11 citations (OpenAlex)

arXiv:1803.09159 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In the recent literature on estimating heterogeneous treatment effects, each proposed method makes its own set of restrictive assumptions about the intervention's effects and which subpopulations to explicitly estimate. Moreover, the majority of the literature provides no mechanism to identify which subpopulations are the most affected--beyond manual inspection--and provides little guarantee on the correctness of the identified subpopulations. Therefore, we propose Treatment Effect Subset Scan (TESS), a new method for discovering which subpopulation in a randomized experiment is most significantly affected by a treatment. We frame this challenge as a pattern detection problem where we efficiently maximize a nonparametric scan statistic (a measure of the conditional quantile treatment effect) over subpopulations. Furthermore, we identify the subpopulation which experiences the largest distributional change as a result of the intervention, while making minimal assumptions about the intervention's effects or the underlying data generating process. In addition to the algorithm, we demonstrate that under the sharp null hypothesis of no treatment effect, the asymptotic Type I and II error can be controlled, and provide sufficient conditions for detection consistency--i.e., exact identification of the affected subpopulation. Finally, we validate the efficacy of the method by discovering heterogeneous treatment effects in simulations and in real-world data from a well-known program evaluation study.

Citation extraction

54
references
99
in-text mentions
54
distinct cited
3
self-citations
16,938
main-text words

appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1E. McFowland III, S. D. Speakman, and D. B. Neill Fast generalized subset scan for anomalous pattern detection self1.00083100%
2S. Athey and G. Imbens (2016) Recursive partitioning for heterogeneous causal effects0.87452100%
3S. Wager and S. Athey (2018) Estimation and inference of heterogeneous treatment effects using random forests0.81142100%
4X. Su, C.-L. Tsai, H. Wang, D. M. Nickerson, and B. Li (2009) Subgroup analysis via recursive partitioning0.73732100%
5A. B. Krueger (1999) Experimental estimates of education production functions0.69381100%
6E. R. Word, J. Johnston, H. P. Bain, and Others (1990) The state of Tennessee's Student/Teacher Achievement Ratio (STAR) project: Technical report 1985–19900.69361100%
7F. Chen and D. B. Neill (2014) Non-parametric scan statistics for event detection and forecasting in heterogeneous social media graphs0.64422100%
8H. I. Weisberg and V. P. Pontes (2015) Post hoc subgroups in clinical trials: Anathema or analytics?0.64422100%
9V. Chernozhukov, I. Fernández-Val, and B. Melly (2013) Inference on counterfactual distributions0.58531100%
10J. Folger and C. Breda (1989) Evidence from project STAR about class size and student achievement0.58531100%

Showing the top 10 of 54 scored citations.