arXiv 12 Feb 2020 · Machine Learning · 7 citations (OpenAlex)
arXiv:2002.05153 · PDF · DOI · OpenAlex · Extracted main text
Recent work on policy learning from observational data has highlighted the importance of efficient policy evaluation and has proposed reductions to weighted (cost-sensitive) classification. But, efficient policy evaluation need not yield efficient estimation of policy parameters. We consider the estimation problem given by a weighted surrogate-loss classification reduction of policy learning with any score function, either direct, inverse-propensity weighted, or doubly robust. We show that, under a correct specification assumption, the weighted classification formulation need not be efficient for policy parameters. We draw a contrast to actual (possibly weighted) binary classification, where correct specification implies a parametric model, while for policy learning it only implies a semiparametric model. In light of this, we instead propose an estimation approach based on generalized method of moments, which is efficient for the policy parameters. We propose a particular method based on recent developments on solving moment problems using neural networks and demonstrate the efficiency and regret benefits of this method empirically.
appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Jiang, B., Song, R., Li, J., and Zeng, D (2019) Entropy learning for dynamic treatment regimes | 1.000 | 8 | 3 | 100% |
| 2 | Athey, S. and Wager, S (2017) Efficient policy learning | 1.000 | 5 | 3 | 100% |
| 3 | Bennett, A., Kallus, N., and Schnabel, T (2019) Deep generalized method of moments for instrumental variable analysis self | 0.874 | 7 | 2 | 100% |
| 4 | Zhou, X., Mayer-Hamblett, N., Khan, U., and Kosorok, M. R (2017) Residual weighted learning for estimating individualized treatment rules | 0.874 | 5 | 2 | 100% |
| 5 | Beygelzimer, A. and Langford, J (2009) The offset tree for learning with partial labels | 0.811 | 4 | 2 | 100% |
| 6 | Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R (2012) Estimating individualized treatment rules using outcome weighted learning | 0.811 | 4 | 2 | 100% |
| 7 | Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.737 | 3 | 2 | 100% |
| 8 | Zhou, Z., Athey, S., and Wager, S (2018) Offline multi-action policy learning: Generalization and optimization | 0.737 | 3 | 2 | 100% |
| 9 | Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning | 0.585 | 3 | 1 | 100% |
| 10 | Van der Vaart, A. W (2000) Asymptotic statistics | 0.511 | 4 | 2 | 25% |
Showing the top 10 of 28 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Policy Learning with Adaptively Collected Data | 0.405 | 1 | 1 |