EconBase
← All papers

Efficient Policy Learning from Surrogate-Loss Classification Reductions

Andrew Bennett, Nathan Kallus

arXiv 12 Feb 2020 · Machine Learning · 7 citations (OpenAlex)

arXiv:2002.05153 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Recent work on policy learning from observational data has highlighted the importance of efficient policy evaluation and has proposed reductions to weighted (cost-sensitive) classification. But, efficient policy evaluation need not yield efficient estimation of policy parameters. We consider the estimation problem given by a weighted surrogate-loss classification reduction of policy learning with any score function, either direct, inverse-propensity weighted, or doubly robust. We show that, under a correct specification assumption, the weighted classification formulation need not be efficient for policy parameters. We draw a contrast to actual (possibly weighted) binary classification, where correct specification implies a parametric model, while for policy learning it only implies a semiparametric model. In light of this, we instead propose an estimation approach based on generalized method of moments, which is efficient for the policy parameters. We propose a particular method based on recent developments on solving moment problems using neural networks and demonstrate the efficiency and regret benefits of this method empirically.

Citation extraction

28
references
71
in-text mentions
28
distinct cited
4
self-citations
7,035
main-text words

appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Jiang, B., Song, R., Li, J., and Zeng, D (2019) Entropy learning for dynamic treatment regimes1.00083100%
2Athey, S. and Wager, S (2017) Efficient policy learning1.00053100%
3Bennett, A., Kallus, N., and Schnabel, T (2019) Deep generalized method of moments for instrumental variable analysis self0.87472100%
4Zhou, X., Mayer-Hamblett, N., Khan, U., and Kosorok, M. R (2017) Residual weighted learning for estimating individualized treatment rules0.87452100%
5Beygelzimer, A. and Langford, J (2009) The offset tree for learning with partial labels0.81142100%
6Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R (2012) Estimating individualized treatment rules using outcome weighted learning0.81142100%
7Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.73732100%
8Zhou, Z., Athey, S., and Wager, S (2018) Offline multi-action policy learning: Generalization and optimization0.73732100%
9Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning0.58531100%
10Van der Vaart, A. W (2000) Asymptotic statistics0.5114225%

Showing the top 10 of 28 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Policy Learning with Adaptively Collected Data0.40511