Masahiro Kato, Masatoshi Uehara, Shota Yasui
arXiv 26 Feb 2020 · Statistics — Machine Learning · 12 citations (OpenAlex)
arXiv:2002.11642 · PDF · DOI · OpenAlex · Extracted main text
We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, a covariate shift often exists, i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using a nonparametric estimator of the density ratio between the historical and evaluation data distributions. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments.
appendix boundary found by appendix_command · 36% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Pearl, J. and Bareinboim, E (2014) External validity: From do-calculus to transportability across populations | 0.928 | 5 | 3 | 80% |
| 2 | Kanamori, T., Suzuki, T., and Sugiyama, M (2012) Statistical analysis of kernel-based least-squares density-ratio estimation | 0.794 | 6 | 4 | 50% |
| 3 | van der Vaart, A. W (1998) Asymptotic statistics | 0.737 | 10 | 4 | 40% |
| 4 | Rubin, D. B (1987) Multiple Imputation for Nonresponse in Surveys | 0.644 | 3 | 2 | 67% |
| 5 | Athey, S. and Wager, S (2017) Efficient policy learning | 0.644 | 2 | 2 | 100% |
| 6 | Dahabreh, I. J., Robertson, S. E., Tchetgen, E. J., Stuart, E. A., a… (2019) Generalizing causal inferences from individuals in randomized trials to all trial‐eligible individuals | 0.644 | 2 | 2 | 100% |
| 7 | Kallus, N. and Uehara, M (2019) Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning self | 0.644 | 2 | 2 | 100% |
| 8 | Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.644 | 2 | 2 | 100% |
| 9 | Dudḱ, M., Langford, J., and Li, L (2011) Doubly Robust Policy Evaluation and Learning | 0.644 | 2 | 2 | 100% |
| 10 | Narita, Y., Yasui, S., and Yata, K (2019) Efficient counterfactual learning from bandit feedback self | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 58 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.