EconBase
← All papers

Off-Policy Evaluation and Learning for External Validity under a Covariate Shift

Masahiro Kato, Masatoshi Uehara, Shota Yasui

arXiv 26 Feb 2020 · Statistics — Machine Learning · 12 citations (OpenAlex)

arXiv:2002.11642 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, a covariate shift often exists, i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using a nonparametric estimator of the density ratio between the historical and evaluation data distributions. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments.

Citation extraction

58
references
98
in-text mentions
58
distinct cited
3
self-citations
7,943
main-text words

appendix boundary found by appendix_command · 36% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Pearl, J. and Bareinboim, E (2014) External validity: From do-calculus to transportability across populations0.9285380%
2Kanamori, T., Suzuki, T., and Sugiyama, M (2012) Statistical analysis of kernel-based least-squares density-ratio estimation0.7946450%
3van der Vaart, A. W (1998) Asymptotic statistics0.73710440%
4Rubin, D. B (1987) Multiple Imputation for Nonresponse in Surveys0.6443267%
5Athey, S. and Wager, S (2017) Efficient policy learning0.64422100%
6Dahabreh, I. J., Robertson, S. E., Tchetgen, E. J., Stuart, E. A., a… (2019) Generalizing causal inferences from individuals in randomized trials to all trial‐eligible individuals0.64422100%
7Kallus, N. and Uehara, M (2019) Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning self0.64422100%
8Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.64422100%
9Dudḱ, M., Langford, J., and Li, L (2011) Doubly Robust Policy Evaluation and Learning0.64422100%
10Narita, Y., Yasui, S., and Yata, K (2019) Efficient counterfactual learning from bandit feedback self0.64422100%

Showing the top 10 of 58 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Semi-Supervised Treatment Effect Estimation with Unlabeled Covariates for Prediction-Powered Causal Inference0.92854
2Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate Choice0.79464
3Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games0.73732
4A Practical Guide of Off-Policy Evaluation for Bandit Problems0.73732
5Distributionally Robust Policy Learning with Wasserstein Distance0.73732
6Prediction-Powered Causal Inference by Automatic Debiased Machine Learning and Semi-Supervised Riesz Regression0.64441
7Off-Policy Evaluation of Bandit Algorithm from Dependent Samples under Batch Update Policy0.40511
8Direct Debiased Machine Learning via Bregman Divergence Minimization0.40511
9Nearest Neighbor Matching as Least Squares Density Ratio Estimation and Riesz Regression0.40511
10PUATE: Efficient ATE Estimation from Treated (Positive) and Unlabeled Units0.00054