EconBase
← All papers

Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits

Ruohan Zhan, Vitor Hadad, David A. Hirshberg, Susan Athey

arXiv 3 Jun 2021 · Statistics — Machine Learning · 22 citations (OpenAlex)

arXiv:2106.02029 · PDF · DOI · OpenAlex · Extracted main text

Abstract

It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment assignment policies to guide future innovation or experiments. However, policy evaluation is challenging if the target policy differs from the one used to collect data, and popular estimators, including doubly robust (DR) estimators, can be plagued by bias, excessive variance, or both. In particular, when the pattern of treatment assignment in the collected data looks little like the pattern generated by the policy to be evaluated, the importance weights used in DR estimators explode, leading to excessive variance. In this paper, we improve the DR estimator by adaptively weighting observations to control its variance. We show that a t-statistic based on our improved estimator is asymptotically normal under certain conditions, allowing us to form confidence intervals and test hypotheses. Using synthetic data and public benchmarks, we provide empirical evidence for our estimator's improved accuracy and inferential properties relative to existing alternatives.

Citation extraction

32
references
53
in-text mentions
32
distinct cited
4
self-citations
6,192
main-text words

appendix boundary found by appendix_command · 24% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning0.92843100%
2Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments self0.79410650%
3Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L (2013) Openml: Networked science in machine learning0.6443267%
4Agarwal, A., Basu, S., Schnabel, T., and Joachims, T (2017) Effective evaluation using logged bandit feedback from multiple loggers0.64422100%
5Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs0.64422100%
6Imbens, G. W (2004) Nonparametric estimation of average treatment effects under exogeneity: A review0.64422100%
7Imbens, G. W. and Rubin, D. B (2015) Causal inference in statistics, social, and biomedical sciences0.64422100%
8Su, Y., Dimakopoulou, M., Krishnamurthy, A., and Dudḱ, M (2020) Doubly robust off-policy evaluation with shrinkage0.64422100%
9Wang, Y.-X., Agarwal, A., and Dudk, M (2017) Optimal and adaptive off-policy evaluation in contextual bandits0.64422100%
10Zhang, K. W., Janson, L., and Murphy, S. A (2020) Inference for batched bandits0.51121100%

Showing the top 10 of 32 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Policy Learning with Adaptively Collected Data0.87472
2Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity0.84333
3Demistifying Inference after Adaptive Experiments0.51121
4Double/Debiased Machine Learning for Dynamic Treatment Effects via $g$-Estimation0.40511
5Dynamic Selection in Algorithmic Decision-making0.40511
6Adversarial Estimators0.40511
7Best Arm Identification with Contextual Information under a Small Gap0.40511
8Contextual Bandits in a Survey Experiment on Charitable Giving: Within-Experiment Outcomes versus Policy Learning0.40511
9Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate Choice0.40511
10Estimating Causal Effects from Data Generated by Stochastic Algorithms0.40511