Ruohan Zhan, Vitor Hadad, David A. Hirshberg, Susan Athey
arXiv 3 Jun 2021 · Statistics — Machine Learning · 22 citations (OpenAlex)
arXiv:2106.02029 · PDF · DOI · OpenAlex · Extracted main text
It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment assignment policies to guide future innovation or experiments. However, policy evaluation is challenging if the target policy differs from the one used to collect data, and popular estimators, including doubly robust (DR) estimators, can be plagued by bias, excessive variance, or both. In particular, when the pattern of treatment assignment in the collected data looks little like the pattern generated by the policy to be evaluated, the importance weights used in DR estimators explode, leading to excessive variance. In this paper, we improve the DR estimator by adaptively weighting observations to control its variance. We show that a t-statistic based on our improved estimator is asymptotically normal under certain conditions, allowing us to form confidence intervals and test hypotheses. Using synthetic data and public benchmarks, we provide empirical evidence for our estimator's improved accuracy and inferential properties relative to existing alternatives.
appendix boundary found by appendix_command · 24% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Dudḱ, M., Langford, J., and Li, L (2011) Doubly robust policy evaluation and learning | 0.928 | 4 | 3 | 100% |
| 2 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments self | 0.794 | 10 | 6 | 50% |
| 3 | Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L (2013) Openml: Networked science in machine learning | 0.644 | 3 | 2 | 67% |
| 4 | Agarwal, A., Basu, S., Schnabel, T., and Joachims, T (2017) Effective evaluation using logged bandit feedback from multiple loggers | 0.644 | 2 | 2 | 100% |
| 5 | Agrawal, S. and Goyal, N (2013) Thompson sampling for contextual bandits with linear payoffs | 0.644 | 2 | 2 | 100% |
| 6 | Imbens, G. W (2004) Nonparametric estimation of average treatment effects under exogeneity: A review | 0.644 | 2 | 2 | 100% |
| 7 | Imbens, G. W. and Rubin, D. B (2015) Causal inference in statistics, social, and biomedical sciences | 0.644 | 2 | 2 | 100% |
| 8 | Su, Y., Dimakopoulou, M., Krishnamurthy, A., and Dudḱ, M (2020) Doubly robust off-policy evaluation with shrinkage | 0.644 | 2 | 2 | 100% |
| 9 | Wang, Y.-X., Agarwal, A., and Dudk, M (2017) Optimal and adaptive off-policy evaluation in contextual bandits | 0.644 | 2 | 2 | 100% |
| 10 | Zhang, K. W., Janson, L., and Murphy, S. A (2020) Inference for batched bandits | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 32 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.