arXiv 23 Oct 2020 · Machine Learning · 2 citations (OpenAlex)
arXiv:2010.13554 · PDF · DOI · OpenAlex · Extracted main text
The goal of off-policy evaluation (OPE) is to evaluate a new policy using historical data obtained via a behavior policy. However, because the contextual bandit algorithm updates the policy based on past observations, the samples are not independent and identically distributed (i.i.d.). This paper tackles this problem by constructing an estimator from a martingale difference sequence (MDS) for the dependent samples. In the data-generating process, we do not assume the convergence of the policy, but the policy uses the same conditional probability of choosing an action during a certain period. Then, we derive an asymptotically normal estimator of the value of an evaluation policy. As another advantage of our method, the batch-based approach simultaneously solves the deficient support problem. Using benchmark and real-world datasets, we experimentally confirm the effectiveness of the proposed method.
appendix boundary found by appendix_command · 39% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | van der Laan, M. J. and Lendle, S. D (2014) Online targeted learning | 1.000 | 10 | 3 | 100% |
| 2 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2019) Confidence intervals for policy evaluation in adaptive experiments | 1.000 | 9 | 3 | 100% |
| 3 | Kato, M., Ishihara, T., Honda, J., and Narita, Y (2002) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm self | 0.874 | 7 | 2 | 100% |
| 4 | van der Laan, M. J (2008) The construction and analysis of adaptive group sequential designs | 0.874 | 5 | 2 | 100% |
| 5 | Luedtke, A. R. and van der Laan, M. J (2016) Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy | 0.874 | 5 | 2 | 100% |
| 6 | Narita, Y., Yasui, S., and Yata, K (2019) Efficient counterfactual learning from bandit feedback | 0.874 | 5 | 2 | 100% |
| 7 | Hahn, J., Hirano, K., and Karlan, D (2011) Adaptive experimental design using the propensity score | 0.811 | 4 | 2 | 100% |
| 8 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters | 0.737 | 3 | 2 | 100% |
| 9 | Yang, Y. and Zhu, D (2002) Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates | 0.737 | 3 | 2 | 100% |
| 10 | Zheng, W. and van der Laan, M. J (2011) Cross-validated targeted minimum-loss-based estimation | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 36 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Adaptive Doubly Robust Estimator from Non-stationary Logging Policy under a Convergence of Average Probability | 0.405 | 1 | 1 |