Masahiro Kato, Kyohei Okumura, Takuya Ishihara, Toru Kitagawa
arXiv 8 Jan 2024 · Machine Learning · 1 citations (OpenAlex)
arXiv:2401.03756 · PDF · DOI · OpenAlex · Extracted main text
This study investigates the contextual best arm identification (BAI) problem, aiming to design an adaptive experiment to identify the best treatment arm conditioned on contextual information (covariates). We consider a decision-maker who assigns treatment arms to experimental units during an experiment and recommends the estimated best treatment arm based on the contexts at the end of the experiment. The decision-maker uses a policy for recommendations, which is a function that provides the estimated best treatment arm given the contexts. In our evaluation, we focus on the worst-case expected regret, a relative measure between the expected outcomes of an optimal policy and our proposed policy. We derive a lower bound for the expected simple regret and then propose a strategy called Adaptive Sampling-Policy Learning (PLAS). We prove that this strategy is minimax rate-optimal in the sense that its leading factor in the regret upper bound matches the lower bound as the number of experimental units increases.
appendix boundary found by appendix_command · 35% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kato, M., Ishihara, T., Honda, J., and Narita, Y (2002) Adaptive experimental design for efficient treatment effect estimation: Randomized allocation via contextual bandit algorithm, 2… self | 0.928 | 5 | 4 | 80% |
| 2 | van der Laan, M. J (2008) The construction and analysis of adaptive group sequential designs | 0.874 | 6 | 4 | 67% |
| 3 | Athey, S. and Wager, S (2021) Policy learning with observational data | 0.737 | 3 | 3 | 67% |
| 4 | Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S (2021) Confidence intervals for policy evaluation in adaptive experiments | 0.737 | 3 | 3 | 67% |
| 5 | Kato, M., McAlinn, K., and Yasui, S (2021) The adaptive doubly robust estimator and a paradox concerning logging policy self | 0.737 | 3 | 2 | 100% |
| 6 | Hahn, J., Hirano, K., and Karlan, D (2011) Adaptive experimental design using the propensity score | 0.693 | 6 | 4 | 33% |
| 7 | Kaufmann, E., Cappé, O., and Garivier, A (2016) On the complexity of best-arm identification in multi-armed bandit models | 0.669 | 10 | 3 | 30% |
| 8 | Glynn, P. and Juneja, S (2004) A large deviations perspective on ordinal optimization | 0.644 | 4 | 2 | 50% |
| 9 | Bubeck, S., Munos, R., and Stoltz, G (2011) Pure exploration in finitely-armed and continuous-armed bandits | 0.585 | 10 | 2 | 30% |
| 10 | Zhan, R., Ren, Z., Athey, S., and Zhou, Z (2022) Policy learning with adaptively collected data, 2022 | 0.585 | 3 | 3 | 33% |
Showing the top 10 of 101 scored citations.