Kaito Ariu, Masahiro Kato, Junpei Komiyama, Kenichiro McAlinn, Chao Qin
arXiv 16 Sep 2021 · Econometrics · 3 citations (OpenAlex)
arXiv:2109.08229 · PDF · DOI · OpenAlex · Extracted main text
We consider the "policy choice" problem -- otherwise known as best arm identification in the bandit literature -- proposed by Kasy and Sautmann (2021) for adaptive experimental design. Theorem 1 of Kasy and Sautmann (2021) provides three asymptotic results that give theoretical guarantees for exploration sampling developed for this setting. We first show that the proof of Theorem 1 (1) has technical issues, and the proof and statement of Theorem 1 (2) are incorrect. We then show, through a counterexample, that Theorem 1 (3) is false. For the former two, we correct the statements and provide rigorous proofs. For Theorem 1 (3), we propose an alternative objective function, which we call posterior weighted policy regret, and derive the asymptotic optimality of exploration sampling.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kasy and Sautmann (2021) Adaptive Treatment Assignment in Experiments for Policy Choice | 1.000 | 35 | 8 | 100% |
| 2 | Shang, de Heide, Menard, Kaufmann, and Valko (2020) Fixed-confidence guarantees for Bayesian best-arm identification | 1.000 | 20 | 5 | 100% |
| 3 | Russo (2016) Simple Bayesian Algorithms for Best Arm Identification | 1.000 | 20 | 4 | 100% |
| 4 | Carpentier and Locatelli (2016) Tight (Lower) Bounds for the Fixed Budget Best Arm Identification Bandit Problem | 0.874 | 5 | 2 | 100% |
| 5 | Glynn and Juneja (2004) A large deviations perspective on ordinal optimization | 0.737 | 3 | 2 | 100% |
| 6 | Audibert, Bubeck, and Munos (2010) Best Arm Identification in Multi-Armed Bandits | 0.511 | 2 | 1 | 100% |
| 7 | Lattimore and Szepesv^^c3^^a1ri (2020) Bandit Algorithms | 0.511 | 2 | 1 | 100% |
| 8 | Chernoff (1959) Sequential Design of Experiments | 0.405 | 1 | 1 | 100% |
| 9 | Even-Dar, Mannor, and Mansour (2006) Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems | 0.405 | 1 | 1 | 100% |
| Kaufman2016complexity | unmatched citation key Kaufman2016complexity | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 15 scored citations. 1 of these could not be matched to a bibliography entry, so only the citation key is shown.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.