EconBase
← All papers

Policy Choice and Best Arm Identification: Asymptotic Analysis of Exploration Sampling

Kaito Ariu, Masahiro Kato, Junpei Komiyama, Kenichiro McAlinn, Chao Qin

arXiv 16 Sep 2021 · Econometrics · 3 citations (OpenAlex)

arXiv:2109.08229 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We consider the "policy choice" problem -- otherwise known as best arm identification in the bandit literature -- proposed by Kasy and Sautmann (2021) for adaptive experimental design. Theorem 1 of Kasy and Sautmann (2021) provides three asymptotic results that give theoretical guarantees for exploration sampling developed for this setting. We first show that the proof of Theorem 1 (1) has technical issues, and the proof and statement of Theorem 1 (2) are incorrect. We then show, through a counterexample, that Theorem 1 (3) is false. For the former two, we correct the statements and provide rigorous proofs. For Theorem 1 (3), we propose an alternative objective function, which we call posterior weighted policy regret, and derive the asymptotic optimality of exploration sampling.

Citation extraction

14
references
95
in-text mentions
15
distinct cited
0
self-citations
9,955
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kasy and Sautmann (2021) Adaptive Treatment Assignment in Experiments for Policy Choice1.000358100%
2Shang, de Heide, Menard, Kaufmann, and Valko (2020) Fixed-confidence guarantees for Bayesian best-arm identification1.000205100%
3Russo (2016) Simple Bayesian Algorithms for Best Arm Identification1.000204100%
4Carpentier and Locatelli (2016) Tight (Lower) Bounds for the Fixed Budget Best Arm Identification Bandit Problem0.87452100%
5Glynn and Juneja (2004) A large deviations perspective on ordinal optimization0.73732100%
6Audibert, Bubeck, and Munos (2010) Best Arm Identification in Multi-Armed Bandits0.51121100%
7Lattimore and Szepesv^^c3^^a1ri (2020) Bandit Algorithms0.51121100%
8Chernoff (1959) Sequential Design of Experiments0.40511100%
9Even-Dar, Mannor, and Mansour (2006) Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems0.40511100%
Kaufman2016complexityunmatched citation key Kaufman2016complexity0.40511100%

Showing the top 10 of 15 scored citations. 1 of these could not be matched to a bibliography entry, so only the citation key is shown.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Best Arm Identification with Contextual Information under a Small Gap0.64422
2Optimizing Adaptive Experiments: A Unified Approach to Regret Minimization and Best-Arm Identification0.40511
3Bandit Algorithms for Policy Learning: Methods, Implementation, and Welfare-performance0.40511