Anders Bredahl Kock, David Preinerstorfer, Bezirgen Veliyev
arXiv 21 Dec 2018 · Econometrics · publishedJournal of the American Statistical Association (2020)
arXiv:1812.09408 · PDF · DOI · OpenAlex · Extracted main text
Consider a setting in which a policy maker assigns subjects to treatments, observing each outcome before the next subject arrives. Initially, it is unknown which treatment is best, but the sequential nature of the problem permits learning about the effectiveness of the treatments. While the multi-armed-bandit literature has shed much light on the situation when the policy maker compares the effectiveness of the treatments through their mean, much less is known about other targets. This is restrictive, because a cautious decision maker may prefer to target a robust location measure such as a quantile or a trimmed mean. Furthermore, socio-economic decision making often requires targeting purpose specific characteristics of the outcome distribution, such as its inherent degree of inequality, welfare or poverty. In the present paper we introduce and study sequential learning algorithms when the distributional characteristic of interest is a general functional of the outcome distribution. Minimax expected regret optimality results are obtained within the subclass of explore-then-commit policies, and for the unrestricted class of all policies.
appendix boundary found by appendix_command · 32% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Manski, C. F. and A. Tetenov (2016) Sufficient trial size to inform clinical practice | 1.000 | 5 | 4 | 100% |
| 2 | Vakili, S., A. Boukouvalas, and Q. Zhao (2018) Decision Variance in Online Learning | 1.000 | 5 | 3 | 100% |
| 3 | Kock, A. B. and M. Thyrsgaard (2017) Optimal sequential treatment allocation self | 0.928 | 4 | 3 | 100% |
| 4 | Maillard, O.-A (2013) Robust Risk-Averse Stochastic Multi-armed Bandits, in | 0.928 | 4 | 3 | 100% |
| 5 | Manski, C. F (2004) Statistical treatment rules for heterogeneous populations | 0.928 | 4 | 3 | 100% |
| 6 | Zimin, A., R. Ibsen-Jensen, and K. Chatterjee (2014) Generalized risk-aversion in stochastic multi-armed bandits | 0.928 | 4 | 3 | 100% |
| 7 | Degenne, R. and V. Perchet (2016) Anytime optimal algorithms in stochastic multi-armed bandits, in | 0.874 | 6 | 2 | 100% |
| 8 | Kock, A. B., D. Preinerstorfer, and B. Veliyev (2020) a): Functional Sequential Treatment Allocation with Covariates self | 0.843 | 3 | 3 | 100% |
| 9 | Vakili, S. and Q. Zhao (2016) Risk-Averse Multi-Armed Bandit Problems Under Mean-Variance Measure | 0.811 | 4 | 2 | 100% |
| 10 | Auer, P., N. Cesa-Bianchi, and P. Fischer (2002) Finite-time analysis of the multiarmed bandit problem | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 110 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.