EconBase
← All papers

Functional Sequential Treatment Allocation

Anders Bredahl Kock, David Preinerstorfer, Bezirgen Veliyev

arXiv 21 Dec 2018 · Econometrics · publishedJournal of the American Statistical Association (2020)

arXiv:1812.09408 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Consider a setting in which a policy maker assigns subjects to treatments, observing each outcome before the next subject arrives. Initially, it is unknown which treatment is best, but the sequential nature of the problem permits learning about the effectiveness of the treatments. While the multi-armed-bandit literature has shed much light on the situation when the policy maker compares the effectiveness of the treatments through their mean, much less is known about other targets. This is restrictive, because a cautious decision maker may prefer to target a robust location measure such as a quantile or a trimmed mean. Furthermore, socio-economic decision making often requires targeting purpose specific characteristics of the outcome distribution, such as its inherent degree of inequality, welfare or poverty. In the present paper we introduce and study sequential learning algorithms when the distributional characteristic of interest is a general functional of the outcome distribution. Minimax expected regret optimality results are obtained within the subclass of explore-then-commit policies, and for the unrestricted class of all policies.

Citation extraction

110
references
210
in-text mentions
110
distinct cited
3
self-citations
12,747
main-text words

appendix boundary found by appendix_command · 32% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Manski, C. F. and A. Tetenov (2016) Sufficient trial size to inform clinical practice1.00054100%
2Vakili, S., A. Boukouvalas, and Q. Zhao (2018) Decision Variance in Online Learning1.00053100%
3Kock, A. B. and M. Thyrsgaard (2017) Optimal sequential treatment allocation self0.92843100%
4Maillard, O.-A (2013) Robust Risk-Averse Stochastic Multi-armed Bandits, in0.92843100%
5Manski, C. F (2004) Statistical treatment rules for heterogeneous populations0.92843100%
6Zimin, A., R. Ibsen-Jensen, and K. Chatterjee (2014) Generalized risk-aversion in stochastic multi-armed bandits0.92843100%
7Degenne, R. and V. Perchet (2016) Anytime optimal algorithms in stochastic multi-armed bandits, in0.87462100%
8Kock, A. B., D. Preinerstorfer, and B. Veliyev (2020) a): Functional Sequential Treatment Allocation with Covariates self0.84333100%
9Vakili, S. and Q. Zhao (2016) Risk-Averse Multi-Armed Bandit Problems Under Mean-Variance Measure0.81142100%
10Auer, P., N. Cesa-Bianchi, and P. Fischer (2002) Finite-time analysis of the multiarmed bandit problem0.81142100%

Showing the top 10 of 110 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Treatment recommendation with distributional targets1.00073
2Regularizing Fairness in Optimal Policy Learning with Distributional Targets0.84343
3Functional Sequential Treatment Allocation with Covariates0.794205
4Who Should Get Vaccinated? Individualized Allocation of Vaccines Over SIR Network0.40511
5Policy Choice in Time Series by Empirical Welfare Maximization0.40511
6Treatment Choice with Nonlinear Regret0.40511
7Identification and Inference for Welfare Gains without Unconfoundedness0.40511
8Asymptotic Representations for Sequential Decisions, Adaptive Experiments, and Batched Bandits0.40511
9Policy Learning with Distributional Welfare0.40511
10Bandit Algorithms for Policy Learning: Methods, Implementation, and Welfare-performance0.40511