EconBase
← All papers

Enhancing Preference-based Linear Bandits via Human Response Time

Shen Li, Yuyang Zhang, Zhaolin Ren, Claire Liang, Na Li, Julie A. Shah

arXiv 9 Sep 2024 · Machine Learning

arXiv:2409.05798 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Interactive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html

Citation extraction

79
references
222
in-text mentions
79
distinct cited
0
self-citations
8,134
main-text words

appendix boundary found by appendix_command · 46% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1M. Azizi, B. Kveton, and M. Ghavamzadeh (2022) Fixed-budget best-arm identification in structured bandits0.97112592%
2K.-S. Jun, L. Jain, B. Mason, and H. Nassif (2021) Improved confidence bounds for the linear logistic model and applications to bandits0.96510590%
3T. Fiez, L. Jain, K. G. Jamieson, and L. Ratliff (2019) Sequential experimental design for transductive linear bandits0.8434375%
4K. Xiang Chiong, M. Shum, R. Webb, and R. Chen (2023) Combining choice and response time data: A drift-diffusion model of mobile advertisements0.8307457%
5S. M. Smith and I. Krajbich (2018) Attention and choice across domains0.7946450%
6X. Yang and I. Krajbich (2023) A dynamic computational model of gaze and choice in multi-attribute decisions0.7547543%
7E.-J. Wagenmakers, H. L. J. Van Der Maas, and R. P. P. P. Grasman (2007) An ez-diffusion model for response time and accuracy0.75019842%
8J. A. Clithero (2018) Response times in economics: Looking through the lens of sequential sampling models0.7374275%
9G. Fisher (2017) An attentional drift diffusion model over binary-attribute choice0.7373367%
10C. Tao, S. Blanco, and Y. Zhou (2018) Best arm identification in linear bandits with linear dimension dependency0.7373367%

Showing the top 10 of 79 scored citations.