Shen Li, Yuyang Zhang, Zhaolin Ren, Claire Liang, Na Li, Julie A. Shah
arXiv 9 Sep 2024 · Machine Learning
arXiv:2409.05798 · PDF · DOI · OpenAlex · Extracted main text
Interactive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html
appendix boundary found by appendix_command · 46% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | M. Azizi, B. Kveton, and M. Ghavamzadeh (2022) Fixed-budget best-arm identification in structured bandits | 0.971 | 12 | 5 | 92% |
| 2 | K.-S. Jun, L. Jain, B. Mason, and H. Nassif (2021) Improved confidence bounds for the linear logistic model and applications to bandits | 0.965 | 10 | 5 | 90% |
| 3 | T. Fiez, L. Jain, K. G. Jamieson, and L. Ratliff (2019) Sequential experimental design for transductive linear bandits | 0.843 | 4 | 3 | 75% |
| 4 | K. Xiang Chiong, M. Shum, R. Webb, and R. Chen (2023) Combining choice and response time data: A drift-diffusion model of mobile advertisements | 0.830 | 7 | 4 | 57% |
| 5 | S. M. Smith and I. Krajbich (2018) Attention and choice across domains | 0.794 | 6 | 4 | 50% |
| 6 | X. Yang and I. Krajbich (2023) A dynamic computational model of gaze and choice in multi-attribute decisions | 0.754 | 7 | 5 | 43% |
| 7 | E.-J. Wagenmakers, H. L. J. Van Der Maas, and R. P. P. P. Grasman (2007) An ez-diffusion model for response time and accuracy | 0.750 | 19 | 8 | 42% |
| 8 | J. A. Clithero (2018) Response times in economics: Looking through the lens of sequential sampling models | 0.737 | 4 | 2 | 75% |
| 9 | G. Fisher (2017) An attentional drift diffusion model over binary-attribute choice | 0.737 | 3 | 3 | 67% |
| 10 | C. Tao, S. Blanco, and Y. Zhou (2018) Best arm identification in linear bandits with linear dimension dependency | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 79 scored citations.