Yeshwanth Cherapanamjeri, Constantinos Daskalakis, Gabriele Farina, Sobhan Mohammadpour
arXiv 17 Oct 2025 · Machine Learning
arXiv:2510.15839 · PDF · DOI · OpenAlex · Extracted main text
Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learning from Human Feedback (RLHF). However, a crucial shortcoming of many of these techniques is the Independence of Irrelevant Alternatives (IIA) assumption, which collapses all human preferences to a universal underlying utility function, yielding a coarse approximation of the range of human preferences. On the other hand, statistical and computational guarantees for models avoiding this assumption are scarce. In this paper, we investigate the statistical and computational challenges of learning a correlated probit model, a fundamental RUM that avoids the IIA assumption. First, we establish that the classical data collection paradigm of pairwise preference data is fundamentally insufficient to learn correlational information, explaining the lack of statistical and computational guarantees in this setting. Next, we demonstrate that best-of-three preference data provably overcomes these shortcomings, and devise a statistically and computationally efficient estimator with near-optimal performance. These results highlight the benefits of higher-order preference data in learning correlated utilities, allowing for more fine-grained modeling of human preferences. Finally, we validate these theoretical guarantees on several real-world datasets, demonstrating improved personalization of human preferences.
appendix boundary found by appendix_command · 28% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kenneth E Train (2009) Discrete choice methods with simulation | 0.737 | 3 | 3 | 67% |
| 2 | R Duncan Luce (1959) Individual choice behavior. Vol. 4 | 0.644 | 2 | 2 | 100% |
| 3 | Moshe E. Ben-Akiva (1973) Structure of passenger travel demand models | 0.511 | 2 | 2 | 50% |
| 4 | Austin R Benson, Ravi Kumar, and Andrew Tomkins (2016) On the relevance of irrelevant alternatives. In Proceedings of the 25th international conference on world wide web. 963–973 | 0.511 | 2 | 2 | 50% |
| 5 | Daniel McFadden (1980) Econometric models for probabilistic choice among products | 0.511 | 2 | 2 | 50% |
| 6 | Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg,… (2017) Deep reinforcement learning from human preferences | 0.405 | 1 | 1 | 100% |
| 7 | Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson… (2022) Constitutional AI: Harmlessness from AI Feedback | 0.405 | 1 | 1 | 100% |
| 8 | Gerard Debreu (1960) Reviewed Work: Individual choice behavior: A theoretical analysis | 0.405 | 1 | 1 | 100% |
| 9 | Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Tu… (2024) Foundational Challenges in Assuring Alignment and Safety of Large Language Models | 0.405 | 1 | 1 | 100% |
| 10 | Guillermo Gallego and Ruxian Wang (2014) Multiproduct price optimization and competition under the nested logit model with product-differentiated price sensitivities | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 47 scored citations.