arXiv 26 Dec 2025 · Machine Learning
arXiv:2512.21917 · PDF · DOI · OpenAlex · Extracted main text
Aligning large language models (LLMs) to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., a logistic Bradley-Terry link). Misspecification of this link can bias inferred rewards and misalign learned policies. We study preference alignment under an unknown and unrestricted link function. We show that realizability of $f$-divergence-constrained reward maximization in a policy class induces a semiparametric single-index binary choice model, where a scalar policy-dependent index captures all dependence on demonstrations and the remaining preference distribution is unrestricted. Rather than assuming this model has identifiable finite-dimensional structural parameters and estimating them, as in econometrics, we focus on policy learning with the reward function implicit, analyzing error to the optimal policy and allowing for unidentifiable nonparametric indices. We develop preference optimization algorithms robust to the unknown link and prove convergence guarantees in terms of generic function complexity measures. We demonstrate this empirically on LLM alignment. Code is available at https://github.com/causalml/spo/
appendix boundary found by appendix_command · 40% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Roger W. Klein and Richard H. Spady (1993) An efficient semiparametric estimator for binary response models | 1.000 | 7 | 5 | 100% |
| 2 | Joel L. Horowitz (1992) A smoothed maximum score estimator for the binary response model | 1.000 | 7 | 4 | 100% |
| 3 | Stephen R. Cosslett (1983) Distribution-free maximum likelihood estimator of the binary choice model | 1.000 | 5 | 5 | 100% |
| 4 | Jae-Young Kim and David Pollard (1990) Cube root asymptotics | 1.000 | 5 | 3 | 100% |
| 5 | Robert Sherman (1993) The limiting distribution of the maximum rank correlation estimator | 0.928 | 4 | 4 | 100% |
| 6 | Aaron K. Han (1987) Non-parametric estimation of a binary choice model by maximum rank correlation | 0.843 | 3 | 3 | 100% |
| 7 | Charles F. Manski (1975) Maximum score estimation of the stochastic utility model of choice | 0.843 | 3 | 3 | 100% |
| 8 | Charles F. Manski (1985) Semiparametric analysis of discrete response: Asymptotic properties of the maximum score estimator | 0.843 | 3 | 3 | 100% |
| 9 | Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Chelse… (2023) Direct preference optimization: Your language model is secretly a reward model | 0.811 | 4 | 2 | 100% |
| 10 | Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Ra… (2019) Fine-tuning language models from human preferences | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 69 scored citations.