arXiv 2 May 2026 · Machine Learning
arXiv:2605.01311 · PDF · DOI · OpenAlex · Extracted main text
Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence how its output is judged, so raw comparisons of logged scores mix self-selected populations rather than estimating a common quantity of interest. A small randomized experiment can break this bias by overriding model choice, but in practice such experiments are scarce and costly. We study a three-source design that combines a large confounded observational log (OBS) for scale, a small randomized experiment (EXP) for unconfounded scoring, and an offline simulator (SIM) that replays candidate models on cached contexts. Our main result is an identification theorem showing that the randomized experiment and the simulator are together enough to recover causal model values; the observational log enters only afterward, to reduce estimation error rather than to make the causal comparison valid. Six estimator families are evaluated in a controlled semi-synthetic validation and in two real-task cached benchmarks for summarization and coding. No family dominates every regime; relative performance depends on the amount of unbiased EXP supervision and on how closely the target reward aligns with OBS-derived structure.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nathan Kallus, Aahlad Manas Puli, and Uri Shalit (2018) Removing hidden confounding by experimental grounding | 1.000 | 6 | 3 | 100% |
| 2 | Evan T. R. Rosenman, Guillaume Basse, Art B. Owen, and Mike Baiocchi (2023) Combining observational and experimental datasets using shrinkage estimators | 1.000 | 6 | 3 | 100% |
| 3 | Xuelin Yang, Licong Lin, Susan Athey, Michael I. Jordan, and Guido W… (2025) Cross-validated causal inference: a modern method to combine experimental and observational data, 2025 | 1.000 | 6 | 3 | 100% |
| 4 | David Cheng and Tianxi Cai (2021) Adaptive combination of randomized and observational data, 2021 | 1.000 | 5 | 3 | 100% |
| 5 | Xi Lin, Jens Magelund Tarp, and Robin J. Evans (2025) Combining experimental and observational data through a power likelihood | 0.928 | 4 | 3 | 100% |
| 6 | Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelo… (2024) Chatbot arena: An open platform for evaluating LLMs by human preference | 0.843 | 3 | 3 | 100% |
| 7 | Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeh… Teaching machines to read and comprehend | 0.737 | 3 | 3 | 67% |
| 8 | Abigail See, Peter J. Liu, and Christopher D. Manning Get to the point: Summarization with pointer-generator networks | 0.737 | 3 | 3 | 67% |
| 9 | Heejung Bang and James M. Robins Doubly robust estimation in missing data and causal inference models | 0.644 | 2 | 2 | 100% |
| 10 | Elias Bareinboim and Judea Pearl (2016) Causal inference and the data-fusion problem | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 39 scored citations.