EconBase
← All papers

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

Erfan Loghmani

arXiv 30 May 2025 · Machine Learning

arXiv:2506.00152 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Large language models are being widely used across industries to generate content that contributes directly to key performance metrics, such as conversion rates. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering and logistical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data that remains underutilized. In this work, we study the challenges and opportunities of fine-tuning LLMs using observational data. We show that while observational outcomes can provide valuable supervision, directly fine-tuning models on such data can lead them to learn spurious correlations. We present empirical evidence of this issue using various real-world datasets and propose DeconfoundLM, a method that explicitly removes the effect of known confounders from reward signals. Using simulation experiments, we demonstrate that DeconfoundLM improves the recovery of causal relationships and mitigates failure modes found in fine-tuning methods that ignore or naively incorporate confounding variables. Our findings highlight that while observational data presents risks, with the right causal corrections, it can be a powerful source of signal for LLM alignment. Please refer to the project page for code and related resources.

Citation extraction

42
references
62
in-text mentions
42
distinct cited
0
self-citations
6,245
main-text words

appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zikun Ye, Hema Yoganarasimhan, and Yufeng Zheng (2024) Lola: Llm-assisted online learning algorithm for content experiments0.7374350%
2Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning… (2024) Direct preference optimization: Your language model is secretly a reward model0.7373367%
3Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom… (2021) A general language assistant as a laboratory for alignment0.6443267%
4Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters0.6443267%
5Leo Gao, John Schulman, and Jacob Hilton (2023) Scaling laws for reward model overoptimization0.64422100%
6Rafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sushil Sikchi… (2024) Scaling laws for reward model overoptimization in direct alignment algorithms0.64422100%
7Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright… (2022) Training language models to follow instructions with human feedback0.5853333%
8Panagiotis Angelopoulos, Kevin Lee, and Sanjog Misra (2024) Causal alignment: Augmenting language models with a/b tests0.5112250%
9Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie… (2023) Pythia: A suite for analyzing large language models across training and scaling0.5112250%
10Nishanth Dikkala, Greg Lewis, Lester Mackey, and Vasilis Syrgkanis (2020) Minimax estimation of conditional moment models0.5112250%

Showing the top 10 of 42 scored citations.