Omri Feldman, Amar Venugopal, Jann Spiess, Amir Feder
arXiv 17 Feb 2026 · cs.CL
arXiv:2602.15730 · PDF · DOI · OpenAlex · Extracted main text
Understanding the causal effects of text on downstream outcomes is a central task in many applications. Estimating such effects requires researchers to run controlled experiments that systematically vary textual features. While large language models (LLMs) hold promise for generating text, producing and evaluating controlled variation requires more careful attention. In this paper, we present an end-to-end pipeline for the generation and causal estimation of latent textual interventions. Our work first performs hypothesis generation and steering via sparse autoencoders (SAEs), followed by robust causal estimation. Our pipeline addresses both computational and statistical challenges in text-as-treatment experiments. We demonstrate that naive estimation of causal effects suffers from significant bias as text inherently conflates treatment and covariate information. We describe the estimation bias induced in this setting and propose a solution based on covariate residualization. Our empirical results show that our pipeline effectively induces variation in target features and mitigates estimation error, providing a robust foundation for causal effect estimation in text-as-treatment settings.
appendix boundary found by appendix_command · 64% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ash, Elliott and Hansen, Stephen (2023) Text algorithms in economics | 0.644 | 2 | 2 | 100% |
| 2 | Bricken, Trenton and Templeton, Adly and Batson, Joshua and Chen, Br… (2023) Towards Monosemanticity: Decomposing Language Models With Dictionary Learning | 0.644 | 2 | 2 | 100% |
| 3 | Feder, Amir and Keith, Katherine A and Manzoor, Emaad and Pryzant, R… (2022) Causal inference in natural language processing: Estimation, prediction, interpretation and beyond self | 0.644 | 2 | 2 | 100% |
| 4 | Fong, Christian and Grimmer, Justin (2016) Discovery of treatments from text corpora | 0.644 | 2 | 2 | 100% |
| 5 | Martin, Olivia and Amar Venugopal (2026) Participation and Representation in Local Government Speech self | 0.644 | 2 | 2 | 100% |
| 6 | Lin, Johnny (2023) Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks | 0.511 | 2 | 2 | 50% |
| 7 | Zheng, Carolina and Beltran-Velez, Nicolas and Karlekar, Sweta and S… (2025) Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders self | 0.511 | 2 | 2 | 50% |
| 8 | Nie, Xinkun and Wager, Stefan (2021) Quasi-oracle estimation of heterogeneous treatment effects | 0.511 | 2 | 1 | 100% |
| 9 | Bills, Steven and Cammarata, Nick and Mossing, Dan and Tillman, Henk… (2023) Language models can explain neurons in language models | 0.405 | 1 | 1 | 100% |
| 10 | Angelopoulos, Panagiotis and Lee, Kevin and Misra, Sanjog (2024) Value aligned large language models | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 44 scored citations.