Haodong Liang, Krishnakumar Balasubramanian, Lifeng Lai
arXiv 2 Oct 2024 · Statistics — Machine Learning
arXiv:2410.01265 · PDF · DOI · OpenAlex · Extracted main text
We explore the capability of transformers to address endogeneity in in-context linear regression. Our main finding is that transformers inherently possess a mechanism to handle endogeneity effectively using instrumental variables (IV). First, we demonstrate that the transformer architecture can emulate a gradient-based bi-level optimization procedure that converges to the widely used two-stage least squares $(\textsf{2SLS})$ solution at an exponential rate. Next, we propose an in-context pretraining scheme and provide theoretical guarantees showing that the global minimizer of the pre-training loss achieves a small excess loss. Our extensive experiments validate these theoretical findings, showing that the trained transformer provides more robust and reliable in-context predictions and coefficient estimates than the $\textsf{2SLS}$ method, in the presence of endogeneity.
appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei (2023) Transformers as statisticians: Provable in-context learning with in-context algorithm selection | 0.950 | 7 | 3 | 86% |
| 2 | J.M. Wooldridge (2015) Introductory Econometrics: A Modern Approach | 0.737 | 3 | 2 | 100% |
| 3 | Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason… (2023) Looped transformers as programmable computers | 0.644 | 2 | 2 | 100% |
| 4 | Liu Yang, Kangwook Lee, Robert Nowak, and Dimitris Papailiopoulos (2024) Looped transformers are better at learning learning algorithms | 0.644 | 2 | 2 | 100% |
| 5 | Ruiqi Zhang, Spencer Frei, and Peter L Bartlett (2024) Trained transformers learn linear models in-context | 0.644 | 2 | 2 | 100% |
| 6 | Joel A. Tropp An Introduction to Matrix Concentration Inequalities | 0.511 | 2 | 2 | 50% |
| 7 | Joshua D. Angrist and Alan B. Krueger (2001) Instrumental variables and the search for identification: From supply and demand to natural experiments | 0.511 | 2 | 1 | 100% |
| 8 | Joshua D Angrist and Jörn-Steffen Pischke (2009) Mostly harmless econometrics: An empiricist's companion | 0.511 | 2 | 1 | 100% |
| 9 | Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant (2022) What can transformers learn in-context? a case study of simple function classes | 0.511 | 2 | 1 | 100% |
| 10 | Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova (2019) Bert: Pre-training of deep bidirectional transformers for language understanding | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 49 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression | 0.405 | 1 | 1 |