EconBase
← All papers

Transformers Handle Endogeneity in In-Context Linear Regression

Haodong Liang, Krishnakumar Balasubramanian, Lifeng Lai

arXiv 2 Oct 2024 · Statistics — Machine Learning

arXiv:2410.01265 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We explore the capability of transformers to address endogeneity in in-context linear regression. Our main finding is that transformers inherently possess a mechanism to handle endogeneity effectively using instrumental variables (IV). First, we demonstrate that the transformer architecture can emulate a gradient-based bi-level optimization procedure that converges to the widely used two-stage least squares $(\textsf{2SLS})$ solution at an exponential rate. Next, we propose an in-context pretraining scheme and provide theoretical guarantees showing that the global minimizer of the pre-training loss achieves a small excess loss. Our extensive experiments validate these theoretical findings, showing that the trained transformer provides more robust and reliable in-context predictions and coefficient estimates than the $\textsf{2SLS}$ method, in the presence of endogeneity.

Citation extraction

49
references
64
in-text mentions
49
distinct cited
3
self-citations
7,216
main-text words

appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei (2023) Transformers as statisticians: Provable in-context learning with in-context algorithm selection0.9507386%
2J.M. Wooldridge (2015) Introductory Econometrics: A Modern Approach0.73732100%
3Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason… (2023) Looped transformers as programmable computers0.64422100%
4Liu Yang, Kangwook Lee, Robert Nowak, and Dimitris Papailiopoulos (2024) Looped transformers are better at learning learning algorithms0.64422100%
5Ruiqi Zhang, Spencer Frei, and Peter L Bartlett (2024) Trained transformers learn linear models in-context0.64422100%
6Joel A. Tropp An Introduction to Matrix Concentration Inequalities0.5112250%
7Joshua D. Angrist and Alan B. Krueger (2001) Instrumental variables and the search for identification: From supply and demand to natural experiments0.51121100%
8Joshua D Angrist and Jörn-Steffen Pischke (2009) Mostly harmless econometrics: An empiricist's companion0.51121100%
9Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant (2022) What can transformers learn in-context? a case study of simple function classes0.51121100%
10Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova (2019) Bert: Pre-training of deep bidirectional transformers for language understanding0.40511100%

Showing the top 10 of 49 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression0.40511