Nathan Canen, Ted Enamorado
arXiv 3 Aug 2026 · Econometrics
arXiv:2608.02909 · PDF · Extracted main text
Prediction-based methods, including Large Language Models (LLMs) and other machine learning techniques, are often used to construct measures of political phenomena that are difficult to quantify directly, such as policy positions in manifestos or emotions expressed on social media. In many applications, these prediction-generated measures are used as explanatory variables in regression models, even though they are measured with error. This leads to biased estimates. In this paper, we propose a simple solution to these biases: instrumental variables constructed from multiple measures created on independent splits of the original data. This approach is theoretically valid, easy to implement, and does not require new data. Through simulations, we show that this approach recovers estimates close to the true values, even in relatively small samples, while the standard approach can produce substantial bias in practice. We illustrate the method by revisiting two applications: whether gendered speech affects legislative outcomes in the German Parliament, and whether political risk influences poverty alleviation programs in China.
appendix boundary found by appendix_command · 85% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Battaglia, Laura and Christensen, Timothy and Hansen, Stephen and Sa… (2025) Inference for Regression with Variables Generated by AI or Machine Learning | 0.693 | 8 | 1 | 100% |
| 2 | Yang, Mochen and McFowland III, Edward and Burtch, Gordon and Adomav… (2022) Achieving reliable causal inference with data-mined variables: A random forest approach to the measurement error problem | 0.693 | 8 | 1 | 100% |
| 3 | Burtch, Gordon and McFowland III, Edward and Yang, Mochen and Adomav… (2026) EnsembleIV: Creating Instrumental Variables from Ensemble Learners for Robust Statistical Inference with ML-Generated Variables | 0.693 | 7 | 1 | 100% |
| 4 | Lin, Shengqiao (2025) Addressing Risk by Doing Good: Business Response to Government Policy Initiative | 0.693 | 7 | 1 | 100% |
| 5 | Duan, Junting and Pelger, Markus (2026) Inference with AI-Generated Covariates | 0.693 | 6 | 1 | 100% |
| 6 | Gillen, Ben and Snowberg, Erik and Yariv, Leeat (2019) Experimenting with measurement error: Techniques with applications to the caltech cohort study | 0.693 | 5 | 1 | 100% |
| 7 | Ash, Elliott and Krümmel, Johann and Slapin, Jonathan B (2025) Gender and reactions to speeches in German parliamentary debates | 0.659 | 7 | 1 | 86% |
| 8 | Wooldridge, Jeffrey M (2016) Introductory econometrics a modern approach | 0.644 | 4 | 1 | 100% |
| 9 | Angelopoulos, Anastasios N and Bates, Stephen and Fannjiang, Clara a… (2023) Prediction-powered inference | 0.585 | 3 | 1 | 100% |
| 10 | Braghieri, Luca and Eichmeyer, Sarah and Levy, Ro’ee and Mobius, Mar… (2024) Article-Level Slant and Polarization of News Consumption on Social Media | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 31 scored citations.