arXiv 16 Sep 2026 · Econometrics
arXiv:2609.18172 · PDF · Extracted main text
Applied instrumental variables (IV) practice reports a first-stage F, now often the conditional F of Sanderson and Windmeijer (2016), and reads a large value as license to interpret the second stage. We show that no first-stage diagnostic can provide it. With a scalar instrument, a scalar treatment, and covariates entered linearly, the 2SLS estimand splits into a signal that a saturated specification would target and a contamination, the covariance between curvature in the instrument propensity and a covariate level function. The same nuisance sits in both terms, so it biases the estimand and inflates the reported strength at once. When the instrument is nearly collinear with the covariates the signal vanishes and the strength is manufactured. When the strength is honest the curvature still biases the estimand through the outcome, where no first-stage number reaches it. We prove that no functional of the joint distribution of instrument, treatment, and covariates can detect this second bias, and we give a directed test built from reduced-form regressions, a corrected estimator, and a reportable contamination share. Two applications show the modes. A husband's insurance instrument (Olson, 1998) has a conditional F above 36,000, yet correcting a linear income control more than doubles the estimate. The instrument of Nunn and Wantchekon (2011) loses most of its first stage once geography enters flexibly.
appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Blandhol, C., Bonney, J., Mogstad, M., Torgovitsky, A (2022) When is TSLS actually LATE? NBER Working Paper 29709 | 1.000 | 8 | 4 | 100% |
| 2 | Montiel Olea, J.L., Pflueger, C (2013) A robust test for weak instruments | 0.843 | 3 | 3 | 100% |
| 3 | Sanderson, E., Windmeijer, F (2016) A weak instrument $F$-test in linear IV models with multiple endogenous variables | 0.843 | 3 | 3 | 100% |
| 4 | Goldsmith-Pinkham, P., Hull, P., Kolesár, M (2024) Contamination bias in linear regressions | 0.811 | 4 | 2 | 100% |
| 5 | Nunn, N., Wantchekon, L (2011) The slave trade and the origins of mistrust in Africa | 0.811 | 4 | 2 | 100% |
| 6 | Olson, C.A (1998) A comparison of parametric and semiparametric estimates of the effect of spousal health insurance coverage on weekly hours worke… | 0.811 | 4 | 2 | 100% |
| 7 | Soczyński, T., Sun, L., Uysal, S.D (2026) A practical guide to instrumental variables methods with heterogeneous treatment effects | 0.737 | 3 | 2 | 100% |
| 8 | Andrews, I., Stock, J.H., Sun, L (2019) Weak instruments in instrumental variables regression: theory and practice | 0.644 | 2 | 2 | 100% |
| 9 | Newey, W.K., McFadden, D (1994) Large sample estimation and hypothesis testing, in: Handbook of Econometrics, vol. 4 | 0.511 | 2 | 1 | 100% |
| 10 | Soczyński, T (2022) When should we (not) interpret linear IV estimands as LATE? Working paper, Brandeis University | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 29 scored citations.