EconBase
← All papers

Mostly Harmless Machine Learning: Learning Optimal Instruments in Linear IV Models

Jiafeng Chen, Daniel L. Chen, Greg Lewis

arXiv 12 Nov 2020 · Econometrics · 7 citations (OpenAlex)

arXiv:2011.06158 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We offer straightforward theoretical results that justify incorporating machine learning in the standard linear instrumental variable setting. The key idea is to use machine learning, combined with sample-splitting, to predict the treatment variable from the instrument and any exogenous covariates, and then use this predicted treatment and the covariates as technical instruments to recover the coefficients in the second-stage. This allows the researcher to extract non-linear co-variation between the treatment and instrument that may dramatically improve estimation precision and robustness by boosting instrument strength. Importantly, we constrain the machine-learned predictions to be linear in the exogenous covariates, thus avoiding spurious identification arising from non-linear relationships between the treatment and the covariates. We show that this approach delivers consistent and asymptotically normal estimates under weak conditions and that it may be adapted to be semiparametrically efficient (Chamberlain, 1992). Our method preserves standard intuitions and interpretations of linear instrumental variable methods, including under weak identification, and provides a simple, user-friendly upgrade to the applied economics toolbox. We illustrate our method with an example in law and criminal justice, examining the causal effect of appellate court reversals on district court sentencing decisions.

Citation extraction

45
references
79
in-text mentions
45
distinct cited
4
self-citations
8,968
main-text words

appendix boundary found by appendix_command · 70% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.87462100%
2–- (1992) Comment: Sequential moment restrictions in panel data0.84333100%
3Chamberlain, G (1987) Asymptotic efficiency in estimation with conditional moment restrictions0.81142100%
4Andrews, I., Stock, J. H. and Sun, L (2019) Weak instruments in iv regression: Theory and practice0.73732100%
5–- and Pischke, J.-S (2008) Mostly harmless econometrics: An empiricist's companion0.73732100%
6Angrist, J. and Frandsen, B (2019) Machine labor0.73732100%
7Belloni, A., Chen, D., Chernozhukov, V. and Hansen, C (2012) Sparse models and methods for optimal instruments with an application to eminent domain self0.73732100%
8Ai, C. and Chen, X (2003) Efficient estimation of models with conditional moment restrictions containing unknown functions0.64422100%
9–- and –- (2007) Estimation of possibly misspecified semiparametric conditional moment restriction models with different conditioning variables0.64422100%
10–- and –- (2012) The semiparametric efficiency bound for models of sequential moment restrictions containing unknown functions0.64422100%

Showing the top 10 of 45 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Machine learning the first stage in 2SLS: Practical guidance from bias decomposition and simulation0.971125
22.5cm Identification and Inference with Machine-Learned Instruments0.92844
3Improving Inference from Simple Instruments through Compliance Estimation0.81142
4Semiparametric Bayesian Inference for a Conditional Moment Equality Model0.51121
5Program Evaluation with Remotely Sensed Outcomes0.51122
6A Distance Covariance-based Estimator0.40511