Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, James Robins
arXiv 30 Jul 2016 · Statistics — Machine Learning · 103 citations (OpenAlex)
arXiv:1608.00060 · PDF · DOI · OpenAlex · Extracted main text
Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, average treatment effects, average lifts, and demand or supply elasticities. In fact, estimates of such causal parameters obtained via naively plugging ML estimators into estimating equations for such parameters can behave very poorly due to the regularization bias. Fortunately, this regularization bias can be removed by solving auxiliary prediction problems via ML tools. Specifically, we can form an orthogonal score for the target low-dimensional parameter by combining auxiliary and main ML predictions. The score is then used to build a de-biased estimator of the target parameter which typically will converge at the fastest possible 1/root(n) rate and be approximately unbiased and normal, and from which valid confidence intervals for these parameters of interest may be constructed. The resulting method thus could be called a "double ML" method because it relies on estimating primary and auxiliary predictive models. In order to avoid overfitting, our construction also makes use of the K-fold sample splitting, which we call cross-fitting. This allows us to use a very broad set of ML predictive methods in solving the auxiliary and main prediction problems, such as random forest, lasso, ridge, deep neural nets, boosted trees, as well as various hybrids and aggregators of these methods.
appendix boundary found by appendix_titled_section at “Appendix: Proofs of Results” · 70% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Belloni, A., D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse models and methods for optimal instruments with an application to eminent domain | 1.000 | 8 | 3 | 100% |
| 2 | van der Vaart, A. W (1998) Asymptotic Statistics | 1.000 | 5 | 3 | 100% |
| 3 | Newey, W (1994) The asymptotic variance of semiparametric estimators self | 0.874 | 9 | 2 | 100% |
| 4 | Belloni, A., V. Chernozhukov, and C. Hansen (2014) Inference on treatment effects after selection amongst high-dimensional controls | 0.843 | 3 | 3 | 100% |
| 5 | Athey, S., G. Imbens, and S. Wager (2016) Approximate residual balancing: De-biased inference of average treatment effects in high-dimensions | 0.811 | 4 | 2 | 100% |
| 6 | Belloni, A., V. Chernozhukov, I. Fernández-Val, and C. Hansen (2017) Program evaluation with high-dimensional data | 0.811 | 4 | 2 | 100% |
| 7 | Neyman, J (1959) Optimal asymptotic tests of composite statistical hypotheses | 0.811 | 4 | 2 | 100% |
| 8 | Farrell, M (2015) Robust inference on average treatment effects with possibly more covariates than observations | 0.737 | 3 | 2 | 100% |
| 9 | Belloni, A., V. Chernozhukov, and L. Wang (2014) Pivotal estimation via square-root lasso in nonparametric regression | 0.737 | 3 | 2 | 100% |
| 10 | Zhang, C. and S. Zhang (2014) Confidence intervals for low-dimensional parameters with high-dimensional data | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 91 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.