Lester Mackey, Vasilis Syrgkanis, Ilias Zadik
arXiv 1 Nov 2017 · Machine Learning · 8 citations (OpenAlex)
arXiv:1711.00342 · PDF · DOI · OpenAlex · Extracted main text
Double machine learning provides $\sqrt{n}$-consistent estimates of parameters of interest even when high-dimensional or nonparametric nuisance parameters are estimated at an $n^{-1/4}$ rate. The key is to employ Neyman-orthogonal moment equations which are first-order insensitive to perturbations in the nuisance parameters. We show that the $n^{-1/4}$ requirement can be improved to $n^{-1/(2k+2)}$ by employing a $k$-th order notion of orthogonality that grants robustness to more complex or higher-dimensional nuisance parameters. In the partially linear regression setting popular in causal inference, we show that we can construct second-order orthogonal moments if and only if the treatment residual is not normally distributed. Our proof relies on Stein's lemma and may be of independent interest. We conclude by demonstrating the robustness benefits of an explicit doubly-orthogonal estimation procedure for treatment effect.
appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2017) Double/debiased/neyman machine learning of treatment effects | 0.979 | 16 | 7 | 94% |
| 2 | Belloni, A., Chernozhukov, V., Val, I. F., and Hansen, C Program evaluation and causal inference with high dimensional data | 0.405 | 1 | 1 | 100% |
| 3 | Berk, R., Brown, L., Buja, A., Zhang, K., and Zhao, L (2013) Valid post-selection inference | 0.405 | 1 | 1 | 100% |
| 4 | Javanmard, A. and Montanari, A (2015) De-biasing the Lasso: Optimal Sample Size for Gaussian Designs | 0.405 | 1 | 1 | 100% |
| 5 | Neyman, J (1961) C(α) tests and their use | 0.405 | 1 | 1 | 100% |
| 6 | Tibshirani, R. J., Taylor, J., Lockhart, R., and Tibshirani, R (2016) Exact post-selection inference for sequential regression procedures | 0.405 | 1 | 1 | 100% |
| 7 | Zhang, C. H. and Zhang, S Confidence intervals for low dimensional parameters in high dimensional linear models | 0.405 | 1 | 1 | 100% |
| 8 | van de Geer, S., Buhlmann, P., Ritov, Y., and Dezeure, R (2014) On asymptotically optimal confidence regions and tests for high-dimensional models | 0.405 | 1 | 1 | 100% |
| 9 | Newey, W. and McFadden, D.l (1994) Chapter 36 large sample estimation and hypothesis testing | 0.000 | 4 | 2 | 0% |
| 10 | Flanders, H (1973) Differentiation under the integral sign | 0.000 | 4 | 1 | 0% |
Showing the top 10 of 14 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.