Dylan J. Foster, Vasilis Syrgkanis
arXiv 25 Jan 2019 · Mathematics — Statistics Theory
arXiv:1901.09036 · PDF · Extracted main text
We provide non-asymptotic excess risk guarantees for statistical learning in a setting where the population risk with respect to which we evaluate the target parameter depends on an unknown nuisance parameter that must be estimated from data. We analyze a two-stage sample splitting meta-algorithm that takes as input arbitrary estimation algorithms for the target parameter and nuisance parameter. We show that if the population risk satisfies a condition called Neyman orthogonality, the impact of the nuisance estimation error on the excess risk bound achieved by the meta-algorithm is of second order. Our theorem is agnostic to the particular algorithms used for the target and nuisance and only makes an assumption on their individual performance. This enables the use of a plethora of existing results from machine learning to give new guarantees for learning with a nuisance component. Moreover, by focusing on excess risk rather than parameter estimation, we can provide rates under weaker assumptions than in previous works and accommodate settings in which the target parameter belongs to a complex nonparametric class. We provide conditions on the metric entropy of the nuisance and target classes such that oracle rates of the same order as if we knew the nuisance parameter are achieved.
appendix boundary found by appendix_command · 34% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | M. R. Kosorok (2008) Introduction to empirical processes and semiparametric inference | 0.928 | 5 | 4 | 80% |
| 2 | P. L. Bartlett, O. Bousquet, and S. Mendelson (2005) Local rademacher complexities | 0.928 | 4 | 3 | 100% |
| 3 | P. M. Robinson (1988) Root-n-consistent semiparametric regression | 0.928 | 4 | 3 | 100% |
| 4 | V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W.… (2018) Double/debiased machine learning for treatment and structural parameters | 0.885 | 13 | 6 | 69% |
| 5 | J. M. Robins and A. Rotnitzky (1995) Semiparametric efficiency in multivariate regression models with missing data | 0.874 | 6 | 4 | 67% |
| 6 | V. Chernozhukov, J. C. Escanciano, H. Ichimura, W. K. Newey, and J.… Locally robust semiparametric estimation | 0.874 | 6 | 2 | 100% |
| 7 | M. J. van der Laan and S. Dudoit (2003) Unified cross-validation methodology for selection among estimators and a general cross-validated adaptive epsilon-net estimator… | 0.874 | 6 | 2 | 100% |
| 8 | V. Chernozhukov, M. Goldman, V. Semenova, and M. Taddy (2017) Orthogonal machine learning for demand estimation: High dimensional causal inference in dynamic panels | 0.874 | 5 | 2 | 100% |
| 9 | V. Chernozhukov, D. Nekipelov, V. Semenova, and V. Syrgkanis (1806) Plug-in regularized estimation of high-dimensional parameters in nonlinear semiparametric models | 0.863 | 14 | 7 | 64% |
| 10 | A. Maurer and M. Pontil (2009) Empirical Bernstein bounds and sample variance penalization | 0.843 | 4 | 3 | 75% |
Showing the top 10 of 135 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.