Jikai Jin, Lester Mackey, Vasilis Syrgkanis
arXiv 3 Jul 2025 · Statistics — Machine Learning
arXiv:2507.02275 · PDF · Extracted main text
Structure-agnostic causal inference studies how well one can estimate a treatment effect given black-box machine learning estimates of nuisance functions (like the impact of confounders on treatment and outcomes). Here, we find that the answer depends in a surprising way on the distribution of the treatment noise. Focusing on the partially linear model of \citet{robinson1988root}, we first show that the widely adopted double machine learning (DML) estimator is minimax rate-optimal for Gaussian treatment noise, resolving an open problem of \citet{mackey2018orthogonal}. Meanwhile, for independent non-Gaussian treatment noise, we show that DML is always suboptimal by constructing new practical procedures with higher-order robustness to nuisance errors. These ACE procedures use structure-agnostic cumulant estimators to achieve $r$-th order insensitivity to nuisance errors whenever the $(r+1)$-st treatment cumulant is non-zero. We complement these core results with novel minimax guarantees for binary treatments in the partially linear model. Finally, using synthetic demand estimation experiments, we demonstrate the practical benefits of our higher-order robust estimators.
appendix boundary found by appendix_command · 26% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sivaraman Balakrishnan, Edward H Kennedy, and Larry Wasserman (2023) The fundamental limits of structure-agnostic functional estimation | 0.950 | 7 | 4 | 86% |
| 2 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning | 0.941 | 6 | 5 | 83% |
| 3 | Jikai Jin and Vasilis Syrgkanis (2024) Structure-agnostic optimality of doubly robust learning for treatment effect estimation self | 0.874 | 6 | 5 | 67% |
| 4 | Lester Mackey, Vasilis Syrgkanis, and Ilias Zadik (2018) Orthogonal machine learning: Power and limitations self | 0.805 | 23 | 8 | 52% |
| 5 | Victor Chernozhukov, Whitney K Newey, and Rahul Singh (2022) Automatic debiased machine learning of causal and structural effects | 0.644 | 3 | 2 | 67% |
| 6 | Peter M Robinson (1988) Root-n-consistent semiparametric regression | 0.644 | 2 | 2 | 100% |
| 7 | T Tony Cai and Mark G Low (2011) Testing composite hypotheses, hermite polynomials and optimal estimation of a nonsmooth functional | 0.511 | 4 | 2 | 25% |
| 8 | Rick Durrett (2019) Probability: theory and examples, volume 49 | 0.511 | 2 | 2 | 50% |
| 9 | Charles J Stone (1982) Optimal global rates of convergence for nonparametric regression | 0.511 | 2 | 2 | 50% |
| 10 | Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2011) Inference for high-dimensional sparse econometric models | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 32 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Learning bounds for doubly-robust covariate shift adaptation | 0.644 | 2 | 2 |
| 2 | On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models | 0.405 | 1 | 1 |