arXiv 8 Dec 2025 · Statistics ā Methodology
arXiv:2512.07083 · PDF · Extracted main text
Double Machine Learning is often justified by nuisance-rate conditions, yet finite-sample reliability also depends on the conditioning of the orthogonal-score Jacobian. This conditioning is typically assumed rather than tracked. When residualized treatment variance is small, the Jacobian is ill-conditioned and small systematic nuisance errors can be amplified, so nominal confidence intervals may look precise yet systematically under-cover. Our main result is an exact identity for the cross-fitted PLR-DML estimator, with no Taylor approximation. From this identity, we derive a stochastic-order bound that separates oracle noise from a conditioning-amplified nuisance remainder and yields a sufficiency condition for root-n-inference. We further connect the amplification factor to semiparametric efficiency geometry via the Riesz representer and use a triangular-array framework to characterize regimes as residual treatment variation weakens. These results motivate an out-of-fold diagnostic that summarizes the implied amplification scale. We do not propose universal thresholds. Instead, we recommend reporting the diagnostic alongside cross-learner sensitivity summaries as a fragility assessment, illustrated in simulation and an empirical example.
appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C⦠(2018) Double/debiased machine learning for treatment and structural parameters | 1.000 | 8 | 3 | 100% |
| 2 | D'Amour, A., Ding, P., Feller, A., Lei, L., & Sekhon, J (2021) Overlap in observational studies with high-dimensional covariates | 0.843 | 3 | 3 | 100% |
| 3 | Chernozhukov, V., Newey, W. K., & Singh, R (2022) Automatic debiased machine learning of causal and structural effects | 0.737 | 3 | 2 | 100% |
| 4 | Kennedy, E. H (2024) Semiparametric doubly robust targeted double machine learning: A review | 0.737 | 3 | 2 | 100% |
| 5 | Newey, W. K (1990) Semiparametric efficiency bounds | 0.737 | 3 | 2 | 100% |
| 6 | Robinson, P. M (1988) Root-$N$-consistent semiparametric regression | 0.737 | 3 | 2 | 100% |
| 7 | Belsley, D. A., Kuh, E., & Welsch, R. E (1980) Regression Diagnostics: Identifying Influential Data and Sources of Collinearity | 0.644 | 2 | 2 | 100% |
| 8 | Golub, G. H., & Van Loan, C. F (2013) Matrix Computations | 0.644 | 2 | 2 | 100% |
| 9 | Khan, S., & Tamer, E (2010) Irregular identification, support conditions, and inverse weight estimation | 0.644 | 2 | 2 | 100% |
| 10 | Rosenbaum, P. R., & Rubin, D. B (1983) The central role of the propensity score in observational studies for causal effects | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 32 scored citations.