David Bruns-Smith, Oliver Dukes, Avi Feller, Elizabeth L. Ogburn
arXiv 27 Apr 2023 · Statistics — Methodology · publishedJournal of the Royal Statistical Society Series B (Statistical Methodology) (2025) · 6 citations (OpenAlex)
arXiv:2304.14545 · PDF · DOI · OpenAlex · Extracted main text
We provide a novel characterization of augmented balancing weights, also known as automatic debiased machine learning (AutoDML). These popular doubly robust or de-biased machine learning estimators combine outcome modeling with balancing weights - weights that achieve covariate balance directly in lieu of estimating and inverting the propensity score. When the outcome and weighting models are both linear in some (possibly infinite) basis, we show that the augmented estimator is equivalent to a single linear model with coefficients that combine the coefficients from the original outcome model and coefficients from an unpenalized ordinary least squares (OLS) fit on the same data. We see that, under certain choices of regularization parameters, the augmented estimator often collapses to the OLS estimator alone; this occurs for example in a re-analysis of the Lalonde 1986 dataset. We then extend these results to specific choices of outcome and weighting models. We first show that the augmented estimator that uses (kernel) ridge regression for both outcome and weighting models is equivalent to a single, undersmoothed (kernel) ridge regression. This holds numerically in finite samples and lays the groundwork for a novel analysis of undersmoothing and asymptotic rates of convergence. When the weighting model is instead lasso-penalized regression, we give closed-form expressions for special cases and demonstrate a “double selection” property. Our framework opens the black box on this increasingly popular class of estimators, bridges the gap between existing results on the semiparametric efficiency of undersmoothed and doubly robust estimators, and provides new insights into the performance of augmented balancing weights.
appendix boundary found by appendix_command · 43% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | D. A. Hirshberg and S. Wager (2021) Augmented minimax linear estimation | 1.000 | 6 | 4 | 100% |
| 2 | R. Singh, L. Xu, and A. Gretton (2020) Kernel methods for causal functions: Dose, heterogeneous, and incremental response curves | 0.928 | 4 | 3 | 100% |
| 3 | V. Chernozhukov, W. K. Newey, and R. Singh (2022) Automatic debiased machine learning of causal and structural effects | 0.914 | 17 | 9 | 76% |
| 4 | J. Robins, M. Sued, Q. Lei-Gomez, and A. Rotnitzky (2007) Comment: Performance of double-robust estimators when" inverse probability" weights are highly variable | 0.899 | 11 | 7 | 73% |
| 5 | E. Ben-Michael, A. Feller, D. A. Hirshberg, and J. R. Zubizarreta (2021) The balancing act in causal inference | 0.894 | 7 | 5 | 71% |
| 6 | S. Athey, G. W. Imbens, and S. Wager (2018) Approximate residual balancing: debiased inference of average treatment effects in high dimensions | 0.855 | 8 | 4 | 62% |
| 7 | P. Kline (2011) Oaxaca-blinder as a reweighting estimator | 0.843 | 4 | 4 | 75% |
| 8 | R. J. LaLonde (1986) Evaluating the econometric evaluations of training programs with experimental data | 0.843 | 15 | 5 | 60% |
| 9 | J. R. Zubizarreta (2015) Stable weights that balance covariates for estimation with incomplete outcome data | 0.843 | 5 | 4 | 60% |
| 10 | J.-C. Deville and C.-E. Särndal (1992) Calibration estimators in survey sampling | 0.843 | 5 | 3 | 60% |
Showing the top 10 of 103 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.