EconBase
← All papers

A New Central Limit Theorem for the Augmented IPW Estimator: Variance Inflation, Cross-Fit Covariance and Beyond

Kuanhao Jiang, Rajarshi Mukherjee, Subhabrata Sen, Pragya Sur

arXiv 20 May 2022 · Mathematics — Statistics Theory · publishedThe Annals of Statistics (2025) · 4 citations (OpenAlex)

arXiv:2205.10198 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Estimation of the average treatment effect (ATE) is a central problem in causal inference. In recent times, inference for the ATE in the presence of high-dimensional covariates has been extensively studied. Among the diverse approaches that have been proposed, augmented inverse probability weighting (AIPW) with cross-fitting has emerged a popular choice in practice. In this work, we study this cross-fit AIPW estimator under well-specified outcome regression and propensity score models in a high-dimensional regime where the number of features and samples are both large and comparable. Under assumptions on the covariate distribution, we establish a new central limit theorem for the suitably scaled cross-fit AIPW that applies without any sparsity assumptions on the underlying high-dimensional parameters. Our CLT uncovers two crucial phenomena among others: (i) the AIPW exhibits a substantial variance inflation that can be precisely quantified in terms of the signal-to-noise ratio and other problem parameters, (ii) the asymptotic covariance between the pre-cross-fit estimators is non-negligible even on the root-n scale. These findings are strikingly different from their classical counterparts. On the technical front, our work utilizes a novel interplay between three distinct tools--approximate message passing theory, the theory of deterministic equivalents, and the leave-one-out approach. We believe our proof techniques should be useful for analyzing other two-stage estimators in this high-dimensional regime. Finally, we complement our theoretical results with simulations that demonstrate both the finite sample efficacy of our CLT and its robustness to our assumptions.

Citation extraction

127
references
244
in-text mentions
127
distinct cited
3
self-citations
13,757
main-text words

appendix boundary found by appendix_titled_section at “Supplementary material” · 17% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1barticle[author] Chernozhukov, VictorV., Chetverikov, DenisD., Demir… (2017) )1.000155100%
2barticle[author] Smucler, EzequielE., Rotnitzky, AndreaA. Robins, Ja… (2019) )1.000104100%
3barticle[author] Bang, HeejungH. Robins, James MJ. M (2005) )1.00083100%
4barticle[author] Sur, PragyaP. Candès, Emmanuel JE. J (2019) )1.00074100%
5barticle[author] Scharfstein, Daniel OD. O., Rotnitzky, AndreaA. Rob… (1999) )0.87462100%
6barticle[author] Tan, ZhiqiangZ (2020) )0.87452100%
7barticle[author] Donoho, David LD. L., Maleki, ArianA. Montanari, An… (2009) )0.84333100%
8barticle[author] Bean, DerekD., Bickel, Peter JP. J., El Karoui, Nou… (2013) )0.81142100%
9barticle[author] El Karoui, NoureddineN., Bean, DerekD., Bickel, Pet… (2013) )0.81142100%
10barticle[author] El Karoui, NoureddineN (2018) )0.81142100%

Showing the top 10 of 127 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Method-of-Moments Inference for GLMs and Doubly Robust Functionals under Proportional Asymptotics0.73732
2Assumption-lean Falsification Tests of Rate Double-Robustness of Double-Machine-Learning Estimators0.40511
3Decomposition of Spillover Effects Under Misspecification: Pseudo-True Estimands and a Local-Global Extension0.40511