EconBase
← All papers

Smaller Confidence Intervals From IPW Estimators via Data-Dependent Coarsening

Alkis Kalavasis, Anay Mehrotra, Manolis Zampetakis

arXiv 2 Oct 2024 · Statistics — Methodology

arXiv:2410.01658 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Inverse propensity-score weighted (IPW) estimators are prevalent in causal inference for estimating average treatment effects in observational studies. Under unconfoundedness, given accurate propensity scores and $n$ samples, the size of confidence intervals of IPW estimators scales down with $n$, and, several of their variants improve the rate of scaling. However, neither IPW estimators nor their variants are robust to inaccuracies: even if a single covariate has an $\varepsilon>0$ additive error in the propensity score, the size of confidence intervals of these estimators can increase arbitrarily. Moreover, even without errors, the rate with which the confidence intervals of these estimators go to zero with $n$ can be arbitrarily slow in the presence of extreme propensity scores (those close to 0 or 1). We introduce a family of Coarse IPW (CIPW) estimators that captures existing IPW estimators and their variants. Each CIPW estimator is an IPW estimator on a coarsened covariate space, where certain covariates are merged. Under mild assumptions, e.g., Lipschitzness in expected outcomes and sparsity of extreme propensity scores, we give an efficient algorithm to find a robust estimator: given $\varepsilon$-inaccurate propensity scores and $n$ samples, its confidence interval size scales with $\varepsilon+1/\sqrt{n}$. In contrast, under the same assumptions, existing estimators' confidence interval sizes are $\Omega(1)$ irrespective of $\varepsilon$ and $n$. Crucially, our estimator is data-dependent and we show that no data-independent CIPW estimator can be robust to inaccuracies.

Citation extraction

48
references
76
in-text mentions
48
distinct cited
0
self-citations
28,248
main-text words

appendix boundary found by appendix_command · 91% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hernan, Miguel A., Robins, James M (2023) Causal Inference: What If0.92843100%
2Imbens, Guido W., Rubin, Donald B (2015) Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction0.8947471%
3Wager, Stefan, Athey, Susan (2018) Estimation and Inference of Heterogeneous Treatment Effects using Random Forests0.81142100%
4McCaffrey, Daniel F., Ridgeway, Greg, Morral, Andrew R (2004) Propensity Score Estimation With Boosted Regression for Evaluating Causal Effects in Observational Studies0.73732100%
5Wager, Stefan (2020) STATS 361: Causal Inference0.73732100%
6Foster, Dylan J., Syrgkanis, Vasilis (2023) Orthogonal Statistical Learning0.6444250%
7Hirano, Keisuke, Imbens, Guido W., Ridder, Geert (2003) Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score0.64422100%
8Shalev-Shwartz, Shai, Ben-David, Shai (2014) Understanding Machine Learning: From Theory to Algorithms0.64422100%
9Chernozhukov, Victor, Chetverikov, Denis, Demirer, Mert, Duflo, Esth… (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters0.5113233%
10Bang, Heejung, Robins, James M (2005) Doubly Robust Estimation in Missing Data and Causal Inference Models0.5113233%

Showing the top 10 of 48 scored citations.