EconBase
← All papers

Calibration Strategies for Robust Causal Estimation: Theoretical and Empirical Insights on Propensity Score-Based Estimators

Sven Klaassen, Jan Rabenseifner, Jannis Kueck, Philipp Bach

arXiv 21 Mar 2025 · Statistics — Machine Learning

arXiv:2503.17290 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The partitioning of data for estimation and calibration critically impacts the performance of propensity score based estimators like inverse probability weighting (IPW) and double/debiased machine learning (DML) frameworks. We extend recent advances in calibration techniques for propensity score estimation, improving the robustness of propensity scores in challenging settings such as limited overlap, small sample sizes, or unbalanced data. Our contributions are twofold: First, we provide a theoretical analysis of the properties of calibrated estimators in the context of DML. To this end, we refine existing calibration frameworks for propensity score models, with a particular emphasis on the role of sample-splitting schemes in ensuring valid causal inference. Second, through extensive simulations, we show that calibration reduces variance of inverse-based propensity score estimators while also mitigating bias in IPW, even in small-sample regimes. Notably, calibration improves stability for flexible learners (e.g., gradient boosting) while preserving the doubly robust properties of DML. A key insight is that, even when methods perform well without calibration, incorporating a calibration step does not degrade performance, provided that an appropriate sample-splitting approach is chosen.

Citation extraction

47
references
87
in-text mentions
47
distinct cited
0
self-citations
11,262
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1L. van der Laan, Z. Lin, M. Carone, and A. Luedtke (2024) Stabilized inverse probability weighting via isotonic calibration, 2024a1.00094100%
2S. Deshpande and V. Kuleshov (2023) Calibrated propensity scores for causal effect estimation1.00053100%
3V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W.… (2018) Double/debiased machine learning for treatment and structural parameters0.874122100%
4D. Ballinari (2024) Calibrating doubly-robust estimators with unbalanced treatment assignment, 20240.87452100%
5L. van der Laan, E. Ulloa-Perez, M. Carone, and A. Luedtke (2023) Causal isotonic calibration for heterogeneous treatment effects0.84333100%
6R. Gutman, E. Karavani, and Y. Shimoni (2024) Improving inverse probability weighting by post-calibrating its propensity scores0.64422100%
7E. Mammen and K. Yu (2007) Additive isotone regression0.64422100%
8M. V. Wüthrich and J. Ziegel (2023) Isotonic recalibration under a low signal-to-noise ratio0.64422100%
9D. Ballinari and N. Bearth (2025) Improving the finite sample estimation of average treatment effects using double/debiased machine learning with propensity score…0.58531100%
10C. Gupta and A. K. Ramdas (2021) Distribution-free calibration guarantees for histogram binning without sample splitting, 20210.51121100%

Showing the top 10 of 47 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Semiparametric inference for impulse response functions using double/debiased machine learning0.40511
2Sensitivity Analysis for Treatment Effects in Difference-in-Differences Models using Riesz Representation0.40511
3Calibeating Prediction-Powered Inference0.40511