Sven Klaassen, Jan Rabenseifner, Jannis Kueck, Philipp Bach
arXiv 21 Mar 2025 · Statistics — Machine Learning
arXiv:2503.17290 · PDF · DOI · OpenAlex · Extracted main text
The partitioning of data for estimation and calibration critically impacts the performance of propensity score based estimators like inverse probability weighting (IPW) and double/debiased machine learning (DML) frameworks. We extend recent advances in calibration techniques for propensity score estimation, improving the robustness of propensity scores in challenging settings such as limited overlap, small sample sizes, or unbalanced data. Our contributions are twofold: First, we provide a theoretical analysis of the properties of calibrated estimators in the context of DML. To this end, we refine existing calibration frameworks for propensity score models, with a particular emphasis on the role of sample-splitting schemes in ensuring valid causal inference. Second, through extensive simulations, we show that calibration reduces variance of inverse-based propensity score estimators while also mitigating bias in IPW, even in small-sample regimes. Notably, calibration improves stability for flexible learners (e.g., gradient boosting) while preserving the doubly robust properties of DML. A key insight is that, even when methods perform well without calibration, incorporating a calibration step does not degrade performance, provided that an appropriate sample-splitting approach is chosen.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | L. van der Laan, Z. Lin, M. Carone, and A. Luedtke (2024) Stabilized inverse probability weighting via isotonic calibration, 2024a | 1.000 | 9 | 4 | 100% |
| 2 | S. Deshpande and V. Kuleshov (2023) Calibrated propensity scores for causal effect estimation | 1.000 | 5 | 3 | 100% |
| 3 | V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W.… (2018) Double/debiased machine learning for treatment and structural parameters | 0.874 | 12 | 2 | 100% |
| 4 | D. Ballinari (2024) Calibrating doubly-robust estimators with unbalanced treatment assignment, 2024 | 0.874 | 5 | 2 | 100% |
| 5 | L. van der Laan, E. Ulloa-Perez, M. Carone, and A. Luedtke (2023) Causal isotonic calibration for heterogeneous treatment effects | 0.843 | 3 | 3 | 100% |
| 6 | R. Gutman, E. Karavani, and Y. Shimoni (2024) Improving inverse probability weighting by post-calibrating its propensity scores | 0.644 | 2 | 2 | 100% |
| 7 | E. Mammen and K. Yu (2007) Additive isotone regression | 0.644 | 2 | 2 | 100% |
| 8 | M. V. Wüthrich and J. Ziegel (2023) Isotonic recalibration under a low signal-to-noise ratio | 0.644 | 2 | 2 | 100% |
| 9 | D. Ballinari and N. Bearth (2025) Improving the finite sample estimation of average treatment effects using double/debiased machine learning with propensity score… | 0.585 | 3 | 1 | 100% |
| 10 | C. Gupta and A. K. Ramdas (2021) Distribution-free calibration guarantees for histogram binning without sample splitting, 2021 | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 47 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Semiparametric inference for impulse response functions using double/debiased machine learning | 0.405 | 1 | 1 |
| 2 | Sensitivity Analysis for Treatment Effects in Difference-in-Differences Models using Riesz Representation | 0.405 | 1 | 1 |
| 3 | Calibeating Prediction-Powered Inference | 0.405 | 1 | 1 |