EconBase
← All papers

Causal Q-Aggregation for CATE Model Selection

Hui Lan, Vasilis Syrgkanis

arXiv 25 Oct 2023 · Statistics — Machine Learning

arXiv:2310.16945 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Accurate estimation of conditional average treatment effects (CATE) is at the core of personalized decision making. While there is a plethora of models for CATE estimation, model selection is a nontrivial task, due to the fundamental problem of causal inference. Recent empirical work provides evidence in favor of proxy loss metrics with double robust properties and in favor of model ensembling. However, theoretical understanding is lacking. Direct application of prior theoretical work leads to suboptimal oracle model selection rates due to the non-convexity of the model selection problem. We provide regret rates for the major existing CATE ensembling approaches and propose a new CATE model ensembling approach based on Q-aggregation using the doubly robust loss. Our main result shows that causal Q-aggregation achieves statistically optimal oracle model selection regret rates of $\frac{\log(M)}{n}$ (with $M$ models and $n$ samples), with the addition of higher-order estimation error terms related to products of errors in the nuisance functions. Crucially, our regret rate does not require that any of the candidate CATE models be close to the truth. We validate our new method on many semi-synthetic datasets and also provide extensions of our work to CATE model selection with instrumental variables and unobserved confounding.

Citation extraction

46
references
119
in-text mentions
46
distinct cited
6
self-citations
7,852
main-text words

appendix boundary found by appendix_command · 27% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kennedy, E. H (2020) Towards optimal doubly robust estimation of heterogeneous causal effects0.9285380%
2Nie, X. and Wager, S (2021) Quasi-oracle estimation of heterogeneous treatment effects0.8558362%
3Han, K. W. and Wu, H (2022) Ensemble method for estimating individualized treatment effects0.8434375%
4Lecué, G. and Rigollet, P (2014) Optimal learning with Q-aggregation0.7946450%
5Foster, D. J. and Syrgkanis, V (2023) Orthogonal statistical learning self0.77817847%
6Alaa, A. and Van Der Schaar, M (2019) Validating causal inference models via influence functions0.7374275%
7Mahajan, D., Mitliagkas, I., Neal, B., and Syrgkanis, V (2022) Empirical analysis of model selection for heterogenous causal effect estimation self0.6444250%
8Poterba, J. M. and Venti, S. F (1994) 401 (k) plans and tax-deferred saving0.64422100%
9Poterba, J. M., Venti, S. F., and Wise, D. A (1995) Do 401 (k) contributions crowd out other personal saving?0.64422100%
10Word, E. et al (1990) Student/teacher achievement ratio (star) tennessee's k-3 class size study. final summary report 1985-19900.64422100%

Showing the top 10 of 46 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Reevaluating Causal Estimation Methods with Data from a Product Release0.40511