Alicia Curth, Mihaela van der Schaar
arXiv 6 Feb 2023 · Statistics — Machine Learning · 1 citations (OpenAlex)
arXiv:2302.02923 · PDF · DOI · OpenAlex · Extracted main text
Personalized treatment effect estimates are often of interest in high-stakes applications -- thus, before deploying a model estimating such effects in practice, one needs to be sure that the best candidate from the ever-growing machine learning toolbox for this task was chosen. Unfortunately, due to the absence of counterfactual information in practice, it is usually not possible to rely on standard validation metrics for doing so, leading to a well-known model selection dilemma in the treatment effect estimation literature. While some solutions have recently been investigated, systematic understanding of the strengths and weaknesses of different model selection criteria is still lacking. In this paper, instead of attempting to declare a global `winner', we therefore empirically investigate success- and failure modes of different selection criteria. We highlight that there is a complex interplay between selection strategies, candidate estimators and the data used for comparing them, and provide interesting insights into the relative (dis)advantages of different criteria alongside desiderata for the design of further illuminating empirical studies in this context.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Schuler, A., Baiocchi, M., Tibshirani, R., and Shah, N (2018) A comparison of methods for model selection when estimating individual treatment effects | 0.950 | 7 | 5 | 86% |
| 2 | Mahajan, D., Mitliagkas, I., Neal, B., and Syrgkanis, V (2022) Empirical analysis of model selection for heterogenous causal effect estimation | 0.950 | 7 | 4 | 86% |
| 3 | Künzel, S. R., Sekhon, J. S., Bickel, P. J., and Yu, B (2019) Metalearners for estimating heterogeneous treatment effects using machine learning | 0.928 | 5 | 3 | 80% |
| 4 | Kennedy, E. H (2020) Optimal doubly robust estimation of heterogeneous causal effects | 0.909 | 8 | 5 | 75% |
| 5 | Alaa, A. and Van Der Schaar, M (2019) Validating causal inference models via influence functions | 0.894 | 7 | 5 | 71% |
| 6 | Curth, A. and van der Schaar, M (1810) Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms self | 0.888 | 10 | 5 | 70% |
| 7 | Nie, X. and Wager, S (2021) Quasi-oracle estimation of heterogeneous treatment effects | 0.874 | 9 | 5 | 67% |
| 8 | Saito, Y. and Yasui, S (2020) Counterfactual cross-validation: Stable model selection procedure for causal inference models | 0.843 | 4 | 4 | 75% |
| 9 | Shalit, U., Johansson, F. D., and Sontag, D (2017) Estimating individual treatment effect: generalization bounds and algorithms | 0.843 | 5 | 4 | 60% |
| 10 | Curth, A. and van der Schaar, M (2021) On inductive biases for heterogeneous treatment effect estimation self | 0.737 | 5 | 3 | 40% |
Showing the top 10 of 40 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Unveiling the Potential of Robustness in Selecting Conditional Average Treatment Effect Estimators | 1.000 | 10 | 4 |
| 2 | 2310.16945 | 0.511 | 2 | 2 |